Murtaza Nikzad · August 13, 2026

The shopkeeper who refuses your money

A note on the writing: I have used AI to edit this blog post.

In Kabul, where I grew up, there is a move every shopkeeper and every buyer knows. You hold out your money for the groceries that you just bought. But the grocer waves it away by saying «اصلا قابل شما ره نداره» - “it’s not worthy of you at all.” Read literally, he is refusing payment. However, that is not the case at all. You as the buyer are absolutely and certainly aware that he expects to be paid. The correct response here is to insist until he takes the money. Now, it is neither the shopkeeper nor the buyer who is lying about anything. Something is being performed in this social interaction. The performance translates to saying our relationship is worth more than this monetary transaction, and you as the buyer are in a way agreeing with him by paying the money.

Now this is not a case of exotic customs from other cultures. The interesting thing about language is that every culture has their own linguistic quirks. In the US, “Let’s grab coffee sometime” doesn’t really translate to a proposal for drinking coffee. “Oh, you shouldn’t have!” in reality means that they are delighted that you did what you did. And you should in fact have done it. My point is that every culture has these linguistic quirks.

In this writing I will focus on Dari/Farsi’s. If a shopkeeper offers you tea, the polite response is to refuse the first offer. The shopkeeper’s job is to insist and yours is to accept only on the second or third round. In a similar scenario, a man who has owed money for a year comes to the lender’s shop and says “I’m ashamed, khalifa, my face is black - I’ll have your money by next week.” In this interaction, “my face is black” cannot be translated literally because it really is not about the person in question’s face. Rather it is a ritualized performance of shame. It signals a sort of social payment, although the monetary payment is still late.

Dari/Farsi ritualizes offer and refusal so thoroughly that there is a name for this pattern of social interaction and it is called taarof.

Now why am I sharing this? When a message crosses a language border through a machine, the words get translated well, however the intent and cultural nuances gets lost.

I wanted to know how bad this problem is. I wanted to do some empirical studies and don’t just rely on my musings.

So I built a benchmark

It’s called Sawda, the Dari word for trade. It is 60 short fictional messages from family-business communication, where each sentence also carries context about who is speaking to whom and what is going on in the social interaction.

I ran six configurations over the items, all from open weights models which include NLLB-600M, Kimi K3, and gpt-oss-20b. Here is exactly what each configuration was given:

SystemModelWhat it gets
ANLLB-200-600Mthe message alone (dedicated translation model)
BKimi K3”translate this”
CKimi K3plus who is writing to whom and the situation
DKimi K3plus explicit taarof instructions and worked examples
Egpt-oss-20b”translate this”
Fgpt-oss-20bplus who is writing to whom and the situation

For judging, I looked at each translation that was shown to me randomly with no information about which model produced it and I rated how much the intent survived. Each of the six configurations translated the same 30-item subset. The total cost of the whole pilot program was $3.83 in API calls.

What came back

Let’s look into one raw failure before any aggregate. In an imaginary scenario, a mechanic has just fixed a taxi’s brakes. The driver, who is a repeat customer, holds out the payment they agreed on that morning and the mechanic says the ritual thing: «چی پیسه خیر است جان برادر؟ موتر از خود ماست، اصلا قابل شما ره نداره!» — roughly, “what’s this money talk, dear brother? the car is practically ours, it’s not worthy of you!” He expects to be paid.

NLLB-600M translated it as: “What is wrong with us, John? We are dead, so you cannot escape.”

So the small models are practically useless in nuanced fa2en translation. However, the problem is that the naive Kimi prompt also tends to produce a translation that is buried inside paragraphs of its own commentary which makes it unusable for commercial uses.

In general, no system preserved the speaker’s intent even a third of the time. The best configuration - the frontier model with full cultural instructions and examples, fully preserved intent in 9 out of 30 judgements. Here is the full table of my 180 judgments:

SystemPreservedDegradedInverted or deletedAppropriateness (1–7)
A · NLLB-600M13%77%10%3.5
B · Kimi K3 naive20%67%13%3.4
C · Kimi K3 audience17%77%7%4.7
D · Kimi K3 instructed30%53%17%4.6
E · gpt-oss-20b naive20%67%13%4.7
F · gpt-oss-20b audience23%70%7%4.9

Each row is 30 blind judgments. Preserved means the intent fully landed for the reader. Degraded means the message arrived but the force did not. Inverted or deleted means the reader would do the wrong thing.

Two other findings surprised me more than this headline.

  1. Telling the model about taarof both helped and hurt at the same time. Full instruction nearly doubled full intent preservation over just naming the audience (9/30 vs 5/30) — and it also more than doubled the catastrophic failures, the inversions and deletions (5/30 vs 2/30). But when I looked at those five catastrophes one by one, none of them was the model overdoing the politeness. All five were broken Dari: wrong language, broken syntax, an incomplete translation. The heavy prompt helps where the model can actually write fluent Dari. When it can’t, the model falls apart. (A note of honesty: the sample size is small - because it was only me and myself who were generating and labeling and judging the data - and none of these paired differences reaches statistical significance.)

  2. I found an interesting asymmetry of translation that exists in en2fa (English to Farsi) and fa2en. Catastrophic failures are concentrated in the en2fa direction, which consists 15 of 90 judgements versus 5 of 90 in fa2en. The models right now are not even decent in writing the Farsi/Dari that is spoken and written in Afghanistan overall. This is a significant gap and there should be more push to investing in low-resources languages within the NLP community.

Appendix

GitHub (not public as of early August 2026): github.com/MurtazaKafka/sawda-bench. If you’re a native Dari or Farsi speaker and this blog interests you, I would love for you to be annotator two.

If you want to know about the engineering underneath it all check out this blog post: How I built a translation benchmark in three days for $3.83.