AI Keeps Replying in English to Spanish Customers
An AI step replies in the wrong language because the chat-completion APIs from OpenAI, Anthropic, and Google have no output-language parameter, so the model infers the language from the whole prompt, and in an automation the customer's own words are often under 2% of it next to an English system prompt, English knowledge-base text, and a quoted English thread. The reliable fix is to decide the reply language in code from a stored preference or a detector run on the customer's own text, pass it to the model as an explicit field, keep schema values out of translation, and check the language of the reply before it is sent.
Your AI replies in English to a Spanish-speaking customer because no major chat API has an output-language setting. Anthropic's Messages API, OpenAI's Chat Completions and Responses APIs, and Google's Gemini generateContent all leave the reply language to the model, which infers it from the entire prompt. In an automation, the customer's words are often a sliver of that prompt, surrounded by an English system prompt, English help-center passages, and a quoted English email thread, so the model follows the majority. The fix that holds is to decide the language in your own code, pass it as an explicit field, and check the reply before it goes out.
Every forum thread on this lands on the same advice: add "always respond in the user's language" to the system prompt. That works in a chat window, where the user's message is the only other thing in the conversation. It is a weak instruction in a pipeline, because it asks the model to work out who "the user" is and which language they used, from a prompt your workflow assembled out of five sources.
Why does the model pick the wrong language?
The model reads everything you send and writes in whatever language the context points to. Anthropic's documentation says it plainly: the model infers the response language from the conversation, and for production you should state the target language explicitly, with a user's runtime choice interpolated into the system prompt rather than inferred.
Now count what is actually in a typical support-reply prompt. These numbers are illustrative, from the shape of flows we build, not a measurement:
| Part of the prompt | Language | Rough tokens |
|---|---|---|
| System prompt with tone, policies, and format rules | English | 800 |
| Three retrieved help-center passages | English | 1,500 |
| Quoted thread: your last reply plus signature and footer | English | 400 |
| Tool result: the customer's next appointment from the CRM | English field names | 150 |
| The customer's new message: "Hola, pueden venir el jueves en lugar del martes? Gracias" | Spanish | 20 |
Twenty tokens out of about 2,870 is 0.7% of the prompt. Strip the quoted thread properly and it is still under 1%. The instruction "reply in the user's language" has to beat 99% of English context, every time, on every model version you will ever pin. Most days it does. The days it does not are the ones a customer screenshots.
Three other conditions make it worse:
- Short messages. "ok gracias", "si", "123 Main St", or a bare order number tell the model almost nothing. It falls back to the language of everything else, which is English.
- Mixed messages. A customer writes Spanish with an English product name or pastes an English error message. The model may treat the pasted part as the user's language.
- Agent tool loops. After two or three tool calls, the most recent content in the context is English JSON from your own systems, and the customer's message is far back.
Where can you actually set the language?
For text replies, nowhere in the API. For speech, you can. That difference matters if you run voice flows, because a wrong transcription language poisons every step after it.
| Provider and endpoint | Language control (as of October 2026) | What it does |
|---|---|---|
| Anthropic Messages | None. Prompt only. | Docs recommend an explicit language in the system prompt. |
| OpenAI Chat Completions, Responses | None. Prompt only. | Language follows the prompt. |
| Google Gemini generateContent | None. Prompt only. | Language follows the prompt. |
OpenAI whisper-1 transcription | language, one ISO-639-1 hint | Omit it and the model auto-detects. |
OpenAI gpt-transcribe | languages, a list of expected codes, such as ["en", "es"] | The response lists detected languages, and returns "languages": [] when it cannot make a reliable prediction. |
That gpt-transcribe response field is the most useful thing in the table. It is a language signal you can store and act on, which is exactly what the text APIs do not give you.
Decide the language in code, then tell the model
Treat the reply language as data your workflow owns, like the customer's time zone, not as something the model guesses per message. This is the order we resolve it in:
- A stored preference wins. A preferred-language field in the CRM, the language a customer picked on your web form, or the option they pressed in the phone menu. If it exists, use it and stop.
- Detect from the customer's own text only. Strip quoted replies, signatures, and disclaimers first; detecting on the raw email body measures your own footer. Run a language-ID library (fastText's lid.176 model recognizes 176 languages; lingua is another common choice) on what is left. Trust it only above about 20 characters and above a confidence threshold you tune on a few hundred of your real messages.
- Fall back to the last confident language for this contact or thread. A customer who wrote three Spanish emails and then sends "ok" still wants Spanish.
- Fall back to the business default only when there is no history at all.
- Write it back. When detection is confident and there was no stored preference, save it to the contact record, so step 1 handles the next message.
Then pass the decision as a hard field next to the record, for example reply_language: es and reply_locale: es-MX, with one line in the system prompt: write the customer-facing reply in the language given in reply_language, regardless of the language of any other text. Keep that field at the end of the prompt with the record, not baked into the system prompt per language. One system prompt for all languages keeps your cached prefix identical across customers, which matters if you rely on prompt caching in your automations.
Locale is worth the extra field. Mexican and Castilian Spanish differ in vocabulary, and Brazilian and European Portuguese differ more. If your customers are mostly in the US, say so: es-US or es-MX gets you the register they expect.
The schema trap: translated enum values
This is the failure that does not look like a language bug. You tell the model to reply in Spanish. It returns JSON with a reply_text in Spanish, which is right, and "intent": "reprogramar" instead of "reschedule", which is wrong. Your switch node has no branch for reprogramar, the record falls through to the default path, and the reschedule never happens. No error is logged anywhere.
The model did what you asked. The instruction to "reply in Spanish" leaked into fields that were never meant for a person to read. How much protection you get depends on the output mode:
| Output mode | Can an enum value come back translated? |
|---|---|
| OpenAI JSON mode | Yes. JSON mode guarantees valid JSON, not schema adherence. |
OpenAI Structured Outputs, strict: true | No. OpenAI states it prevents hallucinated invalid enum values. |
Anthropic output_config.format or strict: true tools | Enums are enforced, but Anthropic documents that capitalization is not guaranteed, so "Reschedule" can arrive with no error. |
Gemini response schema with enum | Enums are supported. Google's docs still tell you to validate values in your application. |
| Plain prompt, "return JSON" | Yes, and it will. |
Three rules close the gap. Use the strict mode your provider offers for every field your code branches on. Compare enum values case-insensitively. And split the output into fields by audience: reply_text in the customer's language, internal_summary and intent in English for your staff and your routing. Name them so the split is obvious to the model. If you have fought AI steps returning broken JSON, this is the same discipline applied to language.
Check the reply before it is sent
The last step is cheap and catches what the other steps miss. Run the same language detector on reply_text and compare it to reply_language:
- Match: send it.
- Mismatch on a reply over about 20 characters: regenerate once with the language field restated at the end of the prompt. If the second attempt also mismatches, route it to a person with both drafts attached. Do not loop.
- Very short reply: skip the check. "Gracias, Ana" will not detect reliably and does not need to.
Log the decision path on every message: where the language came from (stored, detected with its confidence, inherited, default), what was requested, and what the check found. When a customer complains, that line answers "why did they get English" in ten seconds instead of an afternoon. It also tells you whether your detector threshold is set right, because the inherited and default counts climb when it is too strict.
For email flows, the stripping in step 2 is its own problem with its own edge cases. We cover how to strip quoted text from email replies before anything reads them, and the language decision depends on doing that first.
What to do next
Pull the last 200 messages your AI step answered and run a language detector over the customer's text and the reply. Count the mismatches, then count how many of the incoming messages were under 20 characters or carried a quoted thread. That tells you whether your problem is detection, context weight, or both, and whether the reply_language field alone will fix it. If you are building a multilingual assistant or reply flow and want the language handling designed in from the start, that is part of how we build custom AI assistants and software, and you can tell us what your customers write in.
Frequently Asked Questions
SOURCES & CITATIONS
- Multilingual support — Anthropichttps://platform.claude.com/docs/en/build-with-claude/multilingual-support
- Structured outputs — Anthropichttps://platform.claude.com/docs/en/build-with-claude/structured-outputs
- Structured Outputs — OpenAIhttps://developers.openai.com/api/docs/guides/structured-outputs
- Speech to text — OpenAIhttps://developers.openai.com/api/docs/guides/speech-to-text
About Alexey Yushkin
Alexey is the founder of GENERAL INFORMATICS LLC. He designs and ships AI and automation systems for businesses and operators across the US.
Related reading
Want this kind of system in your business?
We build practical AI and automation systems for operators. Send us your current workflow and we will show you what to automate first.
