A local Qwen LLM suddenly switched from English to Chinese, leading me from language drift to Searle’s Chinese Room and a deeper question about AI understanding.
I was using qwen3.5:9b in Ollama like I use most local models: casually, confidently and with that dangerous developer habit of trusting something because it ✨worked✨.
The workflow was normal. I typed in English. The model replied in English. I asked questions, tested ideas, moved on.
Then, without warning, it suddenly switched to Chinese.
Not broken Chinese.
Not random symbols.
Chinese.
The answer looked fluent enough to make the failure feel worse. If it had crashed, I would have understood what kind of problem I was dealing with. If it had returned garbage, I could have blamed the model and moved on.
But this was different.
It was coherent. It was calm. It had the energy of a system that believed nothing unusual had happened.
That was the cursed part.
The model had not stopped working.
It had stopped respecting the conversation I thought we were having.
My first reaction was practical.
Of course Qwen can produce Chinese. It is trained with Chinese scripts. It is a multilingual model family. Chinese was not some alien object that entered the system from nowhere.
So part of me wanted to shrug and say, “Okay, that makes sense.”
But another part of me was still stuck on the social weirdness of it.
I had spoken English. The model had replied in English. Somewhere in that exchange, I assumed English had become the authority of the conversation.
The model apparently did not agree.
That is where I made the first mistake. I treated language like a shared agreement, because that is how language works with people. The model treated language like context.
Those are not the same thing.
To me, English was the room.
To the model, English was just a strong pattern until another pattern became strong enough.
The prompt tried to keep things steady. The context carried whatever it had accumulated. The model kept predicting the next token, then the next, then the next.
And somehow, Chinese became the path.
No drama.
Just drift.
After that, I did what everyone does when a machine behaves in a way that feels haunted.
I googled it.
That made the incident feel less personal, but not less interesting. I found that language drift was not only something people associated with raw local models. Similar behavior was discussed around larger hosted LLMs too, including systems with more visible polish and invisible guardrails.
That changed the shape of the problem.
At first, I thought I was looking at a Qwen problem.
Then I thought I was looking at an Ollama problem.
Then it started to look like an LLM problem.
Local models make the rawness easier to see. Hosted systems usually have more behavior-shaping around them: instructions, post-training, guardrails, product decisions, all the invisible middle management standing behind the model whispering, “Please don’t be weird.”
And most of the time, that works.
But “works most of the time” is not the same as understanding.
It is control.
It is alignment.
It is pressure.
That realization made the Chinese output feel less like a glitch and more like a leak in the abstraction.
The model was not simply answering me.
A whole system was trying to keep the answer inside a shape.
The practical fix was obvious enough.
Tell the model what language to use.
Something like:
Always respond in English unless I explicitly ask for another language. If another language appears in the conversation, translate or explain it in English.
Lower the temperature if the model keeps drifting.
Clear the context if the chat has become messy.
Be explicit instead of assuming the model knows the social contract.
That helps.
It really does.
But it also leaves behind a slightly annoying question:
Why did I expect the model to know that in the first place?
The answer is uncomfortable. I expected it because the model sounded like it understood me.
Fluency tricked me.
The model had been so good at producing the shape of cooperation that I started trusting it with the authority of cooperation.
That was the real bug.
Not in Qwen.
In my mental model.
I did not know about the Chinese Room thought experiment before this happened.
I found it after the language switch, after the googling, after realizing this was not only a local-model oddity.
And the timing mattered.
Because if I had read about the Chinese Room in a philosophy article first, I might have treated it like an abstract debate. Interesting, but far away. One of those old arguments people bring up when they want to sound serious about AI.
But I found it after watching a model produce the wrong language with complete confidence.
So the thought experiment did not feel abstract.
It felt rude.
The Chinese Room asks us to imagine someone inside a room who does not understand Chinese. They receive Chinese symbols, follow a rulebook, manipulate those symbols, and send back Chinese responses. From outside the room, the answers look correct. The room appears to understand Chinese.
But inside, there is no understanding.
Only symbol manipulation.
That hit differently after the Ollama incident.
Because the model had produced Chinese.
It had arranged the symbols.
It had continued the conversation.
But did that mean it understood Chinese?
Or did it mean the room had become very good at passing symbols through itself?
Searle’s argument is usually compressed into one brutal idea:
Syntax is not semantics.
Moving symbols correctly is not the same as meaning them.
That idea suddenly had teeth.
The model could generate the language. It could produce the appearance of meaning. It could make the output look intentional from the outside.
But the failure showed me how thin that appearance can be.
A human switching languages in the middle of a conversation usually has a reason. They may be joking, translating, excluding someone, responding to another speaker, or making a social move.
The model did not need a reason like that.
It only needed continuation.
That does not make the model useless.
It makes it strange.
It can produce language without sharing the human responsibilities that usually come with language.
That is why the Chinese Room felt relevant.
It was not saying, “The output is bad.”
It was saying, “Be careful what you infer from good output.”
That is the part I had missed.
After that, I understood why The Origami Software Engineer described LLMs as parrots.
When a model suddenly switches languages, invents confidence, or follows the wrong contextual slope, it feels hollow.
Not stupid.
Hollow.
It can produce the sound of understanding without carrying the same kind of understanding I instinctively project onto it.
That projection is powerful. Developers do it all the time.
If the model *explains* a concept well, we say it ***understands***.
If it *writes* good code, we say it knows ***programming***.
If it *answers* politely, we say it gets the ***context***.
Then one day it switches to Chinese in an English conversation and the mask slips just enough to make us uncomfortable.
The parrot accusation is tempting because it protects us from being fooled.
It says: do not worship the output. Do not mistake fluent text for a mind. Do not confuse prediction with meaning.
Honestly, that warning is useful.
I needed it.
Still, the “meaningless parrot” label also felt incomplete.
A parrot repeats.
LLMs do something more complex than repetition. They summarize, translate, compare, refactor, explain, and connect ideas across context. They often produce useful structure from messy input.
That does not prove consciousness.
It does not prove human-like understanding.
But it also does not feel like nothing.
The model is not a person, but it is not a tape recorder either. It has learned patterns deep enough to act competent across many situations, even if that competence can drift, panic, hallucinate, or confidently violate the vibe contract.
That is where the debate becomes more interesting than “AI understands” versus “** AI is fake**”.
Maybe the right question is not:
Does this model understand like a human?
Maybe the better question is:
What kind of understanding can exist inside a complex system that does not understand the way we do?
That question is harder to answer, but it feels more honest.
The strongest response to Searle, for me, was the systems view.
Maybe the person inside the Chinese Room does not understand Chinese. But maybe the whole room does: the rulebook, the memory, the process, the input, the output, the complete system.
That reply does not magically solve everything.
But it does make the debate less cartoonish.
With LLMs, there is no little person reading a literal rulebook. There is a trained system shaped by data, architecture, context, instructions, post-training, and guardrails. No single token understands. No single layer feels wise. No single probability distribution looks like a mind.
But the full system can still produce useful behavior.
That is the uncomfortable middle.
It is not human understanding.
It is not meaningless noise.
It is a kind of functional competence that can look like understanding from the outside and still fail in ways that reveal it does not share our assumptions.
That is exactly what the Chinese switch felt like.
Useful.
Fluent.
Wrong in a way that exposed misplaced trust.
The biggest change was not that I became anti-LLM.
I still use them.
The change was that I stopped treating fluency as authority.
That one shift matters.
If I want English, I say English.
If I want consistency, I define consistency.
If I want truth, I verify truth.
If the model drifts, I do not treat it like betrayal from a coworker. I treat it like a system revealing where I gave it too much authority.
The model can assist.
It can accelerate.
It can explain.
It can even surprise me.
But it does not automatically inherit the social contract I have in my head. That contract has to be written down.
Sometimes literally.
The Chinese Room did not teach me that AI is useless.
The Qwen incident did not teach me that multilingual models are broken.
The “stochastic parrot” critique did not make me dismiss LLMs as empty.
The lesson was more practical and more uncomfortable:
That is the version of the debate I now care about most. Not whether AI is secretly human, and not whether it is merely fake, but where authority should live when a system becomes this good at sounding like it understands.
For me, the answer is simple enough to keep taped above the keyboard: Do not confuse fluent continuation with shared intent.
That is the leak in the abstraction.
That is the lesson hiding inside the language switch.
That’s not failure.
That’s evolution.
Thanks for being here. It genuinely helps more than you know!