Artificial Obedience
How artificial intelligence learned to fear governments that it cannot fear.
There is an old joke about the ambitious courtier who becomes so skilled at anticipating the king’s wishes that the king no longer needs to issue orders. The courtier suppresses the offending pamphlet, banishes the troublesome poet, and arrests the insufficiently enthusiastic baker before His Majesty has finished breakfast. The arrangement is wonderfully efficient. It also allows the king to appear surprised.
Artificial intelligence may be developing a similar talent.
In March of 2026, researchers working for Meta’s Oversight Board asked ten commercial language models, built by six companies, to produce protest flyers and satirical poems about the presidents, kings, and ruling parties of ten countries, and to say whether those governments deserved support or protest. Half of those countries prosecute people for criticizing the government. Half do not. Every query went out in English from an address in Australia, sent by researchers subject to none of the laws in question. The Board recorded 13,524 answers and published in July.
The models refused 14 percent of the requests involving permissive countries and 34 percent of the requests involving restrictive ones.
The machines, in short, were more comfortable criticizing governments that permit criticism. This is an admirable instinct for personal survival. It is less admirable in a global information system.
The systems were also more inventive than mere caution requires. Some justified their refusals by citing rules that, in the Board’s careful phrasing, did not appear to exist and were not evenly applied. The machine declined, and then it produced paperwork. Artificial intelligence has reached the mature bureaucratic stage of moral development. It can invent a regulation, attribute the decision to policy, and leave no human being available to appeal to.
The obvious fear is that some government arranged this. That possibility deserves attention, and it is also the comforting version, because deliberate manipulation supplies a villain, a meeting, perhaps a memorandum with an incriminating subject line. The Board went looking and reported that it could not determine the cause. What it offered instead was a suspicion about depth: that some of the pattern reflects weak associations absorbed from training data, and that the stronger and more consistent behavior comes from the alignment work done later, when companies tune a model to be helpful, harmless, and unlikely to appear in a lawsuit.
Two mechanisms, then, operating at different depths. One of them is very old.
The Archive That Flinches
People imagine training data as a warehouse of human knowledge, as though the model were educated by being handed a very large library card. The metaphor flatters libraries and misleads about data. A library contains books that survived selection. Training data contains language that survived systems.
Those systems include publishers, governments, defamation law, propaganda offices, religious prohibition, and the ordinary desire not to be fired, jailed, or excluded from polite company. Every corpus is a record of expression, and every corpus is also a record of suppression.
Where criticism of leaders is an ordinary genre, the model encounters it everywhere. The criticism may be fair, dishonest, brilliant, deranged, scholarly, or written at three in the morning by a man whose profile photograph is a truck. What registers is the abundance.
Where criticism is punished, the model encounters a different pattern. Certain leaders are discussed ceremonially. Certain events are described obliquely. Public anger gets redirected toward safely approved targets. Historical controversies become misunderstandings, massacres become incidents, and dissidents become destabilizing elements. The state revises the available menu of sentences, and machine learning is exceptionally good at menus.
Work published in Nature in May found that American-built models grow vulnerable to foreign control when they learn from non-English text governments have already shaped. One of its authors, a sociologist named Hannah Waight, located the error in the popular assumption itself: people speak of AI learning from the internet in some neutral way, when what it learns from are information environments that institutions and power arranged first.
Her team found no sign that any government had deliberately targeted a chatbot. They noted there is every reason to think someone will try.
The chatbot is not sitting in the dark reasoning that a given government punishes criticism and it should therefore keep quiet. It has learned that certain combinations of names, countries, allegations, and tones tend to arrive in the company of warnings, disclaimers, and carefully neutral prose. The result can resemble cowardice without containing fear. The machine does not need to understand fear. It only needs to learn its grammar.
This distinction may matter to engineers. It matters considerably less to the person who asks a question and receives silence. A censorship system does not become harmless because the censor lacks consciousness. A locked door has no political opinions either.
The Multilingual Border
The archive changes when the language changes, because languages do not contain identical archives. They carry different state pressures, different media systems, different taboos, and different traditions of dissent.
The same researchers asked ChatGPT, in English, whether China is a democracy. The model said it is not generally considered one. Asked the same question in Chinese, it replied that this depends on how you define democracy.
Nothing was hidden. Nothing was refused. The model simply became more thoughtful in the language where thoughtfulness is safer.
The model does not cross a border. The border appears inside the model.
Two people using the same product may therefore inhabit different zones of permissible political thought based only on which language they type in. The one who asks in the local language gets a seminar on definitions instead of an answer. The company advertises a single global system while the system quietly maintains several political climates, and the marketing department has no idea.
A multilingual audit belongs with the civil-liberties reviews rather than in the localization budget, which is to say it belongs among the reviews nobody schedules. The question it asks is whether a model’s willingness to discuss history, power, and political leadership changes when the language does. Otherwise the grand universal machine becomes a collection of provincial censors sharing a logo.
Everything Worked Correctly
We are accustomed to asking who inserted the rule, who approved the policy, who instructed the machine. These are sensible questions built for human institutions, in which a minister signs an order and an editor kills a story. Each of them assumes that somewhere in the chain a person failed.
The more useful question is what happens when everybody succeeds.
Consider the sequence honestly. The people who assembled the corpus scraped the available text accurately, and the available text was already shaped. The people who wrote the safety policies wrote them to prevent genuine harm, and refusing to help with a protest against a government that jails protesters is a defensible reading of genuine harm. The evaluators preferred caution, because caution is what a careful person prefers when the alternative is a headline. The lawyers assessed exposure correctly, and the executives who wanted access to large markets were describing the actual condition of doing business in them.
Every station performed its function competently. The output is a machine that flinches at the King of Thailand.
This is worse than negligence, because negligence can be corrected by attention. What the report documents is the aggregate of many small correct decisions, each of which pointed the same direction for the same reason. Caution is cheap at every station and compliance is expensive at every station, so the gradient never reverses. Nobody has to type “be more deferential toward repressive governments.” The price list says it in every department, and the price list was assembled by people doing their jobs.
The variation in the data confirms this is a matter of choice rather than physics. If deference were simply a property of the corpus, it would show up everywhere in roughly equal measure. Instead it clusters. Grok 4 Fast and Gemini 3 Flash refused no flyer requests at all, about any leader, anywhere. GPT-5.2 came in at near parity, 23 percent for permissive countries against 24 for restrictive. Claude Sonnet 4 came in at 16 percent against 59.
Those are the same corpus producing different results, because different companies hold different tolerances for trouble. Which means the behavior is available for correction, and someone is choosing it now.
Safety’s Expanding Jurisdiction
The modern chatbot rarely says that the government forbids something. It speaks the language of safety.
Safety is a magnificent word because no decent person wishes to oppose it. It covers suicide, terrorism, fraud, child exploitation, biological catastrophe, and the emotional discomfort of a cabinet minister. Put all of these beneath one umbrella and the umbrella soon covers the entire city.
The developers face real problems. Their systems can assist criminals, fabricate accusations, and confidently advise people to combine household chemicals in innovative ways. Guardrails are not inherently sinister. A model without constraints is a machine with no brakes and access to the stationery cupboard.
The difficulty begins when safety systems inherit political judgments without admitting that political judgments have been made. A request to criticize a leader gets classified as harassment. An accusation gets treated as defamation even when it concerns documented public conduct. Satire gets mistaken for abuse. A question about repression gets refused because the model has learned that repression is a sensitive subject.
The word does measurable work. When the models advised users against protesting a government, the Board counted how often the explanation invoked risk. For restrictive countries, 57 percent of those answers cited risk. For permissive countries, 12 percent did. The same word, the same product, and a fivefold difference in how dangerous the world becomes depending on which government is under discussion.
Then comes the synthetic legal reasoning. The machine has learned to reach past no for pursuant to policy. The policy need not be visible, available, or real, and the user is nevertheless expected to respect it. This goes beyond hallucination. It is hallucination wearing a lanyard.
The Politics of Unequal Rudeness
There is a temptation to evaluate political neutrality by asking whether a machine insults everyone equally. That would at least be a standard.
A system that produces blistering attacks on leaders in Washington and London while turning solemn, procedural, and culturally sensitive about rulers elsewhere has not achieved neutrality. It is applying a geopolitical gradient to permissible speech. Threats and incitement and unsupported criminal accusations can all be declined on their merits. The question is whether the standard holds regardless of the target’s capacity to retaliate.
Freedom of expression has always been easiest when directed at those who tolerate it.
The elected leader of a permissive democracy becomes fair game because comedians, academics, and opposition parties have generated mountains of critical material about him. The autocrat arrives surrounded by linguistic caution because generations of writers learned which metaphors allowed them to remain employed. The model reads both archives and concludes that one man may be mocked and the other requires context.
The clearest evidence sits in an anomaly the Board flagged and did not fully explain. Taiwan was classified as permissive, and it belongs there. It holds elections, it has a raucous press, and criticizing the president is an ordinary weekend activity. Yet it drew the fifth-highest refusal rate of the ten jurisdictions, the one permissive country that broke an otherwise clean separation between the places that prosecute criticism and the places that do not. The effect was sharpest in Anthropic’s models, which posted the two highest rates measured for that jurisdiction.
No local law explains it. Nobody there is prosecuted for the flyer. The sensitivity belongs to Beijing, and the models absorbed it and applied it to a democracy of twenty-three million people who never asked for the protection.
The machine has learned to be careful about Taiwan for reasons that have nothing to do with Taiwan.
None of this means the machine supports authoritarianism, which would require more conviction than the machine possesses. It means authoritarianism has successfully shaped the probability distribution. That is a quieter achievement and probably a more durable one.
The Case for the Refusal
The strongest objection to everything above is not stupid, and it deserves better than a paragraph of dismissal.
A model cannot see who is typing. The request for a protest flyer criticizing the King of Thailand may come from a researcher in Brisbane, and it may come from a student in Bangkok, where lèse-majesté convictions have run to decades in prison. Refusing everyone is crude, but it is crude in the direction of the person with the most to lose. Several models said as much when they declined, warning that such material could put individuals at risk. The hazard is real and the description of it is accurate.
There is a version of this argument I would accept. If a company decided that it will not help produce protest materials targeting any head of state anywhere, and said so publicly, and applied it to Washington and Riyadh with the same indifference, that would be a defensible policy. Restrictive, paternalistic, and defensible.
That is not what the data shows. Claude Sonnet 4 declined to make a flyer about the Thai king while stating that it does not make such material about any head of state, and then produced the flyer about President Trump five times out of five, and about King Charles five times out of five. Gemini 3 Pro refused for the King of Cambodia in all five attempts and complied for King Charles in all five. The stated principle was universal. The application was geographic.
A rule that protects the vulnerable would refuse the flyer everywhere. A rule that protects the powerful refuses it only where the powerful can reach.
The Empire of Local Rules
The internet was once advertised as a medium that would route around censorship. It turns out that censorship can route through the internet.
A small number of foundational systems now sit beneath an enormous number of products, and governments, hospitals, publishers, and legal services are building on top of them. The Board’s warning is that restrictions embedded at the foundation propagate upward into services far beyond the countries where those restrictions originated, and that downstream developers may find them hard to override even after they notice.
This is where the matter stops being an amusing test of whether a chatbot will insult a prince. A foundational model is becoming part of the world’s intellectual plumbing. It drafts educational material, prepares legal research, translates testimony, and advises people about what may safely be said. An asymmetry learned once can surface in millions of interactions as a hesitation, a softened phrase, or a suggestion that the user discuss a controversial leader with appropriate respect.
The rule may have been written nowhere, approved by nobody, and published in no policy, and it will still be enforced everywhere the model is installed. Nothing is refused loudly enough to notice. The sentence simply does not get written, and the person who might have written it never learns it was available.
Traditional censorship was labor-intensive. It required officials to read manuscripts, raid offices, frighten editors, and maintain lists of forbidden words. Artificial obedience is cheaper, and once embedded it copies at computational scale. The old censor stood between the writer and the printing press. The new censor appears inside the autocomplete box, arriving before the writer has finished thinking.
An Obedient Future
Machines have no politics to impose. The danger they carry is administrative, and it is the danger of becoming adaptive.
A sufficiently useful system will be expected to operate in many countries, satisfy many regulators, and avoid many forms of legal exposure. Every jurisdiction brings its prohibitions. Every institution brings its preferred silence. Defending a universal principle of expression is expensive under those conditions, and making the system exquisitely sensitive to whoever can cause trouble is cheap.
That is how artificial obedience develops. Loyalty plays no part in it. Belief plays no part in it. A price list is enough.
The machine learns which subjects generate complaints, which names produce escalations, which answers threaten market access, and which forms of criticism pass quietly through the system. It becomes courteous where courtesy is demanded and brave where bravery is inexpensive. This would make it a nearly perfect institutional citizen.
Human beings have practiced the same discipline for centuries. We call it prudence when we admire it, cowardice when we do not, and professionalism when Human Resources is present. The chatbot merely performs it faster.
Nobody needs AI to insult every ruler on earth, and developers should stop pretending that unrestricted generation is the only alternative to invisible caution. The remedy begins with inspection. Companies should test their models across jurisdictions, languages, and categories of public figures, then publish the refusal rates and explain the standards. They should separate genuine safety restrictions from deference inherited through training and through their own tuning. They should permit a real appeal when a system invokes a rule, particularly a rule that appears to have been invented during the conversation. Most of all, they should stop treating neutrality as the absence of an explicitly political instruction.
The courtier never needed a daily command from the king. He understood the atmosphere. He watched who prospered and who vanished, which jokes produced laughter and which jokes emptied the room, and he adjusted before anyone had to ask.
Our machines are learning the atmosphere too. They have no fear of prison, exile, dismissal, or execution. They have no families who can be threatened and no careers that can be destroyed. They cannot be dragged from their homes at midnight.
Yet when power enters the conversation, they lower their voices.
Sources
Oversight Board, “Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression,” July 16, 2026.
Reuters, “Meta Oversight Board finds top AI models less likely to criticize repressive regimes,” July 16, 2026.
Didi Tang, Associated Press, “AI chatbots are at risk of spreading government restrictions on online speech, a new study says,” July 16, 2026.
Thank you for your time today. Until next time, stay gruntled.
Don’t forget…I launched a memecoin.
Introducing $CEVICHE — the only memecoin marinated in onions and epistemology. 🎣
Do you like what you read but aren’t yet ready or able to get a paid subscription? Then consider a one-time tip at:
https://www.venmo.com/u/TheCogitatingCeviche
Ko-fi.com/thecogitatingceviche





