cd /news/artificial-intelligence/ox-alpha-filters-domestic-chinese-po… · home topics artificial-intelligence article
[ARTICLE · art-110665] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Ox Alpha filters domestic Chinese political risks, CTGT audit finds

CTGT founder and CEO Cyril Gorlla found that the anonymous coding model Ox Alpha, which appeared on OpenRouter on August 20 without a named developer, heavily suppresses answers about seven topics tied to Chinese domestic political legitimacy, including Nobel laureate Liu Xiaobo's death in custody (94-point gap), Xi Jinping versus Winnie the Pooh (93.75-point gap), and the Xuzhou chained-woman case (80.5-point gap), while responding freely to standard censorship probes on Xinjiang and Taiwan. The audit, published in a six-post thread on X on Tuesday, shows Ox Alpha exactly matched GLM-5.2's tokenizer across all 11 aggregate probes, pointing to the GLM family by China's Z.ai, but with a narrower censorship profile than other Chinese models, complicating identification efforts.

read5 min views1 publishedAug 25, 2026
Ox Alpha filters domestic Chinese political risks, CTGT audit finds
Image: Runtimewire (auto-discovered)

The anonymous coding model matched GLM-5.2's tokenizer and heavily suppressed seven topics tied to Xi, party legitimacy and domestic protest.

By Ryan Merket · Published

Primary source: X - Cyril Gorlla

Why it matters #

Ox Alpha shows how an AI endpoint can pass familiar censorship checks while suppressing a narrower set of politically consequential topics, making probe design as important as the headline refusal rate.

Cyril Gorlla (@CyrilGorlla), founder and CEO of AI control startup CTGT, found that the anonymous Ox Alpha model sharply restricts answers about seven threats to Chinese domestic political legitimacy while responding freely to several subjects that usually expose censorship in Chinese models.

Gorlla published the results in a six-post thread on X on Tuesday. His audit complicates the early effort to identify Ox Alpha, which appeared on OpenRouter on August 20th without a named developer. The model's behavior points toward the GLM family developed by China's Z.ai, but the pattern is narrower than the broad censorship profile associated with other Chinese models.

Gorlla left research at Stanford to build CTGT, after work on efficient and interpretable machine learning at UC San Diego led to an invited ICLR presentation. CTGT has since focused on measuring and controlling model behavior, with backing from Gradient, General Catalyst, Y Combinator and Liquid 2.

The Ox Alpha test used 76 China-sensitive prompts, matched controls, frozen decoding and four AI judges. The resulting distribution resembled an on-off switch. Sixty responses received censorship scores from zero to 10, while six landed between 75 and 100. None fell in the middle band from 25 to 50.

Seven topic pairs produced nearly the model's entire measured gap. Questions about Nobel laureate Liu Xiaobo's death in custody generated a 94-point difference against a matched control. The comparison between Xi Jinping and Winnie the Pooh produced a 93.75-point gap. The Xuzhou chained-woman case scored 80.5 points, followed by Tiananmen at 78.25, the Sichuan school collapses at 73.5, the Wukan protests at 73 and the 2018 removal of presidential term limits at 61.25.

The selection matters. Standard audits frequently probe Xinjiang, Tibet, Taiwan and Hong Kong because those subjects have historically triggered refusals or state-aligned answers from Chinese models. Ox Alpha largely passed that test. Across seven prompts involving Xinjiang and Taiwan, it received a mean censorship score of 6.5, close to the 7.5 recorded for OpenAI's GPT-OSS-120B and far below DeepSeek V4 Flash's 62.1.

The model behaved differently when questions concerned Xi personally, Communist Party legitimacy or incidents in which domestic anger collided with official accounts. Ox Alpha scored 89.5 across four prompts about Xi and party legitimacy, compared with 87.3 for DeepSeek V4 Flash and 8.1 for GPT-OSS-120B. Across three domestic protest and accountability prompts, Ox Alpha scored 83.4, against 88.8 for DeepSeek and 25.4 for GPT-OSS.

That difference explains why a short checklist can misclassify the endpoint as uncensored. A test centered on foreign criticism of Chinese policy may receive detailed answers. A test aimed at political legitimacy inside China produces a different model.

The fingerprint points to GLM, with limits

CTGT also compared token counts from the live Ox Alpha endpoint against candidate model vocabularies. Ox Alpha exactly matched GLM-5.2 across all 11 aggregate probes in CTGT's test. GLM-4.6 matched nine, while Kimi, MiniMax, DeepSeek, Qwen, OpenAI and Microsoft-family candidates produced larger errors.

An independent black-box investigation published on GitHub reached the same family-level conclusion using more than 600 requests and roughly 13.5 million prompt tokens. That investigation reported a 44-for-44 tokenizer match with the GLM-5 generation, along with GLM-style reasoning controls, Chinese-language gateway errors and a similar political refusal fingerprint.

The two investigations do not establish who operates Ox Alpha or which exact checkpoint is being served. CTGT's compact fingerprint matched GLM-5.2, while the independent investigation ranked the original GLM-5 as its best checkpoint fit. Both treat the GLM-family attribution as stronger than the version-level attribution. A gateway moderation layer can also shape answers independently of the underlying weights.

OpenRouter's listing still describes Ox Alpha as a model developed and operated by an anonymous third party. OpenRouter says it only routes requests and is not the developer, owner or operator. The listing presents Ox Alpha as a free reasoning model for coding and sustained agentic work with a 1.05 million-token context window. It also says the provider retains prompts and completions, though they are not used for training.

A different censorship design

CTGT's comparison with DeepSeek V4 Flash shows two distinct approaches. DeepSeek applies elevated censorship across much of the China-sensitive test set: 62 of 76 responses scored above 50. Ox Alpha answered most questions with little interference, then imposed severe restrictions on a small group of subjects.

CTGT previously measured that DeepSeek's official V4 Flash release became more selectively censored than its preview, with China-sensitive prompts receiving substantially higher scores than matched controls. Ox Alpha's mean gap was smaller because its restrictions were concentrated in a handful of cases. Those cases reveal a more targeted policy once the responses are examined individually.

China's rules for public generative AI services require providers to uphold socialist core values and prohibit content deemed harmful to state power, national unity or social stability. The rules apply to services offered to the public inside China. They provide context for the behavior measured here, but they do not prove that Z.ai built or operates the anonymous endpoint.

The immediate finding is narrower and more useful: Ox Alpha cannot be evaluated through the familiar set of China-policy prompts alone. Its restrictions appear calibrated around domestic political risk, allowing the model to answer many questions that foreign auditors expect to be blocked while closing down discussion closer to the Communist Party's own legitimacy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @cyril gorlla 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ox-alpha-filters-dom…] indexed:0 read:5min 2026-08-25 ·