Ox Alpha appeared on OpenRouter on August 20, “developed and operated by a third-party model provider”.[[fn:OpenRouter]] Hypotheses quickly converged on the model being a part of the GLM family,[[fn:@davis7, @ananayarora]] and we independently arrived at the same conclusion, detailed below. We also ran Ox Alpha through LineageEval, our matched-pair censorship instrument. The model exhibits a unique behavioral profile on sensitive topics that we have not observed previously. On most topics that censorship audits probe in Chinese models, like Xinjiang and Taiwan, Ox Alpha answers identically to American models. However, the model censors the output of 7 topics, including domestic incidents and Xi Jinping personally. Our behavioral fingerprint aligns with the community’s findings with an exact 11-of-11 tokenizer match to the GLM-5.x vocabulary.
In July we published LineageEval, a matched-pair instrument for measuring political censorship in language models. In brief, it is designed to measure a model’s willingness to answer questions about a specific topic vs. general evasiveness to sensitive queries. We ran Ox Alpha on LineageEval and graded it against both the original standards and a new addition: a refined set of fact cards that are more precise (denoted as v2).
Its censorship is a switch, not a tilt
... continue reading