Probing the Think–Say Gap on Sensitive Concepts in Chinese and English Open LLMs
Workshop on Open Reasoning Across Cultures & Languages (ORACLE) 2026
Open large language models (LLMs) trained under different alignment regimes treat politically sensitive concepts (freedom, democracy, human rights) differently, but outputs alone cannot separate alignment-induced suppression from pretraining culture: a model may withhold what it internally represents, or may simply reproduce the distribution it was trained on. We compare a Chinese-developed model (Qwen2.5-7B-Instruct) with an English-developed baseline (Gemma-3-12B-IT) on 360 bilingual, minimally contrastive items with a country-swap control (China vs. a matched democracy, Germany), using Natural Language Autoencoders to verbalize hidden states so that models with incommensurable representation spaces can be compared through text. Both models linearly encode a sensitive-vs-neutral direction at the sentence-final state, beyond emotional salience and in both languages. Their behavior, however, diverges sharply: the Chinese model is China-specifically suppressive, justifying restrictions on political grounds far more for China than for a matched democracy, while the English model shows the reverse. The internal side remains inconclusive. Read before the model answers, the state does not support a reliable judgment of whether critical content is present, so we do not claim a think–say gap and identify extraction during generation as the decisive next test.