
New research: the strongest open-weight LLMs increasingly come from Chinese labs, and they carry political censorship from training. Ask about Tiananmen, Xinjiang, or Taiwan and the model deflects, hedges, or restates the official line.
Can you remove it? The obvious tool is machine unlearning, the standard way to make a model forget something.
We took Qwen2.5-7B and tried it. No matter how long we trained, it backfired: instead of removing the censorship, the methods made it worse or broke the model.
We explored why the method built to make a model forget is the wrong one here, what actually removes the censorship without breaking the model, and why you can't even tell how censored one of these models is until you fix the way censorship is measured.
Download the full research paper: lab.cloud/research/uncen…
Transformer Lab@transformerlab
English



