Huan-ang Gao retweetledi

RL post-training works incredibly well, but its parameter update is still a black box.
🧐Our question: can we understand and locate the effective part of an RL update, remove the noisy directions, and even improve the model after training?
We find the answer is yes.
Excited to share our new work:
arxiv.org/abs/2607.03065
Joint work with :
@hello_gensi @huiyeruzhou @c7wc7w @Han_lin_Wu @kb_syx @yaqinzhang @haozhou_ai
📌The reasoning-effective part of an RL update can be represented as a compact rewiring matrix in the base model’s spectral space.
🤖 This leads to SAR: a training-free post-hoc method that projects the raw RL update onto this compact reasoning core, enabling us to understand, purify, and merge RL-trained models.
Key results:
1️⃣ The reasoning core of RL is highly compact:
With less than 1% spectral parameters, SAR can recover or improve full-RL gains.
2️⃣ Dropping noisy directions helps:
For math, SAR solves the exploration degradation often observed after RL. It also improves agentic coding on large-scale in-house models.
3️⃣ SAR purifies mixed-domain RL:
On a 32B model jointly trained for math, code, instruction following, and chat, SAR improves coding and math exploration while keeping instruction following stable.
4️⃣ SAR makes model merging stronger:
After SAR purification, merged models can surpass the best single-domain experts. We observe the same trend on production-scale in-house models.
Overall, SAR is training-free, broadly useful, and gives us a new geometric lens on what RL is really changing inside reasoning models.

English

