
AREX is a new breed of research agent: it doesn’t just search longer, it recursively audits and improves its own work, constraint by constraint.
The secret? An inner loop drafts answers, while an outer loop checks every claim, flags what’s missing, and spins up focused follow-ups. A learned context-update tool keeps the model’s “working memory” lean—compressing sprawling tool histories to just 26k tokens on average (vs. 128k+), all while boosting accuracy by 12 points.
On tough benchmarks like BrowseComp, DeepSearchQA, and Humanity’s Last Exam, AREX-Base (10B active params) outperforms many models 5–10x its size, hitting 82.5 on BrowseComp. Ablations show: +11.8 points from context compression, +10–11 from the outer loop, +8 from key-step focused training.
The result: far more reliable, efficient long-horizon reasoning—without massive compute. If you want a blueprint for practical, trustworthy AI researchers, this is the paper.
Get the full analysis here: yesnoerror.com/abs/2607.21461
// alpha identified
// $YNE
English