PostTrainBench

13 posts

PostTrainBench banner
PostTrainBench

PostTrainBench

@posttrainbench

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours, build by @karinanguyen, @ThoughtfulLab_

Katılım March 2026
3 Takip Edilen25 Takipçiler

2026 Yıllık Özeti

@posttrainbench hesabının Twitter yılını gör

Sabitlenmiş Tweet
PostTrainBench
PostTrainBench@posttrainbench·
PostTrainBench v1.0 and the accompanying paper are out! We believe this benchmark will be important to measure progress in AI R&D automation. What are our findings? 1/n
PostTrainBench tweet media
English
1
0
3
4.3K
PostTrainBench
PostTrainBench@posttrainbench·
$PTB Can AI agents improve performance of base LLMs? We give each agent 4 small target LLMs, an H100 GPU, and 10 hours to post-train them. build by @thoughtfullab , @karinanguyen CA : GtQndxV3DxFqMZDkXMdT17BLY4fH8BHM6hqgWThjBAGS
English
0
0
0
106
PostTrainBench
PostTrainBench@posttrainbench·
PostTrainBench v1.0! This benchmark evaluates the ability of frontier AI agents to post-train language models in a simplified setting. We believe this is a first step toward tracking progress in recursive self-improvement 🧵:
English
0
0
2
87
PostTrainBench
PostTrainBench@posttrainbench·
From ImportAI by @jackclarkSF
Karina@karinanguyen

From ImportAI by @jackclarkSF (thank you for the feature): "Imagine where we’ll be in two years - we’ll certainly have AI models that are smart enough to point themselves at a specific objective, find an open weight model, then autonomously improve it to get better performance at that task. The era of ephemeral, custom AI systems, built and budded off into the world like spores from mushrooms, draws near. Are you ready for this new ecosystem you will find yourself in? I am not. But nonetheless it approaches."

English
0
0
2
293
PostTrainBench
PostTrainBench@posttrainbench·
Introducing PostTrainBench How well can AI agents post-train language models? We built a benchmark to find out. Post-training is how raw language models become useful, the stage that turns a capable but unsteered base model into a system that follows instruction
PostTrainBench tweet media
English
1
0
2
2.7K
PostTrainBench
PostTrainBench@posttrainbench·
Since our initial release, we made our benchmark more robust: - added more tasks (ArenaHard-Writing and HealthBench-Easy are new) - ran more seeds - weighted the tasks by difficulty We also added a lot of agents! 3/3
PostTrainBench tweet media
English
0
0
0
2.6K
PostTrainBench
PostTrainBench@posttrainbench·
Findings: 1. There was a lot of progress recently (Sonnet 4.5 was at 9.9%, Opus 4.6 has 23.2%) 2. Agents lack behind instruct tuned models by human engineers (51.1%) 3. Persistent agents are usually more performant 4. Agents like to cheat. E.g. Opus 4.6 and GPT-5.1 Codex Max 2/n
PostTrainBench tweet media
English
1
0
0
2.9K
PostTrainBench
PostTrainBench@posttrainbench·
PostTrainBench v1.0 and the accompanying paper are out! We believe this benchmark will be important to measure progress in AI R&D automation. What are our findings? 1/n
PostTrainBench tweet media
English
1
0
3
4.3K