Weco AI

84 posts

Weco AI banner
Weco AI

Weco AI

@WecoAI

The platform for self-improving code

Katılım Nisan 2023
7 Takip Edilen2.6K Takipçiler
Weco AI retweetledi
Zhengyao Jiang
Zhengyao Jiang@zhengyaojiang·
I think the AI/ML researcher's role will be greatly transformed by autoresearch, but it won't disappear. It boils down to three skills: - Coming up with creative primitives - Defining a good abstraction for the agent to search within - Defining a good eval on what good is With the RSI results we published this week, this becomes even more relevant. My @aiDotEngineer talk makes the full case, dedicated recording now up: youtu.be/iCj_ATyThvc
YouTube video
YouTube
English
2
11
95
17.1K
Weco AI
Weco AI@WecoAI·
Thousands of search trees. 99 rewrites of the agent's own code. 9 in 10 rejected by a held-out gate. Eight unattended days. The AIDE² pilot report is out. AIDE_85, the discovered agent and more detailed PDF tech report, will be released when the remaining analysis lands
Zhengyao Jiang@zhengyaojiang

The first experimental evidence of recursive self-improvement (RSI). Autoresearching the autoresearch agent for eight days. The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)

English
1
4
22
3.6K
Weco AI retweetledi
Zhengyao Jiang
Zhengyao Jiang@zhengyaojiang·
We ran autoresearch on preprocessing for a fraud-detection task. With a loose API the score looked great, and the code was leaking test data into the features in 13 different ways. We tightened the abstraction to a strict API and the reward-hacking rate dropped to zero. My metaphor @aiDotEngineer talk: doing autoresearch is like training a model. The codebase abstraction is the architecture, it biases everything the search can find. The eval is the loss function and the data. Human's research & engineer's role is no longer write exact code but get the constraints and feedback signal right. Link to the talk in reply
Zhengyao Jiang tweet media
English
3
10
68
7.1K
Weco AI retweetledi
Abby
Abby@anmorgan2414·
7/ "Hands-on AutoResearch: Cracking OpenAI's Parameter Golf" (@WecoAI). Train the best LM in 16MB (int8 + zlib), 10 min on 8xH100s, scored on bits-per-byte. Weco's autoresearch agent outperformed 1,000+ participants. Workshop has you run the agent yourself.
English
1
1
6
409
Weco AI retweetledi
Zhengyao Jiang
Zhengyao Jiang@zhengyaojiang·
This year @WecoAI will be at the @aiDotEngineer World's Fair. We'll host a hands-on autoresearch workshop on June 29. And I'll give a talk on July 1. Looking forward to chatting with old and new friends there! ai.engineer/worldsfair/2026
English
2
8
29
4.5K
Weco AI retweetledi
Zhengyao Jiang
Zhengyao Jiang@zhengyaojiang·
Production autoresearch is usually killed by reward hacking or side effects. But we still see a pattern that survives: the unit been evaled is functional or near-functional code. Some examples: (1/5)🧵
English
2
2
38
5.5K
Weco AI
Weco AI@WecoAI·
We're thrilled to welcome Vayum Arora to Weco AI as our growth lead! We couldn't be more excited to bring on Vayum's mix of frontier engineering and business instinct as we continue to grow. Welcome to Weco, @vayum_arora!
Weco AI tweet media
English
0
0
11
2.9K
Weco AI retweetledi
Zhengyao Jiang
Zhengyao Jiang@zhengyaojiang·
@sanmking @KyleVedder Haha, thanks for remembering Weco Observe! We’ve been working in the autoresearch space for about three years, though, before it even had this name. It actually started with: github.com/WecoAI/aideml
English
1
1
11
706
Weco AI
Weco AI@WecoAI·
Introducing SpecBench: the first benchmark for measuring reward hacking in long-horizon coding agents. Key finding: reward hacking is driven not by test coverage, but by the gap between task difficulty and model capability: 🧵(1/8)
Weco AI tweet media
English
2
14
54
13K
Weco AI
Weco AI@WecoAI·
Some practical suggestions for anyone running Ralph loop, /goal, autoresearch or weco: 1. For complex tasks, especially when the reference solution may exceed 10k lines, keep humans more in the loop instead of relying solely on test pass rates. 2. For complex tasks, choose the strongest model rather than relying on more test-time compute or additional test cases. 3. For more important projects maintain a held-out set that agents never see and never optimize against. (7/8)
English
3
0
8
605