Boolean

2 posts

Boolean

Boolean

@boolean_ai

Research lab focused on building evals for frontier models

Katılım Ocak 2026
0 Takip Edilen11 Takipçiler
Boolean retweetledi
Sanyam Satia
Sanyam Satia@sanyamsatia·
FrontierBench is live! Excited to have been a task author and reviewer on this effort, with @boolean_ai as a data partner. The team was rigorous in keeping the quality bar high at scale, across a wide range of domains. Congrats to @ryan_marten and everyone involved on the launch!
Ryan Marten@ryan_marten

We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work. Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort. Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%

English
1
4
21
60K
Boolean retweetledi
Sanyam Satia
Sanyam Satia@sanyamsatia·
We ran a frontend eval from an in-progress internal benchmark. Kimi K3 is not at the same level as current frontier models like GPT 5.6 Sol or Fable 5. It's closest to Opus 4.7 on this eval so it's 3 months behind frontier. An impressive result nonetheless. Evaluating frontier models is going to increasingly require very high taste and in-depth domain expertise.
Arena.ai@arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

English
1
3
11
2.3K