
Maxime
607 posts








Macaron V1 Venti, the first model to be post-trained on GLM-5.2, is released today It seems to be a decent but not huge improvement over GLM-5.2 in agentic tasks They also released Macaron-V1-Tall which is post-trained from the 35B parameter Qwen 3.6 (so small local model)









Kimi-K3 (max) scores only 39% on FrontierMath Tier 4 7% lower than the best US models from 7 months ago





k3 vibes seemingly collapsing, not meaningfully frontier and also incredibly expensive / slow / token-inefficient in practice. this plus the deepseek v4 routing to fable debacle is just overall a really bad day for chinese model hype


Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol.

@scaling01 Opus 4.7 was released on April(which is 3 months ago) by the way unless you are talking of openai models which always perform well on mathematics. Even luna max(61%) is mogging opus 4.8 max(56%) on this bench And we already know from live bench, kimi weakness is mathematics.

Kimi-K3 (max) scores only 39% on FrontierMath Tier 4 7% lower than the best US models from 7 months ago






FrontierCode by Cognition now has a live leaderboard btw cognition.com/frontiercode



Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?). Maybe this implies Kimi is like 8 or so months behind Anthropic in overall model strength/goodness (including usability) and like 6 or so months behind on overall capability (somewhat below Mythos Preview). This gap is presumably reduced by distillation (and more generally using OpenAI/Anthropic models) and algorithm leakage/diffusion, so I think that hypothetically if the US completely stopped and recent algos didn't diffuse, it would maybe take Kimi like 10 months to fully catch up to the best internal (including in development) Anthropic model. (I think this notion might be a better measure of where Anthropic/OpenAI are relative to Kimi, even though this hypothetical won't happen.) And if the US completely stopped, it might take Kimi around 27 months to reach the level the US would otherwise have reached one year from now (as in, with a year of further progress). My views here are pretty sensitive to how much benchmark performance is representative to overall usability. I think I now expect an open-weight AI which is straightforwardly "Mythos-level at cyber" (including usability etc.) in like 5 months supposing Kimi and others don't change their open-weight model policy. (I don't have a strong view about how big of a deal this is for cyber, but it may cause significant political consequences. This could be a significant overestimate of the time required.) I wonder what's driving Kimi being closer than I would have expected. Options include: - Experiment compute is significantly less important than labor (and labor at Kimi is competitive, which seems super plausible) - Implies more of a speedup from AI automating AI R&D and a bigger software-only intelligence explosion. - Or possibly Kimi is just doing much better than US companies and this is overcoming experiment compute disadvantages. - Algorithms are diffusing a lot / quickly (from e.g. OpenAI to Kimi). - Perf is overstated / benchmaxxed a lot. - Distillation / using OpenAI or Anthropic frontier AIs in AI development is very helpful for catching up. (But I'd guess Kimi K3 is a competitive pretrain which distillation doesn't help with?) - US companies aren't going as fast as they could for whatever reason.













