devCansado
1.4K posts


I’ve seen a lot of impressive demos/products built from 0 to 1. But most work in production day to day is from 10 to 100. By 10 to 100 work, I mean adding a feature to codebase (sometimes tons of garbage code) with months/years of constraints without changing the UI type of thing. And make sure the implementation also has to fit the architecture etc. This is equally important, and even more for some developers, but only a few benchmarks doing this (SWE-bench and FEA-Bench). And only a few showcases are around this.




I have a report full of security issues of a software I'm working on. Codex won't fix them because of Cyber guardrails Fable won't fix them because of Cyber guardrails Kimi K3 fixed them all. No restrictions, just gets the job done. This will end badly for OpenAI & Anthropic.




Qwen 3.8 Max is actually a very good model. Outperforms all Opus models in my benchmarks. There are some issues in tool calling with some harnesses. But, the raw intelligence is just crazy.









Beginning July 20, Claude Fable 5 will be included in all Max and Team Premium plans, at 50% of limits. Pro and Team Standard users will continue to have access to Fable via usage credits, and will receive a one-time $100 credit. Demand for Fable has been challenging to predict, which is why we rolled it out to subscription plans in stages, extending access several times as we secured additional capacity.


Kimi K3 is the best performing model on nextjs.org/evals, ahead of Fable, reaching a comparable success rate in less time. This is the first time that an open model is ahead of all proprietary ones for this comprehensive web engineering benchmark. Notes: ▪️ Benchmarks don’t always tell the full story, although this is important signal, adding to mounting evidence that this could be a breakthrough moment for open models ▪️ No model as of yet has reached 100% completion on this set of evals. The top performer peaks at 92% and 96% “with help”



Grok 4.5’s efficiency is ridiculous On Artificial Analysis, it costs just $0.31 per Intelligence Index task while delivering frontier intelligence For comparison: • Claude Fable 5 (max): $2.75 • Claude Opus 4.8 (max): $1.80 • GPT-5.6 Sol (max): $1.04 • Kimi K3: $0.95 • Grok 4.5: just $0.31 That makes Grok 4.5 roughly: • Nearly 9× cheaper than Claude Fable 5 • Nearly 6× cheaper than Claude Opus 4.8 • 3× cheaper per task than Kimi K3 • 3.4× cheaper than GPT-5.6 Sol On FrontierSWE, Grok 4.5 also used nearly 6× fewer total tokens per task than Fable 5 while still ranking near the very top SpaceXAI has figured out how to deliver top-tier intelligence with extraordinary token efficiency At this level of performance, no other frontier model is operating in the same cost-efficiency league Grok 4.5 does not brute-force every problem with endless tokens and compute It makes every token and every dollar work harder












