f
38.5K posts

f
@fb8yvu
Open Minded No hurts feelings

Grok 4.5 now also ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol Under the strictest scoring metric - where a task counts only if it is fully solved with a perfect reward and zero errors......Grok 4.5 finished clearly ahead of every other tested model This matters for real-world coding, automation and difficult engineering work because long-horizon terminal tasks require much more than making partial progress The model has to maintain context, recover from mistakes and successfully complete an entire workflow across hundreds of steps Grok 4.5 is showing serious strength on complex agentic tasks over extended periods of time

































