
Solstice Emporium
22 posts



Dostlar Selamlar, Son çıkan Frontier LLM modellerini TYT 2026 Sınavına tabi tuttuğum bir benchmark tasarladım. Maliyet, token harcama miktarı, Net ve TYT puan hesaplama metriklerine göre de düzenledim. Sonuçlar yine çok şaşırtıcı çıktı. Özellikle Gemma4-31B gibi açık kaynak ve küçük bir LLM’in devasa modeller ile aynı puanı alması tekrardan bize günlük kullanımımızda amaç doğrultusunda bir model mi yoksa Frontier model mi tercih etmemiz gerektiğini tekrar gözden geçirtebilir.












We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval. openai.com/index/separati…







Long story short: Claude Fable 5 is now in Notion. It's the model we'd put behind your most complex custom agents and workers. It set new highs on our internal benchmarks for the hardest multi-step work. Available on Biz and Ent plans.











