Kamil
66 posts

Kamil
@madebykamil
Founder @SmophyAI - building an AI workspace that turns multiple models into real work, not just chat.

Grok 4.5 was the cheapest near-frontier model from a major US lab. It held that for one day. Meta took it. SWE-Bench Pro, score per dollar of output tokens: Muse Spark 1.1 - 14.47 Grok 4.5 - 10.78 Opus 4.8 - 2.77 GPT-5.5 - 1.95 xAI and Meta published these tables separately. Both put Opus 4.8 at 69.2 and GPT-5.5 at 58.6, which is what makes the comparison worth running. Meta leads tool use. Opus leads SWE-Bench Pro. GPT-5.5 leads the other coding benchmarks. Nobody wins everything.



(1) Today we're releasing Muse Spark 1.1 -- a strong agentic and coding model at a very low price. It's available through our new Meta Model API and in Meta AI.
















