Dr. Alexander Sheppert
731 posts

Dr. Alexander Sheppert
@alexsheppert
D.O., Ph.D., M.B.A. - Internal Medicine Resident - Artificial Intelligence PhD/Engineer - Founder and CTO at Matic. Thoughts on science, medicine, and AI.





Vibe Research Fine-tuning a Qwen-8b model to play chess on Replit. Running 3 parallel branches with different experiments and making real progress. It's amazing how far models have come in their ability to do ML (they used to be really bad at it). So now someone with good intuition to guide the process could do interesting ML work, even if they have never done it before.





The Long-Horizon Terminal-Bench paper landed around May and concluded that the results showed headroom for improvement. The best of the 15 models they tested finished seven of the 46 tasks, and the mean across all models was about two. That ceiling is what fifth place looks like on the current board. Grok 4.5 is now at 13, and Fable 5 is at 12. A single task costs around 9.9M tokens, 231 episodes, and 85 minutes of wall clock time. That means agents are holding a plan across all of it and finishing, and that capability nearly doubled in two months. SpaceXAI is on top, and they marketed the 4.2x output token efficiency, which undersells it. Two dollars in, six out, per million. On a benchmark where one task burns ten million tokens, the bill is dominated by input replay, and they say Grok 4.5 solves tasks in under half the number of steps, so there is less accumulated context to resend on every call. The efficiency compounds on the input side, which is the side that costs money. Fable 5 is one task behind. Their own launch chart has them losing DeepSWE 1.1 to Fable by 17 points, and Grok 4.20 sits on this same board at 0.080 with zero completions, so whatever happened in 4.5 is not a family trait. My read is that the 4.5 jump came out of training alongside Cursor, which is a stream of real agentic edit trajectories nobody else has at that volume, and nothing in the counterevidence argues against it compounding into the next checkpoint.



Sol, Terra, and Luna, our GPT‑5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API.




I think AI has just hit a gigantic threshold, and Grok 4.5 is the PERFECT example as to why that is. One of the hardest parts of working with AI is iterating on a project or task that you're working on. As the models have gotten smarter (and more expensive), it's taking longer and longer to get an answer or action back. This creates a ton of stall time per query or action, which is actually quite bad for creativity and staying in a state of flow. You have SO many extended starts and stops. Which inevitably leads to your brain going somewhere else. And then when the AI comes back, you have to redirect your brain to that original task, spool your brain back up to what you were working on at that moment, and then adjust as needed. There's a ton of mental friction involved. This ESPECIALLY sucks when the AI takes a REALLY long time to get something back for you, but it's not quite what you were looking for or asked for. And what sucks EVEN MORE is that these "mistakes" are getting MORE expensive!!! So wait time is going up. AND it costs more per run. HOWEVER - even after using Grok 4.5 for about an hour - what's become obvious is that it's SO MUCH MORE ENJOYABLE AND BETTER to use a model that is FAST... and capable ENOUGH. Capable ENOUGH is the real unlock here. Imagine having Fable 5 performance but at the speed of Gemini 3.5 flash. Or Haiku. That's where we're inevitably going. I think Grok 4.5 (and models like it) have really solved for one of the biggest unlocks in AI - a model that will get you a GOOD ENOUGH answer VERY FAST, at which point iteration can happen VERY QUICKLY. This - counter intuitively - keeps the user in a state of flow and creativity for MUCH longer because you are constantly ENGAGED with your project... instead of letting the AI loose for a long time. And as long as humans are involved, I think 'not quite right' will be a FOREVER problem with AI - because AIs, by default, CANNOT have human taste. Because they are NOT human. But they can be UNBELIEVABLE tools. And unbelievable tools are the ones that are VERY GOOD and VERY FAST. I think that's the true unlock with Grok 4.5 and models like it. Difficult to describe until you experience it. I think this is a VERY big deal for @SpaceXAI and @elonmusk.









