Shaun Smith

2.4K posts

Shaun Smith banner
Shaun Smith

Shaun Smith

@evalstate

https://t.co/rA1UoojwhN https://t.co/76p6mDAfej

united kingdom Katılım Temmuz 2024
889 Takip Edilen1.1K Takipçiler
Shaun Smith
Shaun Smith@evalstate·
This. is. insane. $0.04 per trial. Last weekend I benched Grok 4.5 at 80.5% on tb-2.1 for $159 and was impressed. Today gpt-5.6-luna has beaten it: with flex tier pricing the run cost less than $20. Industry economics changed this week.
Shaun Smith tweet media
English
2
3
18
727
Shaun Smith
Shaun Smith@evalstate·
@xeophon No - not fast. But I've not had any weird errors or disconnects [yet].
English
0
0
2
61
Florian Brand
Florian Brand@xeophon·
i love benching the whale
Florian Brand tweet media
English
7
3
84
2.6K
Shaun Smith
Shaun Smith@evalstate·
@andrew_n_carr Agreed. I've completely changed the way I use subagents these last 2 weeks; for conversational stuff big models are a total waste now.
English
0
0
0
100
Andrew Carr 🤸
Andrew Carr 🤸@andrew_n_carr·
My guess is that the models are getting harder to understand because they're being optimized for sub agent spawning and communication. We are not the intended audience
English
9
3
78
4.7K
Shaun Smith
Shaun Smith@evalstate·
Full ATIF trajectories available in harbor, leaderboard submission PR#184. Cost taken from API Key used for the trial. fast-agent cache efficiency was 93.83%, 1.75% of prompt tokens were priced at long context rates.
English
1
0
0
82
Shaun Smith
Shaun Smith@evalstate·
@SarahLacard Yep, my estimate for a TB 2.1 run is about $10. I'm currently figuring out how reproduceable the published score is and doing a couple of reasoning level sweeps.
English
0
0
1
15
Sarah 🇨🇦 🏳️‍⚧️
@evalstate oh i haven't gotten that fancy with the sauce - only using it in opencode, it's a monster, and cheaper than luna from what i can tell
English
1
0
1
25
Shaun Smith
Shaun Smith@evalstate·
omw... this was $20 well spent.
Shaun Smith tweet media
English
1
0
4
217
Shaun Smith
Shaun Smith@evalstate·
@SarahLacard I have, and am on that now😀. Very nice they are delivering it over the Responses API too (but no Web Sockets).
English
1
0
1
16
Vaibhav (VB) Srivastav
I love my team, colleagues and job! It’s been so much fun the last couple days - just letting my ideas run free and building things I care about!! Oh, what a privilege it is that I get to do what I do \o/ Yes, this post is brought to you by two pints and a negroni
English
5
2
156
6.1K
Shaun Smith
Shaun Smith@evalstate·
@Infoxicador Yeah, it's on the official leaderboard but an earlier version of codex. Token drain is real, I think fast-agent is about half the cost of OpenAIs figures in their report.
English
0
0
1
41
Shaun Smith
Shaun Smith@evalstate·
Efficiency matters you say?
Shaun Smith tweet media
English
2
1
10
621
Vaibhav (VB) Srivastav
One of the things I love most about this video is how well it captures the ethos of OpenAI and what drives the people behind it to wake up every day and deliver some of their life’s best work. Jason really put his heart and soul into this one. Kudos, mate!
jason@jxnlco

one of the most beautiful things about OpenAI is that every employee really has a voice. i wanted to capture what it feels like to work here, what our mission means to me, and why you should join us. so i made this video with Codex, shared it with the team, and they felt it was worth producing and sharing with the world. this is our mission. and it’s why i’m here.

English
2
3
101
8.8K
Shaun Smith
Shaun Smith@evalstate·
@liran_tal Yep, but when tokens are that cheap... (aesthetically it's horrible, but I think after using GLM for a bit we should all switch reasoning traces off for good)!
English
1
0
1
25
Liran Tal
Liran Tal@liran_tal·
@evalstate Interesting! Wouldn't heavy token cost directly contribute towards higher spend though? (not across pricing models but across same tier in general) Like for example Opus 4.8 and seems like Sonnet 5 too are incredibly verbose in thinking process
English
1
0
0
90
Liran Tal
Liran Tal@liran_tal·
Wait what 5.6 Luna is scoring higher than Sonnet 5 ???
Liran Tal tweet media
English
3
0
8
1.5K
Shaun Smith retweetledi
merve
merve@mervenoyann·
Thinking Machines released Inkling Small (🦖) + NVFP4 12B active 276B total params, the model performs better than larger Inkling on coding 🤯 > check out our blog covering benchmarks, performance and deployment huggingface.co/blog/thinkingm… > huggingface.co/collections/th…
merve tweet media
English
7
20
158
16.4K
Shaun Smith
Shaun Smith@evalstate·
@un3valuated Yeah, I'll post some terra benchmarks tomorrow. Sonnet is buried, Flash is buried and at $15 output Moonshot doesn't compete here. Grok is the one to watch now. What a time.
English
0
0
1
31
short circuit
short circuit@un3valuated·
@evalstate My half cheap half intelligent model was Gemini 3 flash for last 8 months, now Luna costs in half. Rest in peace bro, you'll be missed.
English
1
0
1
69