Eran Sandler

15.5K posts

Eran Sandler banner
Eran Sandler

Eran Sandler

@erans

Builder, operator and investor. Infra, AI, and product nerd. Trying to make powerful things simple. Opinions are my own. Building https://t.co/b0sgru9dFz

Katılım Nisan 2007
4.2K Takip Edilen973 Takipçiler
Eran Sandler retweetledi
Chris Barber (in SF)
Chris Barber (in SF)@chrisbarber·
I asked a few people to predict: frontier vs non-frontier AI models for coding at SF startups, as of end of 2028 - which gets more of the token volume, and how big is the price premium per token? (Non-frontier here also includes cheaper closed-weight models from frontier labs, and tokens routed to by a frontier model. The reference class here is SWEs at SF startups, which influences the outcome.) Upper left (frontier medium+ premium, 50%+ share): - Nick Mehta (@nrmehta): "I guess for me it comes down to how much you believe in the bitter lesson and generalization. I guess I have a worldview similar to Demis - that all information and perhaps the universe itself is simply a high dimensional vector space. In this way, I believe the problems to be solved will be so much more powerful at the frontier." - Michael Stajer (@michaelstajer): "new startups building new business models using self improving code linked to business evals… more eval focused and self looping ai coding cycles… ai writing code >> evaluating business results >> changing code >> evaluate >> repeat. Mimic of frontier labs self improving AI, startups who build business models on top of ai self improving code will have outsize advantage. What seem like small gaps between frontier and everyone else become large advantages over hundred/thousands/millions of cycles. Also — I think non-frontier will stall as frontier gets better at preventing distillation." Upper right (frontier medium+ premium, <50% share): - prinz (@deredleritt3r): "1. I think the frontier labs will be way ahead of their competitors, and will be able to charge a high premium for their very best models. 2. I think the frontier labs will also offer cheap/fast models that will be very difficult for others to compete with. 3. I think that the combination of these cheap/fast models becoming really good at basis "meat-and-potatoes" tasks AND potential compute constraints will mean that people will use the cheap/fast models more." - Arun Rao (@sudoraohacker): "Right now I'd lean to OAI and Anthropic keeping their premium and leads in frontier coding and all the OSS models being a distant second for workhorse dumber flows-but the main issue is product work (choosing good domains) and verification, and then all the RSI uncertainty." - Super Dario (@inductionheads): "I think frontier companies will have hybrid services where most tokens are served by smaller models. But won't be per task selection by user. Will be more like automatic per token selection ala moe" - Simon Grimm (@Simon__Grimm): "Sticking to my thinking where markets for AI will be like markets for knowledge work, this is how I'd expect it to cash out." - Herbie Bradley (@herbiebradley): "by 2028 in SWE specifically I think things will get to the point where routing is much more sophisticated and most SWEs are using *by token volume* predominantly non-frontier models (eg more complex versions of the Fable advisor Opus implementor pattern). but in terms of $, most $ might end up being frontier simply because of the disparity in cost? in domains outside SWE eg which are harder to verify I expect 2028 frontier to have more of a USP" - Eran Sandler (@erans): "It's going to be a mix. So price wise and token wise can change because frontier is more expensive and planning phases might consume more tokens. But overall token usage is - frontier might end up being a smaller portion. Probably 40% and 60% will end up being non frontier. If you go to places that are not SF startups - like regular enterprises that will look even less frontier and more something else that is good enough for most tasks" Lower left (frontier small to no premium, 50%+ share): - Soren Larson (@hypersoren): "My intuition is that –– for writing code specifically –– so long as there are venture dollars available, startups will continue to use the leading models to write software. Venture is in the game of big wins not minimizing risk –– no early stage venture investor would say "how about you spend less on tokens." That said, coding models are the basis for AGI and likely agent labor. If you were to expand your question by deleting "for coding" in your prompt, then you implicitly ask about the distribution of agent labor application and the associated economics, which I think would produce a different result" Lower right (frontier small to no premium, <50% share): - P Equity Research (@pequityresearch): "I think most of the high volume coding will be capable on non-frontier models, not requiring as much of the frontier model use. You may have some use cases where frontier excels and thats where a small premium exists per token, but most of the work will be automated with good results via non-frontier." - Finn Hulse (@finn_hulse): "i think coding will be more or less a solved problem by end of 2028. premiums can't be that high; code can only be so "correct" before marginal gains just aren't worth the premium. so against backdrop of plummeting costs i don't think frontier models will command a huge premium. this also matches my general sense that premiums will not get out of hand/there won't be a "god" model because ROI is all that matters, and that ratio always approaches reasonable efficiency. i think startups will continue to be happiest to pay for higher premiums simply because they are at the riskiest end of the curve where ROIs might warrant higher premiums (i.e. it makes no sense to hedge or settle for less than the best when attempting something ~impossible). but by the argument that most models will become good enough to solve any reasonably defined coding problem, a good amount of subagent work will be handed off to cheaper sub-models. i honestly think the big labs are going to have to start natively blending models even more efficiently in a way that sort of blurs the question itself. like a single "frontier" query might consume resources commensurate with 90% smaller models and 10% larger models, just out of the box due to scaling limits/brutal competition/diminishing marginal benefits of extra intelligence at routine substeps." Thanks to @alexolegimas @deredleritt3r @erans @finn_hulse @herbiebradley @hypersoren @inductionheads @michaelstajer @nrmehta @pequityresearch @Simon__Grimm @sudoraohacker.
Chris Barber (in SF) tweet media
English
3
6
40
20.1K
Eran Sandler retweetledi
Igor Zalutski
Igor Zalutski@IgorZIJ·
Today we shipped deployment of agents from repos Now all you need to do to deploy a Flue agent on OpenComputer is install a github app and pick a repo. it will then deploy on every push - just like you'd expect a web app on Vercel or a container on Fly Give Durable Agent Sessions a try! It's a fully managed, model-agnostic agents API Connect your repo here: app.opencomputer.dev
English
1
3
8
846
Eran Sandler retweetledi
Igor Zalutski
Igor Zalutski@IgorZIJ·
Today we shipped support of Flue in OpenComputer Durable Agent Sessions... AGAIN! Flue agents on OpenComputer now work they way @FredKSchott likely meant it: instead of a dedicated "brains" box for the Flue agent loop, they now run in lightweight edge workers. As a result, Flue agents on OpenComputer now respond faster AND consume much less compute. > oc agent deploy --agent "mySlackAgent" Grab your api key here: app.opencomputer.dev
English
2
5
20
2.7K
Eran Sandler
Eran Sandler@erans·
@IgorZIJ also that integration sometimes doesn't render the buttons that integrated harness should have :-/ - I would not have implemented it that way.
English
0
0
1
13
Igor Zalutski
Igor Zalutski@IgorZIJ·
@erans its genuinely perplexing how Google even manages to achieve this level of anti-usability with models that, while no longer frontier competitive, aren't that bad
English
1
0
0
64
Eran Sandler
Eran Sandler@erans·
Gemini in sheets is the SLOWEST AI I have ever seen. I built a faster, smarter one 18 months ago with sonnet 3.7.
English
1
0
0
263
Eran Sandler
Eran Sandler@erans·
@zeeg I have some thoughts (and solutions) on this because Junior's work is not necessarily "human waiting at blinking cursor". Happy to share them in a DM.
English
0
0
0
281
Eran Sandler retweetledi
Peter Steinberger 🦞
Peter Steinberger 🦞@steipete·
OpenAI hired me, not OpenClaw. The OpenClaw Foundation is independent, with sponsors rather than owners - and, for the first time, a full-time team keeping the claw alive and stable. 🦞 Couldn’t have done it without Dave and the team. Kudos!
Dave Morin 🦞@davemorin

Today we’re introducing the OpenClaw Foundation: a nonprofit home for open, independent personal AI. A full-time team. Great partners. One mission: bring personal AI to everyone. Welcome to the age of the lobster.🦞 openclaw.ai/blog/introduci…

English
89
110
3.7K
410.1K
Eran Sandler retweetledi
Dave Morin 🦞
Dave Morin 🦞@davemorin·
Today we’re introducing the OpenClaw Foundation: a nonprofit home for open, independent personal AI. A full-time team. Great partners. One mission: bring personal AI to everyone. Welcome to the age of the lobster.🦞 openclaw.ai/blog/introduci…
English
99
161
1.9K
855.6K
Eran Sandler
Eran Sandler@erans·
AI Agents (coding or not) are not humans. They work on a different time scale which means a lot of assumptions can be made that are not relevant for humans.
English
0
0
1
121
Eran Sandler retweetledi
Igor Zalutski
Igor Zalutski@IgorZIJ·
Today we shipped Schedules as part of Durable Agent Sessions on OpenComputer You can now start your agent sessions automatically on a cron in addition to inbound Slack messages or API calls. Grab your API key and try it! app.opencomputer.dev
English
0
5
11
1.6K
Eran Sandler
Eran Sandler@erans·
@davidcrawshaw try working on kernel modules! Opus 4.8 did the same but after a while I managed to get Anthropic to approve me to the cybersecurity thingy - but that doesn't follow you to Fable and its basically useless to me when working on kernel stuff.
English
0
0
3
244
David Crawshaw
David Crawshaw@davidcrawshaw·
And of course the very first thing off my "try when fable is back" list was some exotic work to improve ssh reliability, which was deemed too unsafe for me to be allowed to do. Being treated like a child is deeply unpleasant.
English
7
0
95
6.3K
Eran Sandler retweetledi
Utpal Nadiger
Utpal Nadiger@utpalnadiger·
Every enterprise we talk to is building their own model agnostic internal agents (stuff like Claude Tag but homegrown and multi channel - @tryramp has inspect, @WorkOS has horizon, @coinbase has mux). It's table stakes now. Today, we shipped a simple way to build a claude tag alternative for your org. Anyone in the team (think sales, customer support, gtm, pm's etc) can @mention it in a channel to start a session and steer it from there. It runs on opencomputer durable agent sessions (docs.opencomputer.dev/agent-sessions…). We'd love any and all feedback. We're very excited for this!
Igor Zalutski@IgorZIJ

Today we shipped Slack support in Durable Agent Sessions You can give every agent you build on OpenComputer its own Slack handle Now you can create an agent like Claude Tag or Devin without writing any code! API and SDK also work ofc Try on opencomputer.dev

English
2
3
23
3.9K
Ryan Carson
Ryan Carson@ryancarson·
I'm starting to hit $15-20k per month in token spend for engineering - just for myself. Next month I'll be looking to implement the kinds of things that Brian is doing here at Coinbase. Most likely switching to GLM 5.2 as default and only using frontier models for harder tasks. I can probably get that $20k down to <$5k pretty easily. I'm pretty sure we'll see everyone doing this. It's just not financially viable to do everything with frontier models This is another reason I think we'll see people move away from choosing a lab for their harness (CC or Codex) and move their code factories to in-house agents like @tryramp or agent labs like @DevinAI @FactoryAI @cursor_ai @AmpCode The labs are not incentivized to drive down your token costs
Brian Armstrong@brian_armstrong

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.

English
183
56
1.1K
311.5K
Eran Sandler
Eran Sandler@erans·
@chrisbarber @davemorin Unless they move to assets capex by renting a fixed monthly cost rig or buying one (and even get it depreciate). Then the game is to extract as many tokens per second
English
1
0
1
55
Chris Barber (in SF)
Chris Barber (in SF)@chrisbarber·
I'd love to see LLM benchmarks move to reporting intelligence per second and intelligence per dollar, perhaps by re-running the eval with time and spend limits For enterprises paying API prices, intelligence per dollar is probably the metric For subsidized users, int/s
English
4
5
14
1.9K
Harshil Mathur
Harshil Mathur@harshilmathur·
I’m sure GLM 5.2 is great, but how are people coding with models that can’t see? When something breaks, are you just copy-pasting errors and describing issues manually? That sounds painful in general, but especially for frontend.
English
123
7
420
65.8K
Eran Sandler
Eran Sandler@erans·
GLM 5.2 (like everyone else) but running a dedicated 8xB200 with NVFP4 checkpoint and MTP can easily get to 180-200 tok/sec (running in high effort). Makes you reconsider what is the actual speed most of us get from GPT 5.5 and Opus 4.8 (around 10 tok/sec-ish most of the time).
English
1
0
4
222
Eran Sandler
Eran Sandler@erans·
@illscience mixing those two on the same rig will make it less efficient and that cache (for the most part) resides in the HBM memory of the GPU which is expensive - which directly affects the cost of your rig.
English
0
0
1
151
Eran Sandler
Eran Sandler@erans·
Infra-wise, coding agents are much more cache-friendly than general agents. They usually work in one repo, through the same harness, with the same system prompt, tools, repo instructions, style guides, tests, and working context repeated over many turns. That creates a large stable prefix. Claude Code / Codex-style agents carry a big fixed prompt + tool surface on every request, and repo work has strong locality: same files, deps, tests, errors, branch, build system. So you can optimize around prompt/KV cache, repo affinity, warm sandboxes, cached indexes, and long-lived workspaces. General agents are messier: email, browser, calendar, spreadsheets, support, research, maybe code. The tools, context, and task shape change more often, so cache reuse and scheduling are harder. So roughly: coding agents = repo-local, session-heavy, stable prefix, high cache reuse general agents = task-diverse, tool-diverse, weaker cache locality Separate from "chatting with an agent" - that depends entirely on how stable the context is.
English
1
0
6
237
Anish Acharya
Anish Acharya@illscience·
Strongly agree w claire - in my experience glm 5.2 is better than opus 4.8 and approaching gpt-5.5… and the fact that you can run it locally, albeit on serious hardware, means the enterprise can take token opex cost and convert it into asset capex cost and essentially run the thing 24 hours a day at zero marginal cost… the implications here are badly under discussed
claire vo 🖤@clairevo

I'm now running GLM 5.2 as my default model in claude code + cursor, and it's cost me *checks notes* $3.36 Today's ep of How I AI is my first reviewing an open weights model, @Zai_org's GLM 5.2 which (so far) is giving me Opus vibes without the opus $$$ I cover - how to set up these models in cursor and cc via @OpenRouter API - front end design sense - performance on a long running autonomous task The moment it won me over? When it put chatprd pink in my docs without me having to ask A huge ty to our special sponsor @mercury - Radically different banking loved by over 300K entrepreneurs Full ep on youtube: youtube.com/watch?v=ZoBfQZ…

English
25
29
312
70.1K