Steve Martocci
4.1K posts

Steve Martocci
@smart
Co-Founder @joinSuppCo, @Splice, @GroupMe, @FlyBlade, and at it again this time as a new dad. Mostly Harmless.




How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.


JUST IN: Capital Factory co-founder and CEO Joshua Baer was killed Tuesday night when a small business jet crashed onto a highway near Laredo, Texas, according to the Austin-based startup accelerator. cbsaustin.com/news/local/cap…


Claude Fable 5 is now available in Devin. Fable 5 earns the #1 spot on FrontierCode, our benchmark for real-world engineering tasks that grades mergeability and quality:






1/ We’ve raised over $1B at a $26B valuation, led by @Lux_Capital, @generalcatalyst, and @8vc. Our enterprise usage has grown >10x since the start of this year, and our run-rate revenue grew to $492 M. We launched Devin two years ago as the first AI software engineer. Since then, cloud agents have gone from niche to mainstream, and today they are the fastest growing way to create software.





This post has 3.5MM views (and counting) and is totally bogus. The author doesn’t know how to read a test result. His cited results (in app behind paywall) show premier protein UNDER the prop 65 limit for lead. All of them in fact meet the safety standard. He doesn’t know the difference between ppb in a powder, which is the amount of lead in a KILOGRAM of material, vs dose per serving - which is what all safety levels are set at. There’s also blatant typos in carrying over the test result (has mixed up premier and ritual readings, from the wrong part of the test). You’re fine eating your protein shakes. Watch what you fall for on the internet. This stuff is driving up anxiety for no reason.











