Alessandro Rodi

79 posts

Alessandro Rodi

Alessandro Rodi

@alrodi

eng @coinbase, card team. my views are my own. i sometimes sink way too many hours into OSRS. proud scottish terrier owner

Montreal, Quebec Katılım Mayıs 2011
132 Takip Edilen57 Takipçiler
Sabitlenmiş Tweet
Alessandro Rodi
Alessandro Rodi@alrodi·
some of the stuff we shipped in H1 on the card team @coinbase mux: an internal tool that lets you code with any harness or cli that coinbase eng have access to. mux also helps with token spend. (see @brian_armstrong’s recent post on AI spend) for example: - use one opus 4.8 session for planning and route to @cursor_ai composer 2.5 for implementation by spawning new mux sessions on different worktrees rolled it out across the company. read more here -> coinbase.com/blog/coding-ha… card agent on slack, which basically turned into another teammate: - every page or monitor alert is triaged by Card agent and it effectively builds an intelligence to help for future pages - you can ping it for investigations and it retains memory from past interactions - runs our weekly ops reviews - logs every bug, triage, page and monitor alert into linear, which doubles as its memory - if it sees a page or monitor it can fix with code, it uses our in-house software factory (Forge @chintanturakhia) to do it partnered with agentX @kylecesmat @siddharthkul in the platform org(@rwitoff) on a slack multiplexer that lets any team spin up a slack bot trivially with only one slack app. AI advisor bot @teghsm and more! product side: - travel portal: 5% bitcoin back on flights, hotels and cars, no monthly cap. lives in the app, built with rocket travel - secured card: post $500 to $5,000 in usdc as collateral to get approved with no credit history, and you still earn rewards on the deposit - web onboarding - card management improvements - many more improvements and more coming :) feels like we were an army of 100, but we're ~9 ICs in the card team with @vladberteanu. the company makes it easy to partner across orgs and ship together, which is most of why any of this works.
English
4
1
6
688
dax
dax@thdxr·
@MarcosHernanz woah! how come you're not more successful?
English
81
33
3.3K
78.5K
froem
froem@karstenfroemm·
The model lobotomy happened a lot faster this time around. SOL and Grok are back to being brainless.
English
1
0
1
26
Alessandro Rodi
Alessandro Rodi@alrodi·
@flowstated awesome. Is that also possible with Cursor cloud agents or its local only ?
English
0
0
0
635
erik
erik@flowstated·
fable 5 🤝 composer 2.5 fast
Stockholm, Sweden 🇸🇪 English
98
51
2.3K
216.1K
Old School RuneScape
Old School RuneScape@OldSchoolRS·
this mansion is gigantic... looks like something fit for the master of unlocking ⛱️
English
30
16
461
49.5K
Alessandro Rodi
Alessandro Rodi@alrodi·
@trq212 @bcherny Feedback for claude mobile : Add a “new repo” button Sometimes I want to create a new private repo under my Github and I want to do it directly in claude
Alessandro Rodi tweet media
English
0
0
1
20
Alessandro Rodi
Alessandro Rodi@alrodi·
you can use AI efficiently
Brian Armstrong@brian_armstrong

How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.

English
0
0
2
105
Alessandro Rodi
Alessandro Rodi@alrodi·
it is fun building here. we have the choice to route to frontier models for complex tasks and also have access to powerful ways of routing for example, in one of our internal tools Mux we can choose to plan with opus 4.8 and route the implementation by spawning sessions on different git worktrees using composer 2.5 @cursor_ai CLI
English
1
0
3
369
Brian Armstrong
Brian Armstrong@brian_armstrong·
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching. Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting to open weight models like GLM 5.2 and Kimi 2.7 through our LLM gateway, while still encouraging engineers to choose the right model for the task. 91% of our employees were never hitting their usage caps, so instead of lowering caps and driving up alerts, we're moving to cheaper defaults. Note that code reviews use a diversity of models, so they can check each other's work. Better Routing – In our custom harnesses, we preprocess prompts and route to the best model for the job, considering cache hits and model pricing. For instance, you may want a frontier model for planning, but not for execution where they can be overkill. Ultimately, humans shouldn't be choosing models - AI can automate this task. Better Caching – Cache misses are the easiest way to drive your cost up. All of our requests are cache aware, so we’re reusing a warm cache wherever possible. For example, our cache hit rate went from 5% → 60% in LibreChat once properly implemented. Keep Context Lean – Start fresh sessions when switching tasks. Scope file context narrowly. Disconnect unused tools. Don't just compact. The goal isn't fewer tokens used, it's fewer tokens wasted. Better Visibility – Our engineers can use as many tokens as they want, from whatever model they want, but we’ve made usage visible – and the more you spend on AI, the more impact we expect. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. Putting this into practice has cut our AI spend nearly in half, while our token usage continues to grow.
Brian Armstrong tweet media
English
475
745
6.2K
4.2M
OpenAI
OpenAI@OpenAI·
GPT-5.6 Sol is our most capable model yet for cybersecurity. It shifts the performance-efficiency frontier for long-horizon security tasks including vulnerability research and exploitation.
OpenAI tweet media
English
106
216
3.4K
1M
OpenAI
OpenAI@OpenAI·
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work. openai.com/index/previewi…
English
3.6K
5.7K
40.7K
18.5M
Alessandro Rodi
Alessandro Rodi@alrodi·
hmm.. maybe try setting your context better "literally ignores half the details" either means your task is too big or your LLM is clueless in your repo, which means you need to work a bit before the LLM knows what to do for example : - keep all signals in @linear (bugs, tasks to do, docs, projects) - have a simple agents.md that points to different context in your repo - let your LLM search your PR history to see the foundation (If you have more than one repo, sourcegraph is good in my experience) you helped me a lot with LeetCode in the past so i’m glad to return the favor :) the loop stuff will only work after manually prompting works really well with your repo and setup
English
0
0
0
41
NeetCode
NeetCode@neetcode1·
I'm starting to feel like im stuck in a loop of my own new model -> hype -> genuine improvements -> i overestimate what it can do -> buggy features or just complete slop Did I do something wrong? 1. I created the most detailed plan w /grill-me-with-docs 2. Set up an elaborate loop so it would continuously review the plan after implementation, until it implemented each step (the biggest problem for me is it literally just ignores half the details) 3. Gave it browser access so it could test and QA The end result: I open it and my browser crashes. And half the functionality is missing. Yeah, AI can definitely replace jr devs now. Idk about you guys, but I never saw a jr dev straight up lie that they implemented something. This was for a greenfield feature on a new page, no existing dependencies.
NeetCode@neetcode1

man i really miss writing code i miss hopping between files, reading error messages, writing shit code and then refactoring it later.. and then refactoring it again and again.. agents are nice but i just spent three days fixing a feature i probably could've just coded from scratch more quickly gpt 5.4 is my model of choice these days, and it's pretty good but i think i fell for the bait that i didnt need to read the code anymore but idk maybe its a skill issue

English
44
12
397
38.9K
Aryan
Aryan@justbyte_·
what’s your favorite VS Code extension that every developer should know about?
Aryan tweet media
English
26
1
27
3K
Alessandro Rodi
Alessandro Rodi@alrodi·
@mitsuhiko also curious - my experience with loop engineering was not successful except for watching CI or monitoring for other mundane tasks. Have not found a way to use loops to produce good code in a project.
English
0
0
0
1.8K
Armin Ronacher ⇌
Armin Ronacher ⇌@mitsuhiko·
I decided to do some experiments with looping over the weekend. The only cases where they work so far for me are a) review b) research c) autoresearch. If someone uses them for actual implementation on a medium sized project, would love to have something to look at!
English
61
6
426
157.6K
Alex Craviotto
Alex Craviotto@craviottx·
This city is amazing, winter will be hard though.
Alex Craviotto tweet media
English
39
24
862
28K