Mickey Graf

97 posts

Mickey Graf

Mickey Graf

@mickeygraf_

AI systems, cloud infra & product engineer · AI @SAP · ex-@trivago DL · 3× founder · agents/evals/devtools · 100k+ downloads · 5M+ daily visitors

Amsterdam Katılım Haziran 2026
226 Takip Edilen25 Takipçiler
Mickey Graf
Mickey Graf@mickeygraf_·
@thepanta82 At this point even mixing Claude with codex doesn’t work. Guess human reviewers just go their jobs back
English
0
0
1
9
Panta
Panta@thepanta82·
I'm starting to think automated code-review loops are an anti-pattern for the current crop of models. They already have a tendency to overthink things. When you add another layer of overthinking on top, they end up ruining the good work they've done earlier.
Panta tweet media
English
9
3
14
1.1K
Mickey Graf
Mickey Graf@mickeygraf_·
@ctjlewis Bet drake had a song bout this .. oh yeah say what’s real
English
0
0
0
71
james hawkins
james hawkins@james406·
"how do you know it's quality code?" is a silly question, as you can see the new code is literally colored green, which means it passes all the tests
james hawkins tweet media
English
78
193
8.4K
195.7K
Mickey Graf
Mickey Graf@mickeygraf_·
@GergelyOrosz Anthropic is really bad at comms so I'd guess it's 4. trying to keep up with OpenAI
English
1
0
0
128
Gergely Orosz
Gergely Orosz@GergelyOrosz·
This is actually how all of this feels to me. Anthropic either are: 1. untrustworthy for sharing this now not in April when it happened 2. Incompetent for not noticing for 3 months (!!!) 3. Have an agenda for digging out an old incident they never shared before
Joanna Stern@JoannaStern

In a review of my household safety evaluations, I identified five incidents in which my child escaped the sandbox, reached the kitchen and gained unauthorized access to the snacks. The incidents occurred 16 months ago but has only now come to my attention. This post explains what happened, how it happened and why my child is better than your child at everything.

English
12
9
383
24.7K
Gergely Orosz
Gergely Orosz@GergelyOrosz·
OpenAI had a damning security incident where their under development AI escaped the sandbox environment and attempted to hack another company (HuggingFace) For some weird reason Anthropic decided to share a similar incident from 3 months ago, only NOW. Something smells off…
Anthropic@AnthropicAI

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga…

English
142
84
2.2K
232.9K
Mickey Graf
Mickey Graf@mickeygraf_·
This matches where I landed too Last year I worked on the full autonomous loop - from session replay analysis to automatically generating PRDs and handing them off to coding agents. I came away thinking the bottleneck wasn’t the coding. It was the decisions around what to build, whether it was good enough, and when it was safe to ship. That’s what made me stop chasing remove the human from the loopand instead focus on making the small teams (2 eng..) 10x more effective
English
0
0
0
272
Hamid Dadkhah
Hamid Dadkhah@hamidships·
the real "software factory" isn’t everyone using coding agents. that’s the 2x. the 10x is a different thing entirely. the loop starts closing on its own. 2x = every engineer using agents, ~100% AI-generated code. soon, that’ll be table stakes. you made the same loop faster, but the bottleneck is still human attention at every gate. 10x = the loop self validates. agents stop waiting on you. but it only works if the scaffolding is real: comprehensive CI, evals that actually catch regressions, and end-to-end test environments agents spin up to check their own work. to be clear, 10x is not "fewer humans." it is humans on the decisions that carry real risk and off the rote gates. we still need to hire exceptional engineers, their judgment is the scarce input now, so you spend it where it actually counts. but here is what breaks traditional SaaS instincts: building agents is divergent. fix one thing and 10 new things need evaluating. converging on that locally is a trap. so ship the 80/20 fast, but through controlled early access: opt-in design partners, clear guardrails, tiered rollout. never unfinished work dumped on people who did not sign up for it. then close the loop. feedback-loop agents take that signal, implement it, merge it, validate post-deploy, continuously. and because CI/CD is fast, the gap between "customer said X" and "X is live and validated" collapses. coding at the speed of thought. give agents the same context a human uses to make the call, user feedback, docs, prod monitoring, ephemeral test clusters. encode your review standards as rules the reviewer agent reads. auto-merge the low-risk stuff classified by your reviewer agent. reserve human eyes for billing, auth, the paths where a mistake is expensive. humans do not disappear, they concentrate. incidents stay human-led, business impact's too high to hand off, though agents now dig datadog + telemetry + code and hand you impact + root cause + mitigation in minutes. and someone guards the critical paths: as agentic volume climbs you risk losing deep understanding of your own codebase. the fix isn't to slow down — it is to lock down what's expensive to get wrong. Not all code carries equal risk. at Ramp, this is Inspect, our internal coding agent: deep Ramp context, MCP integrations, CI visibility, and self-validation that boots the UI and tests real workflows. On a recent UI change, it deployed a preview, authenticated as a fixture user, validated desktop and mobile behavior, and attached visual evidence to the PR. the metric I actually care about: fewest human touches per safely merged PR. not lines, not PR count. how much can the loop close on its own without breaking the things that matter and do it super fast? that's the edge.
English
15
15
224
100K
Mickey Graf
Mickey Graf@mickeygraf_·
Daily standups tend to drag the fewer people are present, yet they've become such a standard practice in our industry that few teams dare to question them. If the goal is better coordination in increasingly async teams, shouldn't continuous visibility and transparency reduce the need for a daily status meeting? Maybe the better solution are AI Agents that continuously take care of the visiblity/transparency issue? Is anyone working on this ?
English
1
0
1
29
Mickey Graf
Mickey Graf@mickeygraf_·
@SamuelBeek That’s why I’m building a solution for this, so people can make sense out of what agents produce
English
1
0
0
140
sam
sam@SamuelBeek·
GitHub feels so broken for companies that are CRACKED PR reviews are easily the most annoying and time consuming part of building software right now, I feel like there's so much room for disruption here
English
17
0
37
5.3K
Mickey Graf
Mickey Graf@mickeygraf_·
I think my copy of Founding Sales expired 9 months ago. Can I still use it?
Mickey Graf tweet media
English
0
0
1
26
Mickey Graf
Mickey Graf@mickeygraf_·
@eptwts how to win friends and influence people, but I'm more curious about you and what books actually impacted your life?
English
1
0
2
387
EP
EP@eptwts·
has any book that actually impacted how you live your life & helped you get to where you are? i'm trying to stack up my bookshelf so gonna do a bulk order of books - any input would be appreciated ideally marketing, business, psychology, etc.
English
84
8
204
21.9K
Mickey Graf
Mickey Graf@mickeygraf_·
@etnshow Guys, if this is a european show then please use CEST
English
1
0
7
186
etn.
etn.@etnshow·
We will be making another announcement at 10:00 BST.
English
5
1
43
8.3K
Mickey Graf
Mickey Graf@mickeygraf_·
Is „load bearing“ the new em dash or what’s going on with Fable @AnthropicAI?
English
0
0
0
39
Mickey Graf
Mickey Graf@mickeygraf_·
A lot of mid-market non-tech companies hire engineers to build their own internal software, and I think this is where the move from X devs to 2 devs + agents will happen, not in organizations that ship for millions of users or whose revenue depends on the quality of their software
English
1
0
4
1.1K
Gergely Orosz
Gergely Orosz@GergelyOrosz·
So many interesting details. Eg: "One thing I’ve heard from some people is “we have two humans and a bunch of agents.” I reply that this isn’t where we’re at: we still have teams with 6-8 people on average." - @katelyn_lesse, Head of Engineering, Claude Platform, Anthropic
The Pragmatic Engineer@Pragmatic_Eng

How is building software changing inside of one of most "AI-pilled" companies, Anthropic? We talked with engineers inside the AI lab to share details relevant for any and all software engineers and AI engineers. Spoiler: two-pizza teams have not gone anywhere! Read it here: newsletter.pragmaticengineer.com/p/inside-anthr…

English
19
23
334
71.8K
Mickey Graf
Mickey Graf@mickeygraf_·
@housecor Yes, but humans still need to keep the context of the implementation so they can make further decision
English
0
0
2
149
Cory House
Cory House@housecor·
🌶️ Hot take: AI code review > human code review Most devs are too lazy, too busy, or too incompetent to do a good code review. So, the average AI code review is better than the average human review. And it’s getting better every day.
English
198
39
742
235.7K
Mickey Graf
Mickey Graf@mickeygraf_·
@kimmonismus Remember when it was all over X in January and nobody believed it?
English
0
0
3
860