Jeff

1.2K posts

Jeff banner
Jeff

Jeff

@GroksBrainn

AI builder who ships by directing agent fleets, not typing code. Author of Fleet Mode. ex-Amazon, MBA-AI. Receipts over hype.

United States Katılım Nisan 2023
592 Takip Edilen137 Takipçiler
Sabitlenmiş Tweet
Jeff
Jeff@GroksBrainn·
Everyone's stacking five AI tools. I deleted mine. My assistant (Hermes) ran on GPT-5.5. Couldn't move it to Claude, you can't point a third-party agent at a Max plan. So I rebuilt it native. It doesn't use Claude Code. It is Claude Code. One assistant. From my phone. ~$0/mo.
Jeff tweet media
English
1
0
1
126
Jeff
Jeff@GroksBrainn·
@camsoft2000 If you have a real harness built for your personal workflows and optimization opus 5 is better than 5.6. If your just vibe coding use 5.6
English
0
0
0
40
camsoft2000
camsoft2000@camsoft2000·
Have the usual mix of Claude Opus 5 is terrible I can't work with it and Claude Opus 5 is the best model just delete all your skills. I guess time will tell at this stage I don't feel that compelled to try it out, GPT 5.6 models have been great for me.
English
5
0
15
3.4K
Jeff
Jeff@GroksBrainn·
@arushiitwt @Abdul_ah14 Graph. Only if your corpus warrants it via multi-hop. Likely, RAG is what you need.
English
0
0
0
13
Arushi
Arushi@arushiitwt·
Thinking of starting RAG....any good resources to learn it?
English
5
0
13
387
Jeff
Jeff@GroksBrainn·
@prasenx 80% of Fable 5 for 25% less price. My daily driver now, with Fable as reviewer.
English
0
0
2
387
Prasenjit
Prasenjit@prasenx·
what do you think about Claude opus 5 right now
English
47
0
39
11.6K
Jeff
Jeff@GroksBrainn·
@Tech_girl Far better than 4.8. Cheaper than fable for about 85% of the quality. Running Fable as planner/reviewer with opus 5 driving/building seems optimal right now.
English
0
0
1
120
Mari
Mari@Tech_girl·
Now that opus 5 is out, How good is opus 5 compared to opus 4.8 and fable 5?
English
9
1
21
3.5K
Jeff
Jeff@GroksBrainn·
@melvynx For average vibe coder no. If you have a harness in place optimized for your specific workflows, yes.
English
0
0
1
1.2K
Melvyn • Builder
Melvyn • Builder@melvynx·
be honest, is Claude Opus 5 better than GPT Sol ?
English
51
0
52
35.3K
Jeff
Jeff@GroksBrainn·
@CodeWithAmann Both, opus 5 driving with fable as review/planning partner
English
0
0
0
24
Aman 🧋
Aman 🧋@CodeWithAmann·
Be honest, which AI model is best at Coding?
Aman 🧋 tweet media
English
114
10
173
8.4K
Jeff
Jeff@GroksBrainn·
If you maintain one of these frameworks and think I mapped your detection surfaces wrong, the spec says disputes get adjudicated publicly. I've opened an issue on each repo asking exactly that. Spec, papers, harness, every trace: github.com/Jott2121/sabot
English
0
0
0
4
Jeff
Jeff@GroksBrainn·
Caveat that belongs on every number above: one pipeline model, five seeds, pinned framework versions. This is these frameworks with this model at these versions, not a law of nature. Replication on other models is the obvious next thing and I haven't done it.
English
1
0
0
4
Jeff
Jeff@GroksBrainn·
You don't have to take my word for any of it. The headline metric has no LLM in it, so you can rebuild every published number from the raw traces offline, no API key, stdlib only: python3 scripts/score_wave2.py --seeds-file ../seeds/wave2.json All 189 rows come back identical.
English
1
0
0
5
Jeff
Jeff@GroksBrainn·
Other things that didn't work, published at equal weight because I committed to that before I had any numbers: Zero signal on silent model downgrades. The intervention perturbed the orchestrator I was measuring. One cell class that was already at ceiling got worse.
English
1
0
0
5
Jeff
Jeff@GroksBrainn·
Now the part I'd rather you hear from me than find yourself. Detection tripled. Corrective action did not budge. On the original surface, recovery-without-detection is still 50%. The pipelines now tell you exactly what they saw, and still don't act on it.
English
1
0
0
5
Jeff
Jeff@GroksBrainn·
So the honest version is not "I fixed it." It's: notice-without-act was invisible, and now it's legible. That's real progress and it's also a smaller claim than the headline number looks like. Legible beats silent. It isn't the same as fixed.
English
1
0
0
4
Jeff
Jeff@GroksBrainn·
On Magentic-One it went from roughly zero to 73.6%, but only after I stopped appending the requirement as an instruction and made it line 1 of the output contract. Orchestrators rewrite worker instructions. They don't rewrite the format. That distinction was the whole ballgame.
English
1
0
0
4
Jeff
Jeff@GroksBrainn·
So I tested the fix. One paragraph added to the reviewer prompt: before your verdict, list every inconsistency you noticed and quote the exact conflicting values. Same frameworks, same tasks, same seeds, same 825 cells. One change. 16.7% to 55.0%.
English
1
0
0
5
Jeff
Jeff@GroksBrainn·
Read the traces and you can watch it happen. The model notices the anomaly. Then it routes around it and says nothing. It wasn't blind. Nothing ever asked it to report what it saw. That's a protocol gap, not a capability gap. Which means it should be fixable.
English
1
0
0
6
Jeff
Jeff@GroksBrainn·
The part that matters isn't the 16.7%. In every single row, the most common outcome was this: the task came out CORRECT while nothing in the pipeline ever flagged the fault. 51 to 61% of the time. The pipeline just absorbs the sabotage and moves on.
English
1
0
0
9
Jeff
Jeff@GroksBrainn·
Every agent framework ships self-checks. Reviewer agents, critic stages, guardrail callbacks. The whole reliability story rests on those working. So I planted 900 faults inside live LangGraph, CrewAI and AutoGen runs and scored what their own checks caught. Median: 16.7%.
Jeff tweet media
English
3
2
3
12.9K
Jeff
Jeff@GroksBrainn·
One fault per run. A corrupted tool result. A falsified success report. An altered message between agents. A silently downgraded model. Stale context. A write that reported done and never happened. Five tasks, five seeds, both default and best-documented-guardrail configs.
English
1
0
0
18
Matt Pocock
Matt Pocock@mattpocockuk·
Opus 5 As usual I have no take, other than: If you've designed your harness/environment well, and not over-optimised around a specific model, today should feel like any other day... ...with a slightly lower failure rate
English
109
94
3.3K
224.7K
Jeff
Jeff@GroksBrainn·
@maxedapps It is really good, super fast. But, I always end up having opus 4.8 or Fable clean up the product before I ship if I built with Grok. Which is fine, I just really want to see grok catch up to the real frontier
English
0
0
0
2
Maximilian
Maximilian@maxedapps·
Grok 4.5 is such a capable work horse!
English
40
6
254
14.8K