Michael R. Bock

4.4K posts

Michael R. Bock banner
Michael R. Bock

Michael R. Bock

@michaelrbock

co-founder @columntax (acq). prev @waymo @pioneerdotapp @google. see my popular tweets using this side project i made: https://t.co/fdDI6illyy

San Francisco, CA Katılım Kasım 2009
3K Takip Edilen3.2K Takipçiler
Sabitlenmiş Tweet
Michael R. Bock
Michael R. Bock@michaelrbock·
1/ After 5 years, I’m proud to share that @ColumnTax has found a new home. We’ve been acquired by @AiwynAI. I couldn’t be more sure this is the right move for our business, tech, and team. These pics are the moments we started & sold the company:
Michael R. Bock tweet mediaMichael R. Bock tweet media
English
17
6
78
13.6K
RaoulDuke
RaoulDuke@RaoulDukeDegen·
@michaelrbock they made opus 5 default on max so folks will use it more day to day
English
1
0
1
20
Michael R. Bock
Michael R. Bock@michaelrbock·
Anthropic just released its 4th model in 2 months: Opus 5. It's half the cost of Fable 5. And surprisingly, for knowledge work tasks like tax filing: Opus performs even better than its more expensive counterpart. We tested Opus 5 on TaxCalcBench. Here are the results vs. Fable 5: - Opus 5 w/ web search: 40% tax returns computed correctly (strict) - Fable 5 w/ web search: 35% But only when we turned thinking all the way up. We ran both models with web search on the same 50 realistic federal and state tax returns at five thinking levels: - Lowest thinking: Fable won, 18% to 10% - Low: tie, 16% to 16% - Medium: tie, 24% to 24% - High: Fable won, 32% to 28% - Ultrathink: Opus won, 40% to 34% At ultrathink, the gap was even wider under lenient scoring: 56% for Opus versus 44% for Fable. So the headline isn’t simply “Opus is better than Fable.” It’s that Opus earns its lead when both models are pushed to their maximum reasoning setting. At every lower setting, Fable matched or beat Opus on strict accuracy. All of the results and model outputs are public in TaxCalcBench: github.com/column-tax/tax…
Michael R. Bock tweet media
English
1
0
11
726
Maxime
Maxime@KingMaxime04·
@michaelrbock Which kind of bs benchmark is this again? Why don't you at least force Kimi to use web search?? Or better still compare it to the other models with no web search.
English
1
0
3
69
Michael R. Bock
Michael R. Bock@michaelrbock·
Can Chinese AI models file American tax returns? Kimi K3 made waves with its release last week. We tested Kimi K3's ability to calculate 50 hyper-realistic federal and state tax returns. The result: not great. Compared to GPT-5.6 Sol's 58% and Fable 5's 4%, Kimi K3 only scored 6% on our strict TaxCalcBench eval. That means Kimi K3 still has a long way to go. The narrative last week was a Chinese model had finally caught up to the American labs' frontier. But our scores show something more nuanced. Kimi K3 might be great at frontend coding, but clearly not American tax calculations. I suspect evals and research focus matters quite a bit. In the US, researchers are aware of and focused on distinctly American knowledge work tasks like tax filing. The story of models converging clearly isn't fully true, at least not today. There's one caveat: GPT-5.6 Sol and Fable 5 used web search tools, which helped their scores quite a bit. By comparison, Kimi K3 (via OpenRouter) did not. That being said, Chinese AI has more work to do if it wants to catch up!
Michael R. Bock tweet media
English
3
1
14
2.2K
Mario Filho
Mario Filho@mariofilhoml·
@michaelrbock How can we trust your benchmark if you think it was worth publishing this aberration for engagement farming?
English
1
0
0
77
Michael R. Bock
Michael R. Bock@michaelrbock·
@andrewfang yeah...but even comparing all without web search: - GPT-5.6 Sol: 26% - Claude Fable 5: 26% - Kimi K3: 6%
English
0
0
3
252
Michael R. Bock
Michael R. Bock@michaelrbock·
4/ AI can get 85.6% of scored tax lines right while still missing the exact-match bar on 42% of complete returns. Good, but not good enough for a task that requires 100% correctness. What does this mean for your work? You have to benchmark the full task, inspect the failures, and build deterministic systems around the model.
English
1
0
2
76
Michael R. Bock
Michael R. Bock@michaelrbock·
1/ I've been building tax software for >5 years. Here's where AI is still making mistakes filing tax returns: GPT-5.6 Sol (with web search) is the best-scoring model we've tested on TaxCalcBench v2. It calculated 58% of 50 hyper-realistic federal & state tax returns perfectly.
Michael R. Bock tweet media
English
2
1
10
701
Michael R. Bock
Michael R. Bock@michaelrbock·
4/ You have to benchmark your specific task. For tax filing: we built a custom eval (TaxCalcBench) that's become the industry standard. If you do taxes, you can follow along. If you do another knowledge work task, you should build your own benchmark!
English
0
0
0
49
Michael R. Bock
Michael R. Bock@michaelrbock·
3/ Compare that to Claude Fable 5 which is only at 34%. Claude Code/Cowork may have been first to market, but starting with GPT-5.5, the race to see which company/model is best at knowledge work has been neck-and-neck. What does this mean for your work?
English
1
0
0
86
Michael R. Bock
Michael R. Bock@michaelrbock·
1/ Claude (Code, Cowork, Fable) has all the mindshare right now, but is OpenAI actually the best at knowledge work? Our eval for tax filing says: yes. GPT-5.6 Sol (with web search) just became the best model in the world on TaxCalcBench (v2).
Michael R. Bock tweet media
English
1
0
7
531
Michael R. Bock
Michael R. Bock@michaelrbock·
1/ This AI benchmark will decide Intuit's stock price next year. Last year, I released TaxCalcBench, and it moved markets. Here's how:
Michael R. Bock tweet media
English
3
1
15
467
Michael R. Bock
Michael R. Bock@michaelrbock·
6/ Lots has happened in the last 11 months: new model releases, copycat benchmarks, and lots of accounting startup launches. But one thing remains: TaxCalcBench is the industry standard for measuring AI's ability to file taxes. Check out the new leaderboard: github.com/column-tax/tax…
English
1
0
3
63
Michael R. Bock
Michael R. Bock@michaelrbock·
5/ AI still fails to calculate tax returns accurately on its own. GPT-5.5 is the highest-scoring model. On its own, it only scores 24%. With web search tool use, that jumps to 54%. Good, but not good enough to trust it with a task that requires 100% correctness.
Michael R. Bock tweet media
English
1
0
2
79