Michael R. Bock@michaelrbock
Can Chinese AI models file American tax returns?
Kimi K3 made waves with its release last week. We tested Kimi K3's ability to calculate 50 hyper-realistic federal and state tax returns.
The result: not great.
Compared to GPT-5.6 Sol's 58% and Fable 5's 4%, Kimi K3 only scored 6% on our strict TaxCalcBench eval.
That means Kimi K3 still has a long way to go.
The narrative last week was a Chinese model had finally caught up to the American labs' frontier.
But our scores show something more nuanced.
Kimi K3 might be great at frontend coding, but clearly not American tax calculations.
I suspect evals and research focus matters quite a bit.
In the US, researchers are aware of and focused on distinctly American knowledge work tasks like tax filing.
The story of models converging clearly isn't fully true, at least not today.
There's one caveat: GPT-5.6 Sol and Fable 5 used web search tools, which helped their scores quite a bit. By comparison, Kimi K3 (via OpenRouter) did not.
That being said, Chinese AI has more work to do if it wants to catch up!