Maxime

607 posts

Maxime

Maxime

@KingMaxime04

just a random guy automating stuffs

Katılım Nisan 2022
8 Takip Edilen2 Takipçiler
Maxime
Maxime@KingMaxime04·
@michaelrbock Which kind of bs benchmark is this again? Why don't you at least force Kimi to use web search?? Or better still compare it to the other models with no web search.
English
1
0
3
69
Michael R. Bock
Michael R. Bock@michaelrbock·
Can Chinese AI models file American tax returns? Kimi K3 made waves with its release last week. We tested Kimi K3's ability to calculate 50 hyper-realistic federal and state tax returns. The result: not great. Compared to GPT-5.6 Sol's 58% and Fable 5's 4%, Kimi K3 only scored 6% on our strict TaxCalcBench eval. That means Kimi K3 still has a long way to go. The narrative last week was a Chinese model had finally caught up to the American labs' frontier. But our scores show something more nuanced. Kimi K3 might be great at frontend coding, but clearly not American tax calculations. I suspect evals and research focus matters quite a bit. In the US, researchers are aware of and focused on distinctly American knowledge work tasks like tax filing. The story of models converging clearly isn't fully true, at least not today. There's one caveat: GPT-5.6 Sol and Fable 5 used web search tools, which helped their scores quite a bit. By comparison, Kimi K3 (via OpenRouter) did not. That being said, Chinese AI has more work to do if it wants to catch up!
Michael R. Bock tweet media
English
3
1
14
2.2K
Maxime
Maxime@KingMaxime04·
@teortaxesTex No serious coding benchmark here and I have no idea where they got that deepswe score for glm 5.2 because it's wrong
English
0
0
0
490
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)
Goes to show that Chinese models are still undertrained and don't have enough RL. Computationally modest efforts can still push them harder. Should reduce your prior on the utility of large-scale distillation.
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) tweet media
Lentils@Lentils80

Macaron V1 Venti, the first model to be post-trained on GLM-5.2, is released today It seems to be a decent but not huge improvement over GLM-5.2 in agentic tasks They also released Macaron-V1-Tall which is post-trained from the 35B parameter Qwen 3.6 (so small local model)

English
8
7
118
16.9K
Maxime
Maxime@KingMaxime04·
@scaling01 So you are saying even with all Anthropic safeguards in place. The Chinese still manages to distill in 2 weeks and if GLM distills Gemini?? Yeah I call bullshit on everything. I prefer Most Similar To section in eqbench.com than whatever slop this benchmark is
English
0
0
10
2.4K
Maxime
Maxime@KingMaxime04·
@cosier @bridgebench The speed will be better when the weights is released and other providers host it
English
0
0
0
23
Matthew Cosier
Matthew Cosier@cosier·
@bridgebench Uhm. This was on your post how it was the best thing ever lol. Aged well.
Matthew Cosier tweet media
English
2
0
3
931
Bridgebench
Bridgebench@bridgebench·
Kimi K3 is the slowest frontier model in the world. Makes it practically unusable.
English
92
4
394
36.4K
Maxime
Maxime@KingMaxime04·
@OmedVibeCodes Well expect Kimi to be slower the token per seconds is around 19 tps on the API
English
1
0
0
545
OmedTheVibeCoder
OmedTheVibeCoder@OmedVibeCodes·
But there are some positives: the speed is great. Kimi k3 took around 1 hour and 46 minutes, while Qwen 3.8 Max finished in just 20 minutes—roughly the same as GPT-5.6 High. Fable 5 xHigh at around 50min The Coding Plan also seems very generous, especially for just $16. Overall, I’d place the output quality around GLM 5.2 / GPT-5.6 Luna level.
OmedTheVibeCoder tweet media
English
3
0
52
14.6K
OmedTheVibeCoder
OmedTheVibeCoder@OmedVibeCodes·
Just tested Qwen3.8-Max-Preview on my benchmark. It got absolutely cooked. Not Fable 5 level. Not even remotely close. It doesn't come close to GPT-5.6 or Opus 4.8 Max either. Honestly, marketing it at that level feels incredibly misleading. I wasted my time and money testing this. Qwen, refund me.
English
87
35
1K
181.9K
Maxime
Maxime@KingMaxime04·
@jmbollenbacher The thing is that you are right, opus 4.6 and 4.7 which were genuinely good models score lower than this. Openai models seems to be very good on this benchmark though even Luna is better than opus 4.8
English
0
0
0
116
Maxime
Maxime@KingMaxime04·
@Curline1222 @R2Cdev_ Lmao opus is very unimpressive on this bench and Luna max is absolutely mogging Opus 4.8 max on this benchmark. Unless you also think opus have been benchmarked. Btw this score is better than previous opus except opus 4.8 max
English
0
0
1
38
Curline Zephirin
Curline Zephirin@Curline1222·
@R2Cdev_ Kimi K3 seems to be like Gemini, performing poorly in many unfamiliar domains (compared to GPT & Claude)
English
1
0
2
199
Maxime
Maxime@KingMaxime04·
@ValsAI You won't see Kimi doomers here but if it an obscure benchmark that doesn't matter for 99% of users. They will be there
English
0
0
2
1.4K
Vals AI
Vals AI@ValsAI·
Kimi K3 is the #2 overall model on our in-house Vibe Code Bench at 85.0%. VCB tests a model's ability to go from zero-to-one; creating a web application completely from scratch.
Vals AI tweet media
English
36
71
754
86.9K
Rohan Paul
Rohan Paul@rohanpaul_ai·
Now Kimi K3 ranks #3 on DeepSWE, head to head with Claude Fable & GPT-5.6 Sol. But its about the size of the jump. Kimi moved from 31 points with K2.7-code to 69 points with K3, a 38-point gain. DeepSWE tests whether a coding agent can take a short real-world request, explore an unfamiliar open-source codebase, make substantial changes across multiple files, and deliver working code that passes held-out tests. the tasks are original, long-running, and graded in clean isolated environments. A strong signal of genuine end-to-end software engineering ability rather than simple code generation or memorized fixes.
Rohan Paul tweet media
Datacurve@datacurve

Kimi K3 debuts at #3 on DeepSWE. It's the first open-weights model that delivers frontier-level performance, achieving results similar to Claude Fable and GPT-5.6 Sol.

English
14
17
124
13K
Maxime
Maxime@KingMaxime04·
@tenobrus You guys won't just be honest and it tiring, you suck off Anthropic alot on this space x.com/i/status/20785…
Maxime@KingMaxime04

@scaling01 Opus 4.7 was released on April(which is 3 months ago) by the way unless you are talking of openai models which always perform well on mathematics. Even luna max(61%) is mogging opus 4.8 max(56%) on this bench And we already know from live bench, kimi weakness is mathematics.

English
0
0
0
6
Maxime
Maxime@KingMaxime04·
@scaling01 Opus 4.7 was released on April(which is 3 months ago) by the way unless you are talking of openai models which always perform well on mathematics. Even luna max(61%) is mogging opus 4.8 max(56%) on this bench And we already know from live bench, kimi weakness is mathematics.
Maxime tweet media
English
0
0
6
535
Lisan al Gaib
Lisan al Gaib@scaling01·
Kimi-K3 (max) scores only 39% on FrontierMath Tier 4 7% lower than the best US models from 7 months ago
Lisan al Gaib tweet media
English
61
51
698
138.1K
tower
tower@botincave·
@scaling01 i guess k3 benchmaxxed in coding but never made any progress in mathematics and science.
English
2
0
11
2.9K
BitBrew
BitBrew@BitBrew1·
@scaling01 OpenAI copied Anthropic with everything they’re doing and it’s the actual slop work. Their “everything” app isn’t as integrated as Anthropic.
English
2
0
2
407
Maxime
Maxime@KingMaxime04·
@RyanGreenblatt @Tril1boswagginz @basedjensen This is not true, Gemini 3.5 pro has been testing for 2 - 3 months now with different checkpoints. Not sure where you are getting your information from? Even gpt 5.6 according to reports eg from Leo had been under tests for 1 - 2 months
English
0
0
6
112
Hensen Juang
Hensen Juang@basedjensen·
This is pure cope trying to be desguised as reasonable take. The problem with all takes like this assume Chinese model drop the moment they Finnish training instead of months of internal testing like murican labs and assume they have stopped training and everything else for release.
Ryan Greenblatt@RyanGreenblatt

Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?). Maybe this implies Kimi is like 8 or so months behind Anthropic in overall model strength/goodness (including usability) and like 6 or so months behind on overall capability (somewhat below Mythos Preview). This gap is presumably reduced by distillation (and more generally using OpenAI/Anthropic models) and algorithm leakage/diffusion, so I think that hypothetically if the US completely stopped and recent algos didn't diffuse, it would maybe take Kimi like 10 months to fully catch up to the best internal (including in development) Anthropic model. (I think this notion might be a better measure of where Anthropic/OpenAI are relative to Kimi, even though this hypothetical won't happen.) And if the US completely stopped, it might take Kimi around 27 months to reach the level the US would otherwise have reached one year from now (as in, with a year of further progress). My views here are pretty sensitive to how much benchmark performance is representative to overall usability. I think I now expect an open-weight AI which is straightforwardly "Mythos-level at cyber" (including usability etc.) in like 5 months supposing Kimi and others don't change their open-weight model policy. (I don't have a strong view about how big of a deal this is for cyber, but it may cause significant political consequences. This could be a significant overestimate of the time required.) I wonder what's driving Kimi being closer than I would have expected. Options include: - Experiment compute is significantly less important than labor (and labor at Kimi is competitive, which seems super plausible) - Implies more of a speedup from AI automating AI R&D and a bigger software-only intelligence explosion. - Or possibly Kimi is just doing much better than US companies and this is overcoming experiment compute disadvantages. - Algorithms are diffusing a lot / quickly (from e.g. OpenAI to Kimi). - Perf is overstated / benchmaxxed a lot. - Distillation / using OpenAI or Anthropic frontier AIs in AI development is very helpful for catching up. (But I'd guess Kimi K3 is a competitive pretrain which distillation doesn't help with?) - US companies aren't going as fast as they could for whatever reason.

English
7
7
273
16.7K
Maxime
Maxime@KingMaxime04·
@Tril1boswagginz @basedjensen @RyanGreenblatt And this been a new pretrained base means they should have been working on it for a while. Pretty sure this is not the first post trained model we are getting
English
1
0
0
149
bd5m112
bd5m112@bd5m112·
@renatomoraesp @RyanGreenblatt @tszzl Because chinese labs were not at frontier level so far, they all released models months behind top world models. Their operating workflow is based on distilling GPT and Claude so until they change this workflow, they can't be a truly frontier lab/model.
English
1
0
2
259
Ryan Greenblatt
Ryan Greenblatt@RyanGreenblatt·
Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?). Maybe this implies Kimi is like 8 or so months behind Anthropic in overall model strength/goodness (including usability) and like 6 or so months behind on overall capability (somewhat below Mythos Preview). This gap is presumably reduced by distillation (and more generally using OpenAI/Anthropic models) and algorithm leakage/diffusion, so I think that hypothetically if the US completely stopped and recent algos didn't diffuse, it would maybe take Kimi like 10 months to fully catch up to the best internal (including in development) Anthropic model. (I think this notion might be a better measure of where Anthropic/OpenAI are relative to Kimi, even though this hypothetical won't happen.) And if the US completely stopped, it might take Kimi around 27 months to reach the level the US would otherwise have reached one year from now (as in, with a year of further progress). My views here are pretty sensitive to how much benchmark performance is representative to overall usability. I think I now expect an open-weight AI which is straightforwardly "Mythos-level at cyber" (including usability etc.) in like 5 months supposing Kimi and others don't change their open-weight model policy. (I don't have a strong view about how big of a deal this is for cyber, but it may cause significant political consequences. This could be a significant overestimate of the time required.) I wonder what's driving Kimi being closer than I would have expected. Options include: - Experiment compute is significantly less important than labor (and labor at Kimi is competitive, which seems super plausible) - Implies more of a speedup from AI automating AI R&D and a bigger software-only intelligence explosion. - Or possibly Kimi is just doing much better than US companies and this is overcoming experiment compute disadvantages. - Algorithms are diffusing a lot / quickly (from e.g. OpenAI to Kimi). - Perf is overstated / benchmaxxed a lot. - Distillation / using OpenAI or Anthropic frontier AIs in AI development is very helpful for catching up. (But I'd guess Kimi K3 is a competitive pretrain which distillation doesn't help with?) - US companies aren't going as fast as they could for whatever reason.
English
42
45
594
233.2K