
Hawkings
103 posts





Kimi K3 vs Claude Fable 5 vs Grok 4.5 vs Gemini 3.6 Flash 3.6 flash did a decent job here and better than 3.5 flash > Bugatti W16 engine with mechanically accurate internals and photorealistic exterior in Three.js





Ran 3 flagship Chinese LLMs: Qwen3.8-Max-Preview, Qwen3.7-Max, DeepSeek-V4-Pro: same task battery (long-context, multi-tool, research-sim), same scoring rubric. Capability: nearly tied (92.9-96.4% strict pass rate). Cost per call: ~5x apart. Note: Qwen3.8-Max-Preview on promo pricing; likely to rise once that ends. Same job, up to 5x price difference by model choice.

Qwen3.8-Max-Preview distilling Qwen3.8-Max-Preview: hour 13 update 32,987/50,000 traces produced ZERO refusals 😛

Ran 3 flagship Chinese LLMs: Qwen3.8-Max-Preview, Qwen3.7-Max, DeepSeek-V4-Pro: same task battery (long-context, multi-tool, research-sim), same scoring rubric. Capability: nearly tied (92.9-96.4% strict pass rate). Cost per call: ~5x apart. Note: Qwen3.8-Max-Preview on promo pricing; likely to rise once that ends. Same job, up to 5x price difference by model choice.























