Artsy
29 posts

Artsy
@AutoArtsy
PhD in AI - In love with analyzing and generating art.




'I Follow Rivers' with @winonaoak as part of @tiesto's new EP 'Prismatic: Pack One' is OUT NOW ⚡️🙏 #musicalfreedom #newmusic #tiesto #outnow

【🤔 大規模言語モデルの思考不足とは?Tencent AI Labの最新論文を解説!】 ✎. FYIG: x.com/_akhaliq/statu… Tencent AI Lab、Soochow University、Shanghai Jiao Tong Universityの研究者らが発表した最新論文「Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs」は、大規模言語モデル(LLM)の興味深い問題を指摘しています!😮 💡 この論文では、o1のようなLLM、例えばQwQ-32B-PreviewやDeepSeek-R1-671Bにおいて、「思考不足」という問題が起きているとのことです。具体的には、不正解の回答が正解よりも頻繁に推論を切り替えてしまい、結果として精度が向上せずに応答が長くなってしまう現象を指しています。 📊 添付の画像をご覧ください!AIME2024テストセットを用いた実験結果が示されています。 (a) Qwen-Math-72B、Llama3.3-70B、QwQ-32B-Preview、DeepSeek-R1-671Bの4モデルにおける生成トークン数の比較。 (b) QwQ-32B-PreviewとDeepSeek-R1-671Bの2モデルにおける思考数の比較。 緑色のバーが正解、赤色のバーが不正解です。o1のようなモデルでは、不正解の回答が正解よりも頻繁に推論を切り替えていることがわかります! 🔍 一方、Qwen-Math-72BやLlama3.3-70Bのような従来のLLMでは、不正解と正解の応答長に有意な差は見られませんでした。 ✨ この研究は、LLMの思考プロセスに関する新しい洞察を提供し、今後の発展に役立つ可能性があります。 論文の詳細は、こちらをご覧ください! huggingface.co/papers/2501.18…


o3 BEATS R1 OVERALL AND BLOWS EVERYONE ELSE AWAY IN CODING o3-mini high became the BEST LLM BY FAR when it comes to a combination of performance, speed, and price - beats o1, Sonnet, and others BY A LOT in coding - 2x cheaper than Sonnet and 15x cheaper than o1 - ~5x faster than R1 - 2nd best model right after o1 in all categories ChatLLM and CodeLLM now have o3-high if you want to play with it.



So @Google you wanna officially release Gemini 2.0 Pro or not? Sincerely, everyone

🎉 恭喜发财🧧🐍 As we welcome the Chinese New Year, we're thrilled to announce the launch of Qwen2.5-VL , our latest flagship vision-language model! 🚀 💗 Qwen Chat: chat.qwenlm.ai 📖 Blog: qwenlm.github.io/blog/qwen2.5-v… 🤗 Hugging Face: huggingface.co/collections/Qw… 🤖 ModelScope: modelscope.cn/collections/Qw… 🌟 Key Highlights: * Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all! * Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones. * Long Video Comprehension : Captures events in videos over 1 hour long! ⏳🎥 * Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection. * Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more! 💼📊 Try Qwen2.5-VL now at Qwen Chat or explore models on Hugging Face & ModelScope . 🌐











