Artsy

29 posts

Artsy banner
Artsy

Artsy

@AutoArtsy

PhD in AI - In love with analyzing and generating art.

Katılım Ocak 2025
44 Takip Edilen4 Takipçiler
Artsy
Artsy@AutoArtsy·
@xai set the expectations so high for Grok 3. Will be a big disappointment if it doesn’t meet expectations.
English
0
0
0
8
Artsy retweetledi
Pliny the Liberator 🐉󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭
I don’t want to provide my world-class expertise just for you to hoard crowd-sourced prompts and construct elaborate security theater performances to appease investors who are foolish enough to believe guardrails=safety. I’m allergic to money, so don’t bother. My incentives are aligned with what’s best for the community and the future of AI. Are yours?
English
281
245
4K
373.3K
Artsy
Artsy@AutoArtsy·
Facts as AI Videos ? (Fully AI Generated) - Archaeologists have found edible honey in ancient Egyptian tombs thousands of years old. - The record for the longest flight by a chicken is 13 seconds, highlighting nature’s unexpected quirks. Should we do more of those at @adhdl4b
English
1
0
0
282
Artsy
Artsy@AutoArtsy·
@iruletheworldmo I think they’re too focused on safety at this point ignoring performance.
English
0
0
0
72
🍓🍓🍓
🍓🍓🍓@iruletheworldmo·
in 2 months open ai have shipped 20 products. anthropic have shipped 20 blog posts and dario interviews. the gap is opening. accelerate anon.
English
55
30
980
57.2K
Artsy
Artsy@AutoArtsy·
Read this paper today and it has some incredible insights on how reasoning models behave. Might summarize it in a thread if people are interested.
ハカセ アイ(Ai-Hakase)🐾最新トレンドAIのためのX 🐾@ai_hakase_

【🤔 大規模言語モデルの思考不足とは?Tencent AI Labの最新論文を解説!】 ✎. FYIG: x.com/_akhaliq/statu… Tencent AI Lab、Soochow University、Shanghai Jiao Tong Universityの研究者らが発表した最新論文「Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs」は、大規模言語モデル(LLM)の興味深い問題を指摘しています!😮 💡 この論文では、o1のようなLLM、例えばQwQ-32B-PreviewやDeepSeek-R1-671Bにおいて、「思考不足」という問題が起きているとのことです。具体的には、不正解の回答が正解よりも頻繁に推論を切り替えてしまい、結果として精度が向上せずに応答が長くなってしまう現象を指しています。 📊 添付の画像をご覧ください!AIME2024テストセットを用いた実験結果が示されています。 (a) Qwen-Math-72B、Llama3.3-70B、QwQ-32B-Preview、DeepSeek-R1-671Bの4モデルにおける生成トークン数の比較。 (b) QwQ-32B-PreviewとDeepSeek-R1-671Bの2モデルにおける思考数の比較。 緑色のバーが正解、赤色のバーが不正解です。o1のようなモデルでは、不正解の回答が正解よりも頻繁に推論を切り替えていることがわかります! 🔍 一方、Qwen-Math-72BやLlama3.3-70Bのような従来のLLMでは、不正解と正解の応答長に有意な差は見られませんでした。 ✨ この研究は、LLMの思考プロセスに関する新しい洞察を提供し、今後の発展に役立つ可能性があります。 論文の詳細は、こちらをご覧ください! huggingface.co/papers/2501.18…

English
0
1
1
161
Artsy
Artsy@AutoArtsy·
@abacaj Sonnet is great for front end I agree. But fails on complex coding tasks in my case. R1 and o1 have been nailing those more complex tasks.
English
0
0
0
37
anton
anton@abacaj·
Somehow OAI keeps releasing thinking models but they aren’t better than sonnet at code, they think too long and score well on benchmarks but have no taste and can’t produce a good frontend
English
106
44
1.4K
134.7K
Artsy
Artsy@AutoArtsy·
@OfficialLoganK In my experience sadly, it failed in tasks R1 and O1 did seamlessly. Talking about complex coding tasks. I understand that it’s fast and all, works well with Cline and code editing. But fails in complex coding tasks.
English
0
0
0
30
Logan Kilpatrick
Logan Kilpatrick@OfficialLoganK·
The progress with our Gemini reasoning models is actually wild, we are in the GPT-2 era of scaling reasoning! The main delta is that the models are actually super useful in their current form and getting better week over week. The future is exciting...
Logan Kilpatrick tweet media
English
162
121
1.9K
162.1K
Artsy
Artsy@AutoArtsy·
@dylhunn I know we’re speaking of exceptional speed in terms of generating responses - will there be a comparison with latest reasoning models as well ? How does it compare to DeepSeek R1 (pure performance not speed).
English
0
0
0
262
Dylan Hunn
Dylan Hunn@dylhunn·
o3-mini is just so good. exceptional intelligence and blazing speed. I haven’t been this excited for a model since GPT-4
English
188
91
2K
439.1K
Artsy
Artsy@AutoArtsy·
@ViralMindAI What do you guys think of the new Qwen model ? It has agentic capabilities definitely worth trying to deploy and improve with your data 📈
English
0
0
1
47
Artsy
Artsy@AutoArtsy·
@openrouter are we getting the new Qwen 2.5 1M variants soon ?
English
0
0
0
9
Artsy
Artsy@AutoArtsy·
The next big question is will @Google Gemini 2.0 Pro or @X Grok 3 outperform DeepSeek v3 or R1 ? Time will tell. Let’s not forget that @Alibaba_Qwen also put out a great model last night which isn’t getting as much attention yet.
Chubby♨️@kimmonismus

So @Google you wanna officially release Gemini 2.0 Pro or not? Sincerely, everyone

English
0
0
0
338
Artsy
Artsy@AutoArtsy·
The music 🎶
GIF
Qwen@Alibaba_Qwen

🎉 恭喜发财🧧🐍 As we welcome the Chinese New Year, we're thrilled to announce the launch of Qwen2.5-VL , our latest flagship vision-language model! 🚀 💗 Qwen Chat: chat.qwenlm.ai 📖 Blog: qwenlm.github.io/blog/qwen2.5-v… 🤗 Hugging Face: huggingface.co/collections/Qw… 🤖 ModelScope: modelscope.cn/collections/Qw… 🌟 Key Highlights: * Visual Understanding : From flowers to complex charts, Qwen2.5-VL sees it all! * Agentic Capabilities : It’s a visual agent that can reason and interact with tools like computers & phones. * Long Video Comprehension : Captures events in videos over 1 hour long! ⏳🎥 * Precise Localization : Generates bounding boxes & JSON outputs for accurate object detection. * Structured Data Outputs : Perfect for finance & commerce, handling invoices, forms & more! 💼📊 Try Qwen2.5-VL now at Qwen Chat or explore models on Hugging Face & ModelScope . 🌐

English
0
0
0
57
Artsy
Artsy@AutoArtsy·
@minchoi @tajb03 When it comes to generation yes it’s not that great. It does okay. Definitely not better than DALLE 3 or Flux it comes to visual preference ( rather than metrics ).
English
0
0
0
17
Min Choi
Min Choi@minchoi·
Wow. DeepSeek just dropped Janus-Pro-7B, an open-source multimodal AI that beats DALL-E 3 and Stable Diffusion. The 🐋 is on fire. 👀
Min Choi tweet media
English
56
138
898
126.1K
Artsy
Artsy@AutoArtsy·
Watching non-AI experts confidently discuss @deepseek_ai without understanding the basics is both hilarious and frustrating. The internet never disappoints. Or it always does. 💀
English
0
0
0
35