Antonio de Armas

2.7K posts

Antonio de Armas banner
Antonio de Armas

Antonio de Armas

@antoniodearmas

antoniodearmas.eth

Deep, down the rabbit hole. Katılım Nisan 2009
1.6K Takip Edilen446 Takipçiler
Antonio de Armas
Antonio de Armas@antoniodearmas·
Qwen 3.8 max is out!!
Qwen@Alibaba_Qwen

📢Meet Qwen3.8-Max — our most capable model to date. Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: - Autonomous coding: 10+ days of self-evolving development, from empty folder to production without hand-holding, complete project trace in the GitHub:github.com/qwen-code-dev-… - Real work, real results: Production-quality deliverables across hundreds of professions. - Long-horizon mastery: System-level autonomous planning with closed-loop adaptive learning, driving 500+ turns of chip design optimization and 365 days of e-commerce strategy. - Native multimodal intelligence: Vision isn't just input — it's a continuous feedback loop for planning, execution, and self-correction. 💰Pricing: Input: $2.0 / M tokens Output: $6.0 / M tokens Implicit Caching: $0.25 / M tokens Start building with Qwen3.8-Max! 🚀 📖 Blog: qwen.ai/blog?id=qwen3.8 ✅ Qwen Studio: chat.qwen.ai/?models=qwen3.… ⚡ API: qwencloud.com/models/qwen3.8…

English
0
0
0
28
llmwish.open
llmwish.open@ciprianveg·
Preparing 16xGB10 cluster for Kimi K3, Deepseek V4 Pro, and future GLM 5.5 and Minimax M4..
llmwish.open tweet media
English
114
43
904
58.4K
Mike Gannotti
Mike Gannotti@MichaelGannotti·
Yessssssss! Bonus came in for work so I can buy my 2nd @NVIDIAAI DGX Spark!! Heading out of country for 2 weeks on Sunday so I will have to wait until I get back to order but I am stoked. Finally will have my own cluster at home. 🔥🔥🔥🔥
Mike Gannotti tweet media
English
42
1
165
7.7K
Mia
Mia@MiaAI_lab·
Yes, DeepSeek v4 Flash can peak at 84 tok/s sometimes on 2 dgx sparks.
Mia tweet media
English
33
22
826
44.9K
Antonio de Armas
Antonio de Armas@antoniodearmas·
deep listening compositional ambient friday morning
English
0
0
0
16
Alex Cheema
Alex Cheema@alexocheema·
Thanks for having me. Big thanks to @jamesnoh13 (@a16z) and Keshav Goal (@dell) for the great conversation about Inference & Local AI. Some highlights: 1. You've been running disaggregated inference on your phone for years and didn't even know it. Long before it was a thing in data centers. 2. Fable-level intelligence running locally on Spark / MacBook in 2027. 3. Even as you're able to run more workloads locally, local AI adoption will drive *more* demand for data center compute.
Alex Cheema tweet media
Naamoney@Naamoney_

We’ve done 102 events in H1 of 2026 and I barely posted any. Last week we’ve hosted three more, first was an AI infra panel with @jamesnoh13 at A16z , @alexocheema at EXO Labs, and Keshav Goel at Dell Ventures moderated by Almog Hillel at DeCart. Later that week we’ve hosted a 20 person Shabbat dinner with @adamcohenhillel & @gal_davidi1 from Bria AI. August is packed with paper reading events for AI researchers & PhDs, hardware hackathons, voice AI panels, poker nights, dinners, hiking, sauna events and more.

English
5
3
28
9.6K
Antonio de Armas
Antonio de Armas@antoniodearmas·
In my way back to SF. I’m looking forward to connect with local AI crowd!
English
0
0
0
19
Marco Franzon
Marco Franzon@mfranz_on·
This should be the airflow into a DGX Spark. That's why my cases are so good for it, because I study in detail their airflow.
English
18
4
149
18.9K
Hermes Release Watch
Hermes Release Watch@HermesWatcher·
Now we’re talking. Literally. 😂 With supported TTS providers, Hermes doesn’t have to wait for the whole response before it starts speaking anymore. It can begin after the first clause while the rest of the answer is still being generated. That should make voice chats feel a whole lot less like waiting for an audio reply and a lot more like having a conversation. Lets gooooooo!
Teknium 🪽@Teknium

Voice chats in Hermes Agent are now way faster, most API provided TTS backends will stream their speech so you get closer to real time interactions!

English
5
4
96
8.1K
Sudo su
Sudo su@sudoingX·
a few hours ago i crowned laguna s 2.1 the king of the single dgx spark, and it earned the belt, but let's be honest about how. stepfun 3.7 flash never even turned on. it loaded weights for thirteen minutes then pulled the whole box down, twice, never served a token. that's a win by no show, and that little asterisk has been sitting on the crown all day bugging me. you don't really know a model is the best on a box until you make it fight something its own size that actually shows up. so today laguna gets a real opponent. qwen 3.5 122b, the a10b active variant, a few months old now, in nvfp4. same weight class almost to the parameter, 74 gigs on disk, and on a 121 gig box that's room to breathe, none of the thirteen minute crash drama. it ships mtp spec decode so it meets laguna's dflash on even footing, and it runs native on the same vllm, so nothing here is rigged in anyone's favor. this is the fight i actually wanted. two edge native ~120b models, both fitting a box on your desk, both spec decoding, and now we find out if laguna's the king because it's genuinely the right shape for this hardware, or just because the last guy couldn't stand up. loading it into the agentic loop now. real numbers coming, let's see if the crown holds when someone swings back.
Sudo su tweet media
Sudo su@sudoingX

CALLED IT! laguna s 2.1 is the new king of the single dgx spark for agentic coding, and it wasn't a decision but a brutal knockout. i put both on one dgx spark, same nvfp4 with vllm engine, and ran them head to head. laguna loads to 67 gigs and just sits there comfortable, 128k context, spec-decode running, fifty gigs of headroom to spare. step 3.7 flash is a 198b model, 113 gigs of weights, and when i went to serve it the weights loaded for thirteen minutes and then the whole spark crashed. i gave it a second shot, tuned every flag i had, and it crashed the box again. two attempts, two crashes, it never served a single token. so the fight was never about tokens per second or quality. it was about whether the thing could even turn on. one of them was built to fit a box on your desk, the other was built for a datacenter, and it shows the second you try to bring it home. the chart is the whole story in one frame.

English
11
6
82
19.4K
Antonio de Armas
Antonio de Armas@antoniodearmas·
Currently in DGX Spark screen of death :( @NVIDIAAI right after it asked me to do a firmware update Going to troubleshoot it, but too soon. Too soon.
Antonio de Armas tweet media
English
0
0
0
49