swm@hachyderm.io banner
swm@hachyderm.io

@stevenmerrill

Musician, Transformer-pilled software engineer, cyclist, weightlifter, and fermentation enthusiast. @[email protected]

Boston, MA Katılım Ekim 2007
325 Takip Edilen894 Takipçiler
Google Gemma
Google Gemma@googlegemma·
We’re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of what’s being fixed and updated in this release: 🧵👇
Google Gemma tweet mediaGoogle Gemma tweet media
English
128
321
3.3K
678.7K
swm@hachyderm.io
[email protected]@stevenmerrill·
I've been back in Codex recently: Sol is a talented planner and Luna at high reasoning is a no-nonsense workhorse that can get so much done while sipping usage. I had Luna helping me with SwiftUI work that Sol planned for a new DynamoDB client I've been building last night.
English
2
0
0
92
swm@hachyderm.io
[email protected]@stevenmerrill·
@TheZachMueller @liquidai I hear you; I am using my Strix Halo tablet as my local AI machine of choice (with a Q5 Qwen3.5 122B quant as about the smartest model I can run at okay speeds with MTP, and Gemma 4 or Qwen 3.6 4-8bit quants as small-but-outstanding-for-their-size workhorses.)
English
0
0
1
52
Zach Mueller
Zach Mueller@TheZachMueller·
@stevenmerrill @liquidai Sure, absolutely. Not what I'm targeting here (I want ~the whole system to exist on the spark). Optimize that and then migrate to the 6000 Blackwell and absolutely it'll just mog it in speed :) {Not saying that's bad, I have a rule to stick to for now}
English
2
0
0
71
Zach Mueller
Zach Mueller@TheZachMueller·
Okay, @liquidai LFM 8B@1B stays. Well done guys. TL;DR custom MCP wrapper + use existing CC MCP auth Currently LFM is ~2x slower than Haiku, but this is unoptimized and just an OOTB vLLM deployment. It's 2x slower purely from `decode` being 40 tok/s One final fable /goal
Zach Mueller tweet media
Zach Mueller@TheZachMueller

8B is promising, even if the claude config wasn’t set to only load in the singular MCP we needed and thus the context exploded immediately. Rerunning with only the specific MCP loaded that’s needed.

English
2
0
17
2.8K
swm@hachyderm.io
[email protected]@stevenmerrill·
@reddit_lies @dwallacewells The edge cache still has it cached, but as soon as you hit origin with a cookie that busts caching, it's clear that it's been deleted from the origin.
English
0
0
2
249
David Wallace-Wells
David Wallace-Wells@dwallacewells·
The Department of Energy appears to have deleted its thermostat-level recommendations to avoid people noticing that their generic recommendation was to set indoor temperatures at the same level which Zohran has asked New Yorkers to maintain during a grid-stressing heatwave.
gwen howerton@kissphoria

oh my god

English
126
2K
12.7K
532K
swm@hachyderm.io
[email protected]@stevenmerrill·
@shiba14857 @rasbt Based on the chart, it looks like he's talking about the current crop of 26-35B total param MoEs, including Qwen 3.6 35B-A3B and Nemotron 3 Nano 30B-A3B (and not pictured, probably because it's less strong at coding tasks: Gemma 4 26B-A4B.)
English
1
0
1
358
Manash
Manash@shiba14857·
@rasbt Are you planning any article on glm 5.2? And here 30B is active parameters of MoE?
English
3
0
1
1.2K
Sebastian Raschka
Sebastian Raschka@rasbt·
Have been taking different local open-weight LLMs for a test drive in different harnesses (Qwen-Code, Codex, Claude Code). 30B Mixture-of-Expert models are kind of a nice sweet spot and can solve challenging problems. And they get roughly 40 tok/sec on a Mac or DGX Spark, which is similar to GPT 5.5 in a Pro subscription and totally useable for everyday work. More interesting is also the harness choice! Claude Code seems to be using 2x many tokens as Codex. Gemma 4 E2B is here just for reference to show that the tasks can't be trivially solved by smaller models. Just finishing a longer write-up about this and will share soon (likely tomorrow)!
Sebastian Raschka tweet media
English
117
105
896
62.7K
Brad Dunlap
Brad Dunlap@braddunlap·
@MKBHD Damn that Mac Studio bro. There goes my dream of affording a 512GB M5 Mac Studio when they release. They'll be like $15K or some garbage.
English
3
0
4
8.8K
Marques Brownlee
Marques Brownlee@MKBHD·
Memory price hikes have finally started hitting Apple. Today, they raised some base prices: Macbook Neo: $599 -> $699 Macbook Air: $1099 -> $1299 Macbook Pro: $1699 -> $1999 Mac Studio: $1999 -> $2499 iPad Air $599 - $749 iPad Pro: $999 -> $1199
English
1.2K
2.6K
44.8K
7.3M
swm@hachyderm.io
[email protected]@stevenmerrill·
@_Sheik_ (Before Gemma 4 QAT weights came out I was using RedHat's NVFP4 quants.)
English
0
0
0
21
swm@hachyderm.io
[email protected]@stevenmerrill·
@_Sheik_ @0xSero You could try things like Gemma 4 QAT w4a16 or maybe NVFP4 (although might only be useful on 50-series NVIDIA.)
English
1
0
2
140
0xSero
0xSero@0xSero·
Why is llama.cpp so much slower than vLLM at long context multi-turn tasks, and at high concurrency.
English
41
5
289
44.5K
Dean W. Ball
Dean W. Ball@deanwball·
I am pleased and honored to announce that, on July 6, I'll be joining @OpenAI as leader of a new team called Strategic Futures. Our mandate will be to help the company's leadership shape frontier AI policy. There is a ton of work to do, and I'm excited to get started.
Dean W. Ball tweet media
English
383
185
3.1K
604.6K
Noam Shazeer
Noam Shazeer@NoamShazeer·
I’m excited to share that I’ll be joining OpenAI and look forward to working with the exceptional team there. It was a difficult decision to move on. I’m incredibly proud of the amazing team at Google and everything we’ve built together. It has been an honor and a pleasure to work with all of you.
English
988
864
16.3K
9.5M
swm@hachyderm.io
[email protected]@stevenmerrill·
@nnennahacks Congrats! Would love to get some stickers when you have a date for that meetup in Boston!
English
1
0
1
47
Nnenna 👩🏽‍💻✨
Nnenna 👩🏽‍💻✨@nnennahacks·
Update from me: I’m an @OpenAIDevs Codex Ambassador now. When I attended the gpt 5.5 party at OpenAI in SF, I met @sama, @romainhuet and a passionate community of AI builders. That experience set the tone for me to wanna bring Codex to my city. Thanks @reach_vb, @gabrielchua + @sriramkiron for this opportunity to bring Codex to Boston, one of the most ambitious, AI-forward cities in the US. 🎉
Nnenna 👩🏽‍💻✨ tweet mediaNnenna 👩🏽‍💻✨ tweet media
English
54
11
274
16.6K
Julia Turc
Julia Turc@juliarturc·
Rooting for open-source animal obesity today
English
6
2
66
4.8K
swm@hachyderm.io
[email protected]@stevenmerrill·
@chrisfree @MarkVillacampa Same. I had Fable do the first draft of a macOS Swift app Wednesday night and write me an explainer. (Thankfully I have the working code and the explainer to continue.)
English
0
0
2
27
Dimitris Papailiopoulos
Dimitris Papailiopoulos@DimitrisPapail·
Even if you tone it down it still routes you to Opus 4.8
Dimitris Papailiopoulos tweet media
English
2
0
3
781
[email protected] retweetledi
Philipp Schmid
Philipp Schmid@_philschmid·
More Gemma 4! New QAT Gemma 4 checkpoints with similar performance while using ~4x less memory! It comes with a new mobile quantization format that reduces memory footprint of Gemma 4 E2B to just 1GB. Quantization-Aware Training (QAT) simulates low-precision operations during training to allow loss-less quantization afterwards for smaller, faster models while maintaining accuracy. Available on @huggingface and directly runnable.
Philipp Schmid tweet media
English
19
54
591
32.1K