Sean Uddin

926 posts

Sean Uddin banner
Sean Uddin

Sean Uddin

@Metafintek

Serial Founder | Builder | ENTP Author of "Architects: Building the Exit to Techno-feudalism" 0-1 Systems Designer Own The Means of Creation #OTMOC

Metaverse Katılım Ocak 2022
322 Takip Edilen823 Takipçiler
Sean Uddin retweetledi
Lotto
Lotto@LottoLabs·
Bros we’re going to have an abundance of benchmarks for you to run on localmaxxing These are specific for local models and reflect more realistic local model use
Lotto tweet media
English
4
3
26
2.5K
Sridhar A
Sridhar A@sridharfyi·
where do you think is the best place to scout exceptional founders right now? 🔹 san francisco 🔹 new york city 🔹 toronto 🔹 london vote below 👇 if your city isn't listed (berlin, bangalore, singapore, paris, etc.), reply with it and tell me why.
English
29
0
22
3.1K
Sean Uddin
Sean Uddin@Metafintek·
@craigweiss Opus 5 found a lot of errors in work that was orchestrated by fable 5. As a result cost me 3 weeks of work. Have to cancel. Not worth the money to beta test vague shit. Codex +glm5. 2 is fine. And we are actually building a fully local ai os anyway no need for any subs soon
English
1
0
3
19.4K
Craig Weiss
Craig Weiss@craigweiss·
wow, claude’s reign has ended
English
210
31
1.9K
1M
Sean Uddin retweetledi
Michelle Maxwell ™
Michelle Maxwell ™@MichelleMaxwell·
This is so me and my hubby, Anyone else? 😆
English
666
1.4K
12.1K
800.8K
Sean Uddin retweetledi
Glenn Sonna
Glenn Sonna@GlennSonna·
Pushed Bonsai-1.7B from @PrismML further. 32 → 64 tok/s on the same OnePlus 13. 2×. Parallelized attention and hoisted the conversion out of the KV walk. Bit-for-bit identical output. Pure CPU. I’m sure we can go harder.
GIF
Glenn Sonna@GlennSonna

Bonsai-1.7B from @PrismML: a 1-bit model decoding at 32 tok/s on a OnePlus 13. Q1_0. 237 MB. Pure CPU. That’s a 1.7B model at chat speed on a phone. Who needs a GPU?

English
4
17
138
17.5K
Kiri
Kiri@Kyrannio·
I’m so very excited for this but I do wish they used a different font or at least did some sort of enhancement to the lettering
Kiri tweet media
English
21
4
91
2.9K
Sean Uddin
Sean Uddin@Metafintek·
@butshaunn @sabakarimm Here in lies the rub. Some people raise off the bag of a fag packet dog shit decks and no product. Some great products Are not chosen. VCs are not the arbiters of quality products.
English
0
0
1
6
Shaun Gold | Venture Comedy
On the contrary: "Can you walk me through it right now" not "send me the deck"-I want to see if you are a threat to my portfolio company in the same vertical. "Who's your lead?" means they're considering being it-We don't lead. "What would you do with more than you're asking?"-Do you have any idea how to deploy capital? "Let me intro you to someone on my team"-I don't want to deal with this. They can ignore you from now on.
English
1
0
2
250
Saba Karim
Saba Karim@sabakarimm·
Green flags a VC is actually going to invest: "Can you walk me through it right now" not "send me the deck" "Who's your lead?" means they're considering being it "What would you do with more than you're asking?" "Let me intro you to someone on my team" They reply in hours, not weeks They bring up the deal to you next time, unprompted Speed is the tell. Real interest moves fast. Everything slow is a soft no.
English
5
2
75
5.7K
Sean Uddin retweetledi
Xenova
Xenova@xenovacom·
A swarm of GPT 5.6 Sol agents spent over 40 hours optimizing a Kimi K3-like model from 65 to 406 tok/s. This animation follows their collaboration as they discover operator fusions, transform the execution graph, and develop new kernel algorithms.
English
50
84
952
73.1K
Sean Uddin
Sean Uddin@Metafintek·
@jefielding Well far be it from me to say we would be well suited to work together. But my project is everywear.id local AI with ambition to be on every wearable as the AI App store.
English
0
0
0
0
Jenny Fielding
Jenny Fielding@jefielding·
This week I’ve been asked to: -Sign an NDA to see a pitch deck -Pay a higher price on a Safe than if I had invested last week -Waive my pro rata to make room for a multi stage fund. Nah…
English
54
4
193
26.2K
Sean Uddin
Sean Uddin@Metafintek·
We need to see longer context tasks but modularised, we are designing such benchmarks for everywear.id practical use means breaking up larger tasks on smaller machines, into digestible chunks. We do this in the set up phase automatically but will have advanced settings for tinkerers. raw Toks is just a flex.
English
0
0
1
24
Lotto
Lotto@LottoLabs·
Building a bunch of local first terminal bench style benchmarks designed for local models first TB2.1 are great and they are proxies for how well local models will perform on a bunch of tasks but they’re designed for sota models These benchmarks for local models will be simplified in design and more relevant to smaller models and how well they perform at typical tasks I’ve seen users run as assistants and research bots, there will still be coding tasks but they are designed to reflect actual realistic use cases, including partial scoring
English
3
0
11
1.2K
Sean Uddin retweetledi
Sudo su
Sudo su@sudoingX·
turns out i wasn't the only one quietly running this. i posted asking if nemotron 3 nano omni was as good as it felt, and the answer came back fast, people using it and calling it a favorite, a model the local ai crowd clearly loves and almost nobody posts about out loud. and it just got better. nvidia ships it in nvfp4, so anyone on a dgx spark can run it native on the gb10 through vllm. the most underrated multimodal model that fits on a single box, now in the format that runs best on it. huggingface.co/nvidia/Nemotro…
Sudo su tweet media
Sudo su@sudoingX

anon, if you've got a dgx spark and you do any document, video, or image classification work, the most underrated model out there is nemotron 3 nano omni. i've run it, it's genuinely useful and almost nobody's talking about it. nvidia's open multimodal model, all in one: > images, video, audio, ocr, documents, text, code > 30b total, only 3b active (moe), stays fast > 33 gb footprint on the spark at q8, 256k context > 1300 tok/s prefill, chews through docs and images, 264 gen at short context, ~57 at depth underrated gem for real multimodal work on a dgx spark box. i keep reaching for it. huggingface.co/nvidia/Nemotro…

English
12
5
126
16.2K
Sudo su
Sudo su@sudoingX·
local ai model recommendations for everyone. every hardware tier, all benched on my own machines. 8 to 12gb, the gpu already in your laptop: > bonsai 27b, a 1bit crush of qwen 3.6 27b down to 3.9gb. runs the full agent loop at 128k context on 8gb. real agentic work on hardware you forgot you had. hf: https//huggingface.co/prism-ml/Bonsai-27B-gguf rtx 3090, 24gb, the used market king: > qwen 3.6 27b dense q4 for raw quality at ~40 tok/s hf: https//huggingface.co/unsloth/Qwen3.6-27B-GGUF dgx spark, 128gb unified, the desk supercomputer: > laguna s 2.1, poolside's 118b moe in nvfp4, 30 to 45 tok/s holding a 128k window. hf: https//huggingface.co/poolside/Laguna-S-2.1-NVFP4 every number here i measured myself. the full benches are on my profile, go dig.
English
35
13
237
24.2K
Sean Uddin
Sean Uddin@Metafintek·
Wow Opus 5 is looking pretty good v Fable 5 at half the price! Good Corpo serfs. :D Of course running local if you #OTMOC is better for you. I personally do not like Opus offering CLOUD services now as well.
Sean Uddin tweet media
English
0
0
0
30
Sean Uddin retweetledi
Kyle Hessling
Kyle Hessling@KyleHessling1·
Alright guys, this is a big one, and I will be launching it on HF in the next few days. Lots of trial and error to get to this point but I’m super happy with the result! TLDR: I finally got a merge to work beautifully that keeps the thinking capability of Qwopus v2 and merges it with the massive coding execution delta of Qwopus Coder! Effectively the biggest strengths of both models merged into one! After trying multiple runs of 3 different merge methods over the last three weeks, this one is finally the ticket. Here’s a few info graphics covering the whole journey:
Kyle Hessling tweet media
English
18
14
170
6.4K
Sean Uddin retweetledi
Lotto
Lotto@LottoLabs·
Localmaxxing.com supports this message
Jensen Huang@JensenHuang

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-Weigh…

English
1
1
14
689
Sean Uddin retweetledi
El CondeNado
El CondeNado@elconde_nad0·
Cuando le pido al señor me ayude...
Español
622
9.3K
52.2K
2.5M
Sean Uddin retweetledi
Brian Roemmele
Brian Roemmele@BrianRoemmele·
WOW! The $8 AI Machine! Something extraordinary just happened and it changes what “local AI” can mean. I am testing it tonight. Thus far it shows many possibilities… So what it this $8 AI device? A developer going by slvDev has forced a 28.9-million-parameter language model onto an ESP32-S3 microcontroller that costs roughly eight dollars. Not a Raspberry Pi. Not a Jetson. An eight-dollar microcontroller. The model runs completely offline, generates coherent short stories at about 9.5 tokens per second, and draws power measured in the same range as a small LED. This is more than a hundred times larger than the previous record for the same class of chip (the earlier 260,000-parameter TinyStories experiments). For perspective, the original ChatGPT sat at 117 million parameters. We are now running a model roughly a quarter of that size on silicon you can buy for the price of two coffees. How the Impossible Became Possible The ESP32-S3 has only 512 KB of fast SRAM, 8 MB of PSRAM, and 16 MB of flash. Conventional wisdom said a model of this size simply would not fit. The breakthrough is architectural, not brute force. Most of a language model’s parameters live in a giant embedding table a lookup table you mostly read from, not compute against. Drawing directly from Google’s Per-Layer Embeddings technique (the same family of ideas used in the Gemma models), the developer moved the bulk of that table roughly 25 million parameters into flash memory and memory-mapped it. The chip only needs to pull about six rows, roughly 450 bytes, for each new token. The remaining dense “thinking” core stays in the fast SRAM (around 560 K of active working memory). The model is stored at 4-bit quantization and occupies about 14.9 MB total. The result is a system that feels almost free to run. The heavy parameters sit quietly in flash and are sampled sparingly. The little core does the real work. It is elegant engineering of the purest kind. What I Am Doing With It Right Now I have the boards on the bench in the garage lab. The first units are already talking short, coherent stories appearing on a tiny wired display, generated entirely on the chip with no Wi-Fi, no API key, no cloud round-trip. Latency is local. Privacy is absolute. Power draw is low enough that battery operation becomes interesting. I am treating these as the first generation of true $8 AI machines. Early tests are focused on three practical directions. - Embedding the model into simple nodes. - Pairing it with local voice front-ends - Exploring whether multiple of these chips can be networked as a lightweight swarm. The model is deliberately limited. It was trained on the Microsoft TinyStories dataset and is excellent at coherent narrative, not at open-ended question answering or tool use. That is a feature, not a bug. It forces us to design systems around what the silicon can actually deliver instead of pretending every edge device needs a frontier model. Real Use Cases That Suddenly Become Practical Once you accept that a capable language model can live for eight dollars and run without the cloud, a new class of devices becomes possible: This is the opposite of the current trajectory that wants every intelligent act to travel through a remote server. It is the beginning of intelligence that is cheap enough, private enough, and local enough to become infrastructure rather than a service. We have spent years watching model sizes explode upward. The more interesting frontier may be the opposite direction: how small, how cheap, and how local can useful intelligence become? An eight-dollar chip that can tell coherent stories is not a toy. It is a proof that the lower bound keeps moving. The open repository is at github.com/slvDev/esp32-ai I will keep testing, measuring, and reporting what these little machines can and cannot do. The age of abundant local intelligence just got a little more real, and it arrived wearing an eight-dollar price tag.
Brian Roemmele tweet media
English
192
480
3.1K
398.2K
Sudo su
Sudo su@sudoingX·
what harness are you actually running your local models through?
English
65
2
34
9.7K