Shorm

196 posts

Shorm

Shorm

@Shmorglof

Katılım Eylül 2022
4 Takip Edilen1 Takipçiler
Steven Cheng
Steven Cheng@stevencheng·
I love this opensource project! Someone just put a 28.9M parameter AI model on an ESP32-S3. A chip that costs only a few dollars. No cloud. No API. No internet. It runs locally at ~9.5 tokens/s and can generate stories on a tiny screen. This changes the way we think about AI hardware. The next generation of AI toys, robots, and IoT devices may not need the cloud at all. Small chips are getting smarter. 🤖 github.com/slvDev/esp32-ai
English
27
85
726
59.1K
Giovanni
Giovanni@GinoCirino·
@petergyang Seems like biology is much harder then mathematics. Which is obvious but kind of interesting to see models solving mathematics and yet we haven't truly solved something that changes the death rate
English
4
0
1
648
Peter Yang
Peter Yang@petergyang·
The number one thing I want AI to fix is to cure cancer once and for all
English
277
111
1.7K
76.1K
Shorm
Shorm@Shmorglof·
@0xSero And as for the cost?
English
0
0
0
348
0xSero
0xSero@0xSero·
Here’s what you need for deepseek-v4-flash The size of a tiny box, silent, cool, doesn’t guzzle power. 1 million context with room to spare for a vision model
0xSero tweet media
English
75
49
1.5K
108.2K
Shorm
Shorm@Shmorglof·
@__tinygrad__ Maybe those will become relevant when you can run them on a real home computer
English
0
0
0
307
the tiny corp
the tiny corp@__tinygrad__·
I feel like it's the closed US LLM labs in one little defensive corner vs the rest of the world on GLM, Kimi, and DeepSeek. The only path forward is open, this revolution is too big for any one company.
English
30
44
856
26.4K
Shorm
Shorm@Shmorglof·
@EricLengyel He forgot to tell it to make it high quality, classic mistake
English
0
0
1
314
Shorm
Shorm@Shmorglof·
@forgebitz It's just another tool, it's essentially just a very high level programming language that can still theoretically have good performance. A fool can only take a good tool so far
English
0
0
1
492
Klaas
Klaas@forgebitz·
the whole "software is solved" crowd always seems to ship either nothing or the most basic stuff like, yeah, to-do lists, social media scheduling apps etc. were solved long before ai
English
130
76
1.5K
69.1K
Shorm
Shorm@Shmorglof·
@DCinvestor Never cause someone has to actually act on that knowledge, and its incredibly cumbersome to make any advancements in todays day and age
English
0
0
0
397
DCinvestor
DCinvestor@DCinvestor·
don’t get me wrong, it’s really cool AI can prove/disprove obscure math proofs but when we are gonna achieve some breakthrough, novel scientific insight which makes us materially better off?
English
106
9
270
33.3K
Shorm
Shorm@Shmorglof·
@orgunbertan What makes them toxic I refuse to use it
English
0
0
1
210
Bertan Orgun
Bertan Orgun@orgunbertan·
reddit game dev subs feel pretty toxic these days. X is a much better space for indie devs to share work.
English
67
12
361
25.5K
Superman
Superman@thesupermanmx·
You can now run a 2.8 TRILLION parameter model on a 4GB GPU. Someone open-sourced a tool that uses "Layer-wise Inference." It only loads one layer onto your GPU at a time. so the VRAM you need depends on the layer size, not the model size. no quantization. no distillation. no pruning. → DeepSeek-V3 (671B) on 12GB → Llama 3.1 405B on 8GB → Kimi K3 (2.8 TRILLION params) on under 4GB → works with almost every open model the biggest model on it needs the LEAST VRAM. K3 is sparse MoE, so it streams only the experts a token actually routes to instead of a whole dense layer. 2.8 trillion parameters running in less VRAM than a 70B.
Superman tweet media
English
165
371
3.2K
257.3K
Saloni
Saloni@saloniiio·
@neerajjj6785 because privacy, control, and offline inference still matters
English
5
0
32
1.7K
Neeraj
Neeraj@neerajjj6785·
why are we even discussing local models when we need a $17K of hardware to run them properly?
English
209
11
314
27.5K
Hakimi Eiqbal
Hakimi Eiqbal@hakimieiqbal·
@trikcode deepseek pricing is already there sol speed is the part that would need proving
English
1
0
5
783
Wise
Wise@trikcode·
I don’t need GPT-6. I need GPT-5.6 Sol at DeepSeek pricing.
English
177
59
2.3K
98.9K
Faisal Khatri
Faisal Khatri@mfaisal_khatri·
@trikcode Why don't you try local llm models? Is there any challenge in using them?
English
6
0
3
3.1K
Shorm
Shorm@Shmorglof·
@solopribuilds Make the API start returning zip bombs and ransomware
English
0
0
1
20
Pri
Pri@solopribuilds·
Interviewer: Your API is getting 17,000 requests per second. Server is choking. Users are complaining. You have 15 minutes to fix it. what do you do?
English
149
0
76
16.6K
Hugging Models
Hugging Models@HuggingModels·
Meet Kimi-K3-0.40B, a compact yet powerful feature extraction model that's taking the ML community by storm. With 17K downloads, it's built on the Kimi-K3 base and fine-tuned for precision. Perfect for turning raw text into rich embeddings. Let's dive in! #AI #NLP
Hugging Models tweet media
English
17
59
834
134.4K
Shorm
Shorm@Shmorglof·
@jun_song Yeah lemme just stop by bestbuy and get a trillion gb of ram
English
0
0
1
80
Shorm
Shorm@Shmorglof·
@BrianRoemmele Is 0.3 tokens per second the best it can get or is there some technological improvement there because if thats it, thats pathetic. Might as say you can run it off a phone because you can technically run it on a CPU and use hard drive as virtual memory
English
0
0
0
63
Brian Roemmele
Brian Roemmele@BrianRoemmele·
BREAKTHROUGH: A full, unmodified 2.78-trillion-parameter Kimi K3 on a consumer laptop by streaming only the activated experts from NVMe. YOU CAN’T RUN KIMI K3 “ON THAT” THEY DECLARED. There are many paths to do it. This is one: Marco Bambini Just Gave Us the Full Kimi K3 on a Laptop Meet Marco Bambini he did something that felt impossible only a day ago. He built WASTE Weight-Aware Streaming Tensor Engine a clean, dependency-free C inference engine that runs the complete, unmodified 2.78-trillion-parameter Kimi K3 model by streaming only the activated experts straight from NVMe. No distillation. No pruning. No cloud. The full open-weight model. We have it running in the lab right now. What Marco Actually Built Kimi K3 is a sparse Mixture-of-Experts system. Only about 4 % of its weights fire on any given token. Marco’s insight was simple and ruthless: the idle experts do not need to live in RAM. They only need to be reachable in time. WASTE keeps the model’s “trunk” (attention, shared components, embeddings) resident in memory — roughly 27 GB on the converted container. The 82,000+ routed experts stay on disk as tightly packed residual vector-quantized records. When the router selects its 16 experts per layer, the engine issues direct, cache-bypassing reads from the internal NVMe and feeds them into a bounded expert cache. The rest of the machine’s RAM becomes working space for that cache. On a 64 GB MacBook Pro with the container on the internal SSD, we are measuring 0.32–0.34 tokens per second at a comfortable memory budget. Prefill sits a little higher. The vision tower works. Logits match the reference implementation to within a few parts in a million. It is the real model. The container itself is 982 GiB after conversion from the original 1.42 TB MXFP4 weights. Minimum RAM floor is just over 29 GB for short context. Push the budget higher and the expert cache hit rate climbs; push too high and you start paging and the speed collapses. The sweet spot on current consumer hardware is clear and measurable. How We Are Testing It We converted the official weights, verified the container, and began systematic runs the same day the engine stabilized. First we confirmed numerical fidelity against the PyTorch reference on short prompts. Then we moved to longer generation, vision inputs, and multi-turn chat using Kimi’s native XTML format. We are measuring wall-clock decode, expert I/O versus compute split, cache hit rates at different RAM budgets, and thermal behavior under sustained load. We are also exercising the OpenAI-compatible server that sits on top of the same C library so we can drop the model into existing agent loops without rewriting anything. Early observations: •Expert I/O dominates the timeline, as expected. On a fast internal NVMe the engine is already near the practical ceiling of the storage subsystem. •The architecture’s sparsity is the entire enabler. A dense model of this size would be dead on arrival for local use. •Context length is currently limited by RAM more than by the model itself. Practical working contexts sit comfortably in the tens of thousands of tokens on 64 GB hardware; the full million-token window will need more memory or smarter KV management. •Thinking tokens are expensive at this speed. Long internal monologues turn into multi-hour runs. For agent work we are already experimenting with tighter control over when full reasoning is requested. We are treating this as a research instrument, not a finished product. Every run teaches us something about expert locality, prefetch opportunities, and how far pure software streaming can push trillion-scale inference on ordinary machines. 1 of 2
Brian Roemmele tweet media
English
82
188
1.3K
144.9K
comma
comma@comma_ai·
08.12.2026 🤏
comma tweet media
158
127
4.5K
648.5K
Shorm
Shorm@Shmorglof·
@sfxnz Run it for free means run it for free if you have a $400,000 personal miniature datacenter already, Im guessing
English
3
0
6
1K
Sufyan
Sufyan@sfxnz·
Everyone who’s saying ‘5.6 luna still beats Deepseek v4 flash official’ doesn’t understand 1: the jump from preview to the official version and what that means for local ai. 2: the fact that it can be run for free.
English
15
0
263
13.2K
Mike Key
Mike Key@1337hero·
What it actually takes to run Kimi K3 at home. About $42k worth of hardware in a photo. Still impressive it came in under 1000 watts of power. And if you really think about it, you could put this whole setup in to a nice little mini rack. Photo credit @ServeTheHome
Mike Key tweet media
English
29
10
176
13.8K