John Moore

116 posts

John Moore banner
John Moore

John Moore

@talemi

Data and AI leader, ~15 years. Obsessed with one open problem in AI: your tools remember you, but not each other.

DFW Katılım Ocak 2026
59 Takip Edilen29 Takipçiler
John Moore
John Moore@talemi·
@JJEnglert Yeah for a good while now I make sure at least a few agents are working overnight. Nothing triggers me in the morning like seeing one had hit a roadblock like right after I went to sleep. Infuriating!
English
0
0
0
8
JJ Englert
JJ Englert@JJEnglert·
If you think you're AI Native, here's a good test: When you're heading to bed tonight, are you asking yourself: "What agents can I be running over night to get a head start on my day tomorrow" If that's an honest thought, you're among the top .01% of AI users right now.
English
135
28
476
39.6K
John Moore
John Moore@talemi·
@axiomofmind @geekdan @BrianRoemmele So you conflate two of the three but that just means that you choose speed and quality. It's clear that's your preference and that means you sacrifice cost to get it. Nothing wrong with that!
English
0
0
1
10
Brian Roemmele
Brian Roemmele@BrianRoemmele·
BREAKTHROUGH: A full, unmodified 2.78-trillion-parameter Kimi K3 on a consumer laptop by streaming only the activated experts from NVMe. YOU CAN’T RUN KIMI K3 “ON THAT” THEY DECLARED. There are many paths to do it. This is one: Marco Bambini Just Gave Us the Full Kimi K3 on a Laptop Meet Marco Bambini he did something that felt impossible only a day ago. He built WASTE Weight-Aware Streaming Tensor Engine a clean, dependency-free C inference engine that runs the complete, unmodified 2.78-trillion-parameter Kimi K3 model by streaming only the activated experts straight from NVMe. No distillation. No pruning. No cloud. The full open-weight model. We have it running in the lab right now. What Marco Actually Built Kimi K3 is a sparse Mixture-of-Experts system. Only about 4 % of its weights fire on any given token. Marco’s insight was simple and ruthless: the idle experts do not need to live in RAM. They only need to be reachable in time. WASTE keeps the model’s “trunk” (attention, shared components, embeddings) resident in memory — roughly 27 GB on the converted container. The 82,000+ routed experts stay on disk as tightly packed residual vector-quantized records. When the router selects its 16 experts per layer, the engine issues direct, cache-bypassing reads from the internal NVMe and feeds them into a bounded expert cache. The rest of the machine’s RAM becomes working space for that cache. On a 64 GB MacBook Pro with the container on the internal SSD, we are measuring 0.32–0.34 tokens per second at a comfortable memory budget. Prefill sits a little higher. The vision tower works. Logits match the reference implementation to within a few parts in a million. It is the real model. The container itself is 982 GiB after conversion from the original 1.42 TB MXFP4 weights. Minimum RAM floor is just over 29 GB for short context. Push the budget higher and the expert cache hit rate climbs; push too high and you start paging and the speed collapses. The sweet spot on current consumer hardware is clear and measurable. How We Are Testing It We converted the official weights, verified the container, and began systematic runs the same day the engine stabilized. First we confirmed numerical fidelity against the PyTorch reference on short prompts. Then we moved to longer generation, vision inputs, and multi-turn chat using Kimi’s native XTML format. We are measuring wall-clock decode, expert I/O versus compute split, cache hit rates at different RAM budgets, and thermal behavior under sustained load. We are also exercising the OpenAI-compatible server that sits on top of the same C library so we can drop the model into existing agent loops without rewriting anything. Early observations: •Expert I/O dominates the timeline, as expected. On a fast internal NVMe the engine is already near the practical ceiling of the storage subsystem. •The architecture’s sparsity is the entire enabler. A dense model of this size would be dead on arrival for local use. •Context length is currently limited by RAM more than by the model itself. Practical working contexts sit comfortably in the tens of thousands of tokens on 64 GB hardware; the full million-token window will need more memory or smarter KV management. •Thinking tokens are expensive at this speed. Long internal monologues turn into multi-hour runs. For agent work we are already experimenting with tighter control over when full reasoning is requested. We are treating this as a research instrument, not a finished product. Every run teaches us something about expert locality, prefetch opportunities, and how far pure software streaming can push trillion-scale inference on ordinary machines. 1 of 2
Brian Roemmele tweet media
English
69
152
1K
102.4K
John Moore
John Moore@talemi·
@thdxr Yeah I think directionally right but anecdotally in my circles it's a lot more sharp - either you get it or not and there is no "kinda" use it. That applies within the tech crowd too - I still know engineers who insist it's not usable and they should manually code everything.
English
0
0
0
8
dax
dax@thdxr·
the impact of ai on the average person right now is similar to google search it's ubiquitous, everyone is using it all the time. but they're not exactly seeing it as life changing so understandably it's very confusing what all the commotion is about
English
96
33
1.3K
61.2K
John Moore
John Moore@talemi·
@axiomofmind @BrianRoemmele I think @LeetOffice is right. It's the worst it'll ever be. I would also extend to if you can do it for "free" on your existing hardware and you are time insensitive then that's budget you free up to do other things on paid pathways.
English
0
0
2
55
@Axiomofmind ⚡
@Axiomofmind ⚡@axiomofmind·
@talemi @BrianRoemmele So you'd rather run a simple task at 0.32 tk/s over a long time instead of just using something faster to achieve the same results in 100x less time?
English
6
0
3
331
John Moore
John Moore@talemi·
@axiomofmind @BrianRoemmele The point is that not all LLM interactions are chat. Some of us have scheduled tasks, loops, graphs, or other long running tasks that don't require us to interact directly. A subset of those aren't time sensitive and therefore speed doesn't matter the same way.
English
1
0
5
329
modifylife
modifylife@FallisAaro51194·
@wicksplay Download it even if you can't run it make a local copy to share in case it gets band
English
1
0
5
5.1K
wick
wick@wicksplay·
Kimi K3 goes open weights tomorrow Heres what to do to prepare: 1) Find $500,000 to spend on hardware
English
160
141
5.4K
804.4K
Tibo
Tibo@thsottiaux·
From Science Fiction to Science Reality. Join the team if you want to work on some of the coolest and most impactful technology.
English
706
62
3.6K
515.5K
John Moore
John Moore@talemi·
@thsottiaux Like many others before me... I had indeed just used a banked reset. The horror!
English
0
0
1
1K
Cursor
Cursor@cursor_ai·
Introducing Cursor Router, our intelligent model router that selects the right model for the task at hand. Router delivers frontier-quality results at 60% lower cost.
English
468
656
10.2K
2.6M
John Moore
John Moore@talemi·
@thdxr Obviously false. Especially with the rate smaller open weights are progressing the norm will be for there to be purpose tuned models embedded everywhere. Not one player. I don't believe a majority of users want ChatGPT having their bank details, their health records, et cetera.
English
0
0
2
2.4K
dax
dax@thdxr·
the agent you're building into your product is most likely not going to work just drop it, what's going to happen is there will be a few agent products that individuals and teams choose to use and they will want everything there. might mean they don't even want your product
English
120
28
1.2K
111.8K
John Moore
John Moore@talemi·
@steipete Idk, seems like both have their place in the tool belt.
English
0
0
0
3.6K
John Moore
John Moore@talemi·
Amazing if all this translates in practice. I love where this is going! Our vision of the future feels closer than ever and your own personal memory server is going to be table stakes soon.
Kimi.ai@Kimi_Moonshot

Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: platform.kimi.ai 🔗 Tech blog: kimi.com/blog/kimi-k3

English
0
0
0
48
John Moore
John Moore@talemi·
@sama Agreed - I've shifted to the same. It's hard to waste time typing it and formatting it when the LLM doesn't give a shit.
English
0
0
5
791
Sam Altman
Sam Altman@sama·
i talk to chatgpt more than i type to it at this point new voice model really crossed a threshold
English
2K
427
14.2K
1.1M
John Moore
John Moore@talemi·
@sama Agreed - cooking big time over there buddy! I also agree about the freedom and agency - that's why we are building Talemi.ai! Would love to discuss 😉
English
0
0
0
334
Sam Altman
Sam Altman@sama·
we did not have our best last 12 months ever, which is mostly my fault, but we are about to have our best 12 months to date. the team is doing amazing work and i think you’ll be very happy with what they’ve got cooking for you. i am happy about this for many reasons, but mostly because i care about our users winning. AI has to be about giving lots of people more freedom, agency, and wealth. we want to do the right thing, but we do not want to scare people into doing our thing.
English
2.3K
890
24.7K
2.6M
John Moore
John Moore@talemi·
@thsottiaux @OpenAI I switched from Claude Code to Codex because of 5.6 Sol - top frontier model and usage limits that last longer than 1 day a week for my $200. It was a no brainer!
English
0
0
0
171
John Moore
John Moore@talemi·
@alohaproxy Solve a valuable problem... can't tell you if Claude or the MBP will be useful or not though.
English
0
0
0
14
adah
adah@alohaproxy·
Just got Claude Max and a MacBook. If you had these two things……how would you make your first $1M?
adah tweet media
English
69
0
122
6.5K
Bindu Reddy
Bindu Reddy@bindureddy·
Coming soon - open and closed source model releases - Opus 5 - Gemini 3.5 Pro ( checkpoints are better) - DeepSeek v4 - Kimi 3 Fable 5 will remain the top model for coding for at-least a month
English
53
17
401
37.3K