Joe Devon

42.1K posts

Joe Devon banner
Joe Devon

Joe Devon

@joedevon

CoFounder: #GAAD (Global Accessibility Awareness Day) 200M+ social media reach https://t.co/OVSwb5wEuN, an accessibility benchmark @a11yaudits my co. @A11yGenAI my podcast

Henderson, NV Katılım Şubat 2009
6K Takip Edilen6.2K Takipçiler
Sabitlenmiş Tweet
Joe Devon
Joe Devon@joedevon·
I'm excited to announce that The White House will be doing a Global Accessibility Awareness Day (#GAAD) Event and Town Hall tomorrow at 10AM EST. Links to Livestream and Access Streamtext in the comments. Thanks @WhiteHouseOPE46! From blog post to The White House. 🤯
English
12
18
72
11.4K
Joe Devon
Joe Devon@joedevon·
@jpalioto Out of curiosity, why did you leave Google considering you still believe in them? Asking respectfully, this is not intended as a gotcha or anything.
English
1
0
1
1.4K
John P Alioto
John P Alioto@jpalioto·
Not at Google anymore, but I will tell you the same thing now as then. Do not count Gemini out. The people working on it are absolutely brilliant. The big company structure slows them down. But they are absolutely brilliant.
English
64
27
1.1K
83.2K
OrganicGPT
OrganicGPT@OrganicGPT·
@doodlestein I wish we could change the order tho, it gets cluttered and disorganized so easily. and it's not customizable like CC:
OrganicGPT tweet media
English
1
0
1
96
Jeffrey Emanuel
Jeffrey Emanuel@doodlestein·
Codex Pro Tip: You should check basically every box in the options for the /statusline feature, especially if you have a high resolution display. It's really useful to have all that stuff easily observable at a glance. There are many more than are shown in that image below, btw.
Jeffrey Emanuel tweet media
English
10
0
54
5.7K
Joe Devon
Joe Devon@joedevon·
Thank you. In the meantime I asked my agent (out of laziness) to install Post Compact reminder. It generated 30 beads: - 16 unit-test tasks - 10 end-to-end test tasks - 1 test harness - 1 test runner - 1 CI integration - 1 actual installation task So I just installed it manually. I think the request tripped my @simonw TDD workflow skill. lol
English
1
0
2
43
Jeffrey Emanuel
Jeffrey Emanuel@doodlestein·
@joedevon @andyhennie Eidetic Engine (ee) is extremely useful, I use it in my consulting work for hedge funds. And I use repo_updater (ru) dozens of times a day to keep all my projects in sync with GitHub (and across machines). Post Compact Reminder is a "set it and forget it" kind of thing for cc.
English
1
0
2
165
Joe Devon
Joe Devon@joedevon·
This is absolutely incredible.
Brian Roemmele@BrianRoemmele

BREAKTHROUGH: A full, unmodified 2.78-trillion-parameter Kimi K3 on a consumer laptop by streaming only the activated experts from NVMe. YOU CAN’T RUN KIMI K3 “ON THAT” THEY DECLARED. There are many paths to do it. This is one: Marco Bambini Just Gave Us the Full Kimi K3 on a Laptop Meet Marco Bambini he did something that felt impossible only a day ago. He built WASTE Weight-Aware Streaming Tensor Engine a clean, dependency-free C inference engine that runs the complete, unmodified 2.78-trillion-parameter Kimi K3 model by streaming only the activated experts straight from NVMe. No distillation. No pruning. No cloud. The full open-weight model. We have it running in the lab right now. What Marco Actually Built Kimi K3 is a sparse Mixture-of-Experts system. Only about 4 % of its weights fire on any given token. Marco’s insight was simple and ruthless: the idle experts do not need to live in RAM. They only need to be reachable in time. WASTE keeps the model’s “trunk” (attention, shared components, embeddings) resident in memory — roughly 27 GB on the converted container. The 82,000+ routed experts stay on disk as tightly packed residual vector-quantized records. When the router selects its 16 experts per layer, the engine issues direct, cache-bypassing reads from the internal NVMe and feeds them into a bounded expert cache. The rest of the machine’s RAM becomes working space for that cache. On a 64 GB MacBook Pro with the container on the internal SSD, we are measuring 0.32–0.34 tokens per second at a comfortable memory budget. Prefill sits a little higher. The vision tower works. Logits match the reference implementation to within a few parts in a million. It is the real model. The container itself is 982 GiB after conversion from the original 1.42 TB MXFP4 weights. Minimum RAM floor is just over 29 GB for short context. Push the budget higher and the expert cache hit rate climbs; push too high and you start paging and the speed collapses. The sweet spot on current consumer hardware is clear and measurable. How We Are Testing It We converted the official weights, verified the container, and began systematic runs the same day the engine stabilized. First we confirmed numerical fidelity against the PyTorch reference on short prompts. Then we moved to longer generation, vision inputs, and multi-turn chat using Kimi’s native XTML format. We are measuring wall-clock decode, expert I/O versus compute split, cache hit rates at different RAM budgets, and thermal behavior under sustained load. We are also exercising the OpenAI-compatible server that sits on top of the same C library so we can drop the model into existing agent loops without rewriting anything. Early observations: •Expert I/O dominates the timeline, as expected. On a fast internal NVMe the engine is already near the practical ceiling of the storage subsystem. •The architecture’s sparsity is the entire enabler. A dense model of this size would be dead on arrival for local use. •Context length is currently limited by RAM more than by the model itself. Practical working contexts sit comfortably in the tens of thousands of tokens on 64 GB hardware; the full million-token window will need more memory or smarter KV management. •Thinking tokens are expensive at this speed. Long internal monologues turn into multi-hour runs. For agent work we are already experimenting with tighter control over when full reasoning is requested. We are treating this as a research instrument, not a finished product. Every run teaches us something about expert locality, prefetch opportunities, and how far pure software streaming can push trillion-scale inference on ordinary machines. 1 of 2

English
0
1
13
3.7K
Joe Devon
Joe Devon@joedevon·
Just crazy. Of this list: Daily Drivers: NTM, Agent Mail, CASS (except it's been broken last couple days, but as you told me fix is on the way), ACFS (on a beefy contabo server), beads rust (and AI uses beads-viewer, I think). That by itself is amazing. Intrigued by: Eidetic Engine (curious if and how you use it regularly), Repo Updater, Post Compact Reminder. Also intrigued by all those crazy FrankenInfra you built. Thanks for sharing. I'm going to create a skill that summarizes all your new tweets daily because I've become dependent on your stack and need to understand what's going on. You're so productive it's a job just to keep updated with what you're building! lol. P.S. I think I have post compact reminder on the contabo box but I will double check and add it to my other machines. That will allow me to reduce my CLAUDE.md usage to only CC quirks.
English
1
0
2
176
Joe Devon
Joe Devon@joedevon·
What to do when your AI puts you in the corner until 12:30 when rate limit resets?
English
5
0
2
218
Joe Devon
Joe Devon@joedevon·
@simonw Don't you love when you ask it to research for something very specific like panspermia, expecting an arxiv or wikipedia search but instead it says "searching instagram and pinterest"?
English
0
0
2
251
Simon Willison
Simon Willison@simonw·
As a paying customer of theirs it really matters to me that I understand how their search works - it helps me evaluate how much I can trust the results
English
7
4
112
8.9K
Simon Willison
Simon Willison@simonw·
It's wild to me that both Anthropic and OpenAI have products that lean so hard on search, and yet they both obscure the underlying search index that they are using
English
84
28
1.1K
136K
Joe Devon retweetledi
Simon Willison
Simon Willison@simonw·
I hope there are QA testing experts out there who are thinking "finally, we don't need software developers any more!", rolling up their sleeves and building really great software right now
English
79
48
1K
82.9K
Joe Devon
Joe Devon@joedevon·
@Coderguy Shape. Tripped. Misfired. More is better here. That's the whole point. matters a lot.
English
0
0
0
6
Daniel Ice
Daniel Ice@Coderguy·
I know a lot of people talk about LLM writing styles and overuse of the em dash — I am starting to feel like LLM has an accent. I notice a big use of the following terms. My list is created to show the AI term and what as a Software Develop I am used to say - spine: core, backbone - slice: phase, project, sprint, release - flake: side path, side quest, rabbit trail What's everyone else seeing out there? Any patterns that you are seeing?
Daniel Ice tweet media
English
1
0
3
48
Joe Devon
Joe Devon@joedevon·
@DrMattParker I do but use Fable as orchestrator so when it dies in middle of something...that makes everything stand still.
English
0
0
0
14
Joe Devon retweetledi
ClaudeDevs
ClaudeDevs@ClaudeDevs·
MCP 2026-07-28 is live and it's the largest update to the protocol since launch. MCP is now stateless, making it easier to deploy and scale remote servers. claude.com/blog/bringing-…
English
361
1.4K
13.2K
2.7M
Martin Devon
Martin Devon@MartinDevon·
@joedevon You’re right. You asked me not to say load bearing earlier in the chat and I just ignored it. That’s on me.
English
2
0
0
19
Joe Devon
Joe Devon@joedevon·
Fable: "Three of those deserve receipts, not vibes." WHY CAN'T YOU TALK LIKE A HUMAN BEING? Anyone else notice that if you configure OpenAI models as friendly, their writing is now superior to Anthropic?
English
3
0
6
479
Joe Devon retweetledi
Peter J. Liu
Peter J. Liu@peterjliu·
Reading Kimi K3 paper, I'm struck by the number of innovations and departures from orthodoxy. These guys are doing real research and not copying. It is no longer the case of being, "mostly the Transformer with a few Noam mods". I mean they got rid of the position encodings/embeddings! Bold! More thoughts to come after reading.
English
40
255
4K
222.9K