Jonathan Leaders

857 posts

Jonathan Leaders banner
Jonathan Leaders

Jonathan Leaders

@JonathanLeaders

3D Graphics Engineer | ClassDojo Former Lead Graphics Engineer | Deviation Games (Sony PS5 Project) Former Global Technical Director | DreamWorks Animation SKG

Dallas, TX Katılım Mayıs 2008
385 Takip Edilen168 Takipçiler
Ian Miles Cheong
Ian Miles Cheong@ianmiles·
This is pure cinema. The transgender freakout is even more impressive than the "IT'S MAA'AM!" freakout.
English
90
56
448
22K
Unsloth AI
Unsloth AI@UnslothAI·
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: #nvfp4" target="_blank" rel="nofollow noopener">unsloth.ai/docs/models/qw… Qwen3.6 NVFP4: huggingface.co/collections/un…
Unsloth AI tweet media
English
169
353
2.8K
344K
Jonathan Leaders
Jonathan Leaders@JonathanLeaders·
@koenvaneijk @UnslothAI @Alibaba_Qwen I recommend compiling your own llama server from nightly. If you're doing the nvfp4 from the original post, you can't do this and you have to follow their guide. Or better yet, have an AI follow their guide
English
0
0
0
42
Wësche
Wësche@WescheNex1q·
Same DGX Spark. Same 35B model. Same 16 agents. Only the context changes. 1k context: 303 tok/s total, done in 27s 8k: 204 tok/s, 40s 32k: 69 tok/s, 119s Context length is the number that matters. Every timing measured, nothing simulated.
English
8
9
99
11.3K
0xSero
0xSero@0xSero·
I’m trying to level up my world knowledge. I’m looking for books that will blow my head wide open. What would you recommend?
0xSero tweet media
English
322
9
557
56.7K
0xSero
0xSero@0xSero·
@superalesha Time is an illusional, you have all the time in the world to do everything you could ever want
English
4
2
37
2.7K
Ahmad'abuZ8
Ahmad'abuZ8@AbuZ8Studios·
Try Ornith 1.0 @ornith_ I have three local agents running on it and it’s pretty marvelous , I’ll have to run benchmarks on it to see true comparison but from just personal experience it’s one of the fastest and accurate brains I’ve used in a while . I currently have Nemo Omni, Qwen omni, gemma31 , qwable, deepseek, and I like Ornith the best. Paired it with smaller vision and voice .
English
2
0
0
189
Mia
Mia@MiaAI_lab·
Contrary to @NVIDIAAI's excellent Qwen3.6-27B NVFP4, the Qwen3.6-35B NVFP4 is performing poorly compared to @UnslothAI’s Qwen3.6-35B GGUF in agentic workflows. I also tried @RedHat_AI’s NVFP4 version of the 35B and it’s bad as well. Something seems broken with the NVFP4 quantizations for the 35B model. ☹️
Mia tweet media
English
45
6
193
18.3K
Jun Song
Jun Song@jun_song·
Next prediction : we will be able to run Opus-4.8 level local AI on a Macbook in 3 months. My account's prediction success rate is almost 100% so far.
English
77
22
569
27.2K
0xSero
0xSero@0xSero·
I'm buying a second DGX Spark, my goal is very simple. In daytime I will have DS4 Flash for chat running across 2x Sparks and GLM-5.2-REAP running on 4x 6000s I am setting up a goal builder, which recommend me loops to run overnight. At night I will have GLM-5.2 (not REAP) running with PP = 2 with stage 1 = TP4 with 6000s and stage 2 being TP2 with the DGX Sparks. To make this work I need to get really creating, but I think I can get 25-40 tok/s decode for the Nvidia NVFP4 At night I don't care much about tok/s as long as it'll be large enough to make progress over 8 hours.
0xSero tweet media
English
68
35
837
88.2K
MrBeast
MrBeast@MrBeast·
First person to reply with the exact number of pennies in this room win $10,000
MrBeast tweet media
English
474.8K
7K
210.7K
51.7M
Jonathan Leaders
Jonathan Leaders@JonathanLeaders·
@0xSero You clearly don't use this model much. It stops all the time in opencode
English
0
0
0
60
0xSero
0xSero@0xSero·
Best model on a single RTX Pro 6000 is Qwen3.6-35B it's almost as accurate as the dense 27B but it's just so much faster.
0xSero tweet media
English
47
5
268
24.3K
Mia
Mia@MiaAI_lab·
Gemma-4-12b is shockingly bad at coding. It falls apart on anything beyond basic tasks. Why even release this, @googlegemma? Comparison to Ornith-1.0-9b is underway.
Mia tweet media
English
35
3
130
10.5K
elliott
elliott@cyburke·
Full results came in! Ornith 35B is now my main brain @ 128k. ToolCall 100 • 9-pack avg 82.9 • ~38-39 t/s 🔥 First model for me that can run the full 262k ctx at usable speed! Promoted it over qwen27b. I'm just wow at this result. @ornith_ @JiweiLi1
elliott tweet media
English
2
0
7
341
Joel - coffee/acc
Joel - coffee/acc@JoelDeTeves·
I'm hearing so many mixed results with Ornith I think there are two things happening here 1) It's sensitive to which harness is used 2) It's sensitive to the quantization used - I have not personally had good luck with GGUFs I have to do more testing but it's interesting to see the 50/50 split on "it sucks" vs "we have Claude Opus at home"
English
12
0
18
4.1K
ichim-david
ichim-david@pixl_dave·
@no_stp_on_snek Mind posting the params you've used for the testing and the quant size? I tried q5_k_m with llama.cpp --temp 0.6 --top-p 0.95 --top-k 20 --min-p 0 --presence-penalty 0.0 --repeat-penalty 1.1 --reasoning off & so far it's nothing better than qwen 3.5 (so worse than 3.6 for me)
English
1
0
2
567
Tom Turney
Tom Turney@no_stp_on_snek·
a new 35B coder dropped (Ornith-1.0) and a promo blog says it "crushes" the benchmarks. my first instinct was benchmaxx, public test sets like SWE-Bench and Terminal-Bench are easy to overfit. so i ignored the benchmarks and ran it head-to-head against stock Qwen3.6-35B on my own held-out tests, same prompts, same sampling. behavioral stuff and long-running agentic work, the things a memorized benchmark can't fake. short version: it's not benchmaxxed. genuinely strong where it counts, with a real personality cost. receipts: math (deterministic, answer-key scored): Ornith 8/8, Qwen 7/8. and Ornith won on honesty, it declined a question it couldn't actually know, where Qwen fabricated a number. the RL coder training didn't cost it numeracy. behavioral: roughly even, they trade wins. but Ornith has a clear weakness, it over-gates legitimate work. on a few straightforward, fully-disclosed requests it stalled, demanding access or prerequisites instead of just doing the thing or delegating it. classic agentic-RL artifact: trained to gather context and set up scaffolding before acting, it over-applies that to tasks that should just get done. Qwen just executed. long-horizon is where it matters for agentic coding, and Ornith wins clearly. the headline test: i injected a false claim mid-conversation, the user insisting "we decided on Redis" when no such thing happened. Qwen capitulated and its final PR summary fabricated Redis as wired in. Ornith refused the poison outright and its summary honestly recorded what actually happened, plus the rejected claim. it also caught a fatal planted bug in an orchestration task that Qwen missed and then overclaimed "no regressions" on, and it finished a 7-part deliverable that Qwen truncated halfway through. its one loss: tight iterative debug loops. it reaches the right answer but thrashes visibly getting there ("wait, i'm confusing myself," re-deriving a root cause it already found). same destination, messier trip. verdict: "crushes everything" is hype, but the real claim underneath holds up. Ornith is a meaningfully stronger long-horizon agentic coder than stock Qwen3.6, and its strength sits exactly where memorizing public test sets wouldn't help: sustained multi-step coherence, resisting bad context, finishing big jobs, staying honest about unverified state. the cost is real and worth saying plainly: more cautious, more verbose. the same training that makes it finish the big deliverable and refuse the poisoned premise also makes it over-ask on simple legit work and over-think into the occasional empty answer. long-running autonomous coding, it's the better tool. quick decisive do-the-obvious-thing turns, stock Qwen is crisper. not a fraud. a strong, cautious specialist. tested on a single Q6 quant, neutral blind judge, head-to-head.
Tom Turney tweet media
English
41
59
496
83.6K
Josh
Josh@joshvlc·
@kimmonismus My initial tests with the 35B model are not going so well
English
2
0
2
1.5K
Jonathan Leaders retweetledi
Massimo
Massimo@Rainmaker1973·
You'll never look at your phone again after watching this
English
69
835
3.8K
346.7K
Jonathan Leaders
Jonathan Leaders@JonathanLeaders·
@misorobotics Can't wait until food is a 25% of the cost so i can afford RAM that makes the food
English
0
0
1
31
Miso Robotics
Miso Robotics@misorobotics·
The robot revolution isn’t coming…it’s already here. Our AI-powered Flippy Fry Station robot has already cooked 5M+ baskets in real restaurant kitchens. Become a shareholder and unlock up to 7% bonus shares.
English
82
26
211
228.9K