Thomas Konings

59 posts

Thomas Konings banner
Thomas Konings

Thomas Konings

@tkon99

Doing cool AI stuff. Consultant by day.

Europe Katılım Mayıs 2012
201 Takip Edilen39 Takipçiler
Sabitlenmiş Tweet
Thomas Konings
Thomas Konings@tkon99·
You don't need $10k worth of GPUs to do cool local AI stuff. Just ran embeddings and feature extraction on my RTX 4070 Super 12GB. NuExtract3 + nomic-embed. GLM 5.2 optimizing the extraction as we speak. Benchmarking different speculative decoding options.
Thomas Konings tweet media
English
1
0
0
108
Thomas Konings
Thomas Konings@tkon99·
@PrismML Not just editing files but navigating reference websites using Playwriter, taking text and assets from them and doing a full modern redesign. Honestly surprised.
English
0
1
2
138
Thomas Konings
Thomas Konings@tkon99·
Finally had the time to test Bonsai by @PrismML out. On my mere RTX 4070 Super I get 45 t/s output. Quick addition to ZCode and it's building websites, the thing is incredibly agentic for the hardware it's running on! It's going to be an awesome year for Local AI.
Thomas Konings tweet media
Thomas Konings@tkon99

This chart from the Bonsai model release by @PrismML will be all over LinkedIn tomorrow :) Intelligence density is going to be a huge benchmark going forward, matching the incentives created by memory shortages. Imagine harnessing all the memory we have out and about already.

English
2
3
11
2K
Thomas Konings
Thomas Konings@tkon99·
This chart from the Bonsai model release by @PrismML will be all over LinkedIn tomorrow :) Intelligence density is going to be a huge benchmark going forward, matching the incentives created by memory shortages. Imagine harnessing all the memory we have out and about already.
PrismML@PrismML

Raw capability determines what a model can do. Intelligence density determines where it can do it. Bonsai 27B moves the Pareto frontier left again: 27B-class capability in a footprint smaller than many full-precision 2B models. By intelligence density, 1-bit Bonsai 27B delivers 0.53 per GB - more than 10x the full-precision baseline and roughly 2.7x the best conventional low-bit alternative.

English
0
6
12
3.8K
Thomas Konings
Thomas Konings@tkon99·
Did some more tinkering and now at 10 slots doing 354 t/s! Cool stuff and should run through my 1000 text extracts quickly.
English
0
0
0
24
Thomas Konings
Thomas Konings@tkon99·
6 seems to be the sweet spot. Cut down the run time from 2.5 hours down to 1 hour. Great to be using local models for extraction tasks, really not something I want to be using premium tokens on.
Thomas Konings tweet media
English
1
0
0
22
Thomas Konings
Thomas Konings@tkon99·
You don't need $10k worth of GPUs to do cool local AI stuff. Just ran embeddings and feature extraction on my RTX 4070 Super 12GB. NuExtract3 + nomic-embed. GLM 5.2 optimizing the extraction as we speak. Benchmarking different speculative decoding options.
Thomas Konings tweet media
English
1
0
0
108
Thomas Konings
Thomas Konings@tkon99·
👀The Codex App for Windows has the codex:// protocol to hotlink to skills, settings, automation, projects and even pass prompts. Seems to be disabled by default but enabling is super easy: gist.github.com/tkon99/13fd3bc…
English
0
0
0
66