James Aita

3.4K posts

James Aita banner
James Aita

James Aita

@jaita

Building with Local AI | DGX Spark EdTech | Health-Tech | House Music 🎹💻🎧🎶 alt @arxyo1 Need more VRAM.

Katılım Kasım 2008
940 Takip Edilen319 Takipçiler
James Aita
James Aita@jaita·
@Teknium @NousResearch That would be my preferred method of using it.. Something like hermes desktop / dashboard / webview but not on the host computer. Telegram does the job now but a more direct approach would be amazing.
English
0
0
1
14
Omar Sanseviero
Omar Sanseviero@osanseviero·
We're thinking of doing a special celebration for Gemma 1 billion downloads Send us the top projects you've seen or built with Gemma!👇
English
39
15
298
17.1K
James Aita
James Aita@jaita·
Ok, qwythos 27B is already causing problems. Tool calls leaking into Hermes agent convos...
English
0
0
0
7
James Aita
James Aita@jaita·
What's the coolest thing you've done with your DGX Spark?
English
0
0
0
8
Jack Price
Jack Price@jackprice·
I want to connect with more founders builders vibe coders AI enthusiasts UI designers If you’re someone building with AI right now, drop it in the comments and let's connect
English
510
9
551
32.5K
James Aita
James Aita@jaita·
@sudoingX I don't understand Hermes desktop. I mean it seems great, but I run Hermes agent off a mac-mini so it's sandboxed and doesn't blow up my main laptop. And hermes and openclaw rm -rf stuff all the time and I gotta go running to my backups to restore them.
English
0
0
1
82
Sudo su
Sudo su@sudoingX·
hermes agent desktop users, how are you finding it so far? and the useful question, what's the one change or feature that would actually solve a problem you're hitting? drop it below, i will push it to the people who build this so real feedback here has a real shot at landing.
English
9
0
13
1.6K
Light Silver
Light Silver@lightsilver323·
drop your project URL you never know who might DM you
English
191
4
69
12.8K
James Falconi
James Falconi@Joi2James·
I’m looking for products with soul. Drop your website/app
English
143
1
55
5.3K
Light Silver
Light Silver@lightsilver323·
Built something cool? Drop the link.
English
103
2
56
7.2K
James Aita
James Aita@jaita·
@SpaceTimeViking Yes, this is an issue with VLLM I think - not sure if your variant has fixed it, there's a bunch of logged issues with VLLM about this. Shows up when you try to run benchmarks (e.g. using llama benchy or sparkrun).
English
0
0
0
6
ÆON FORGE ✨
ÆON FORGE ✨@SpaceTimeViking·
Can you elaborate? I was testing it with a config of up to 128 concurrent seqs on Qwen 27b so.. maybe? I wasn't aware there was a concurrency lock. You just have to make sure you set max seqs flag to the number of concurrent session you want to use. if you set it to 16 anything above that will start to queue up vs run in parallel but if you set it to 32 then you get up to 32 concurrent sessions.
English
1
0
1
40
ÆON FORGE ✨
ÆON FORGE ✨@SpaceTimeViking·
Announcing the latest vLLM specifically built and optimized for the DGX Spark aeon-vllm-ultimate 0.26.0! A lot of new features & stability enhancements. Adds --prefix-match-unit 16 Support for NVFP4_AWQ +All previous enhancements carried forward! github.com/AEON-7/vllm-ul…
English
3
9
111
5.2K
James Aita
James Aita@jaita·
Definitely needing a second DGX spark. One just isn't going to cut it. Ugh...
English
0
0
1
12
James Aita
James Aita@jaita·
@0x0SojalSec wait, how are you running it on a macbook pro, it's still 1/2 TB.
English
0
0
0
13
Md Ismail Šojal 🕷️
Kimi K3 locally Run on MacBook Pro. testing the 1-bit version on a MacBook Pro and the results surprised me. Side-by-side comparisons show it holding its own against much larger closed models. 1-bit delivers impressive reasoning, coding, and creative performance. No cloud dependency. And Thanks to Unsloth’s dynamic quantization, one of the strongest open models available (Kimi K3) is now practical on high-end consumer hardware.
English
0
1
11
948
James Aita
James Aita@jaita·
@sudoingX You're right on the money.. 60-85 gb and you're fine when you add in context and the rest. The question is, has anyone fine tuned Laguna or added a chat template yet to get it acting consistently? DS4-flash is a bit over the line but you can get a solid single lane in.
English
0
0
0
64
Sudo su
Sudo su@sudoingX·
if you own a single dgx spark and want to run the big moes, here's the one rule that saves you a wasted weekend: keep your model weights under about 80 gigs. that's the balance point. weights are only half the bill, the other half is kv cache, and that's where your context actually lives. stay under 80 and you've got 35 to 45 gigs left to scale context and run a spec decoder. cross it and you're fighting the box for every token. three moes i've run on mine and exactly where they land: 1. laguna s 2.1 - 67 gigs, a genuine coding beast. ~35 tok/s with its dflash drafter, up to 45 on sustained code, poolside from scratch and it holds context for days. huggingface.co/poolside/Lagun… 2. qwen 3.5 122b - 74 gigs, the all-rounder. ~35 tok/s with mtp and it holds that line whether it's writing code or prose, no swings. a few months old and still trading punches with the newest king. huggingface.co/unsloth/Qwen3.… 3. stepfun 3.7 flash - the cautionary one. the q4 crams in at 108 gigs, way over the line, so it runs slow with almost no room to breathe, and the nvfp4 won't load at all, it wedged my box twice trying. this is what living above 80 looks like. huggingface.co/stepfun-ai/Ste… the two i'd put on your desk tonight are laguna and qwen, both moe, both under the line, both a joy to run. that's the spark's whole personality, feed it a fat moe that only fires a few billion params a token and it flies. stay under 80, keep the context room, and the box never lets you down.
English
19
5
97
5.5K
Thomas Unise
Thomas Unise@thomasunise·
Anyone who's hyping up Fish Audio to you right now with their new real time model, is absolutely lying. I just implemented it in one of our voice agents and their default voices sound like robots from 2022, the latency is trash, and they are nowhere near the quality of elevenlabs or OpenAI real time I'd get my money back for my api tokens if could Huge bummer tbh
Fish Audio@FishAudio

Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro. >It can clone a voice from 5 seconds of audio >2x faster than Cartesia & 1/6th the cost of Eleven Labs >most expressive model with word level control over emotion, intonation, pacing etc We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production. If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free. Book a demo: s.fish.audio/tmapke To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.

English
4
0
4
278