SGode

11.8K posts

SGode

SGode

@Seb_Gode

Katılım Temmuz 2023
65 Takip Edilen252 Takipçiler
SGode
SGode@Seb_Gode·
@GPTWare @robertomasymas @LottoLabs The current variant is better but it again started to loop 2x for me today :D They are working on it I guess, for some reason no quant is really good at the moment, only the base one is nice
English
0
1
3
17
Lotto
Lotto@LottoLabs·
Gotta figure out how to get this on localmaxxing 🥹 only got 48GB vram
Eiso Kant@eisokant

Today we are releasing Laguna S 2.1. At 118B total parameters, with 8B active per token, it does the work of models several times its size on agentic coding. It is remarkably persistent across long-horizon tasks. And it is small enough to run on a single NVIDIA DGX Spark. It is far more capable than anything we have created before, and I think it redefines what a model in its weight class can do. Laguna S 2.1 is an important model for Poolside. What it represents is even more important. If, five years ago, I had read a book that said that by 2030 everything economically valuable, scientifically interesting, and personally meaningful would be built on intelligence contracted from three or four companies, I would have called it dystopian science fiction. We are at a fork in the road of what kind of world we can have. I believe intelligence should and will become a commodity. The question is whether that intelligence comes from three companies, or from many people who can build it, own it, and shape it. The open ecosystem will not win by being the best in its own category. No one cares who is king of the open-source kingdom. People want the best intelligence for the task they are trying to do, with the right balance of quality, speed, cost, and control. If we want a different future, open models have to be on par with, or better than, their closed equivalents. Laguna S 2.1 is a meaningful step in that direction: capable enough to compete far above its weight class, efficient enough to run on hardware you can own, and open-weight so anyone can build on it. Open-weighting our models is the contribution we can make today toward a world where intelligence can be built and owned by many. And we will keep doing it. I am very proud of this team’s work. A big shout out to everyone at Poolside who made this possible, from infrastructure and data to architecture, pretraining, post-training, evaluations, and inference. Laguna S 2.1 is available today under the OpenMDW-1.1 license, with weights on Hugging Face and access through OpenRouter and our API. We are building toward a future where the most capable intelligence in the world can be owned and shaped by anyone. Laguna S 2.1 is one step. We are going to keep building until that future exists. poolside.ai/blog/introduci…

English
17
0
62
7K
SGode
SGode@Seb_Gode·
@giles_home @AstraiaAI I wonder whether the OpenRouter model was also running the broken NVFP4 variant? It's quite confusing to me how the results in that video are so bad when it doesn't reflect my testing at all. The only thing which sucks is looping at the moment, but they know that already.
English
0
0
1
24
Astraia 🇦🇷🇳🇴
If we are being honest, Qwen 3.6 27B is by far, and beyond any doubt, STILL the best open model under 200B parameters right now. Nothing I have tested even comes close, except for its own finetunes.
English
41
11
344
23.2K
Sudo su
Sudo su@sudoingX·
everyone's running laguna s 2.1 on completely different hardware, and i want the full picture in one thread. if you've got it running, drop the exact variant (nvfp4, fp8, gguf, mlx, whichever), your hardware, the quant, the tok/s you're actually seeing, and one line on how it feels. i'll start. laguna s 2.1 nvfp4 on one dgx spark, 128gb, 45 tok/s sustained on code with the dflash drafter, holds the full context, best moe on a desk run i've had. no cherry picking, ugly numbers welcome. let's build the community sheet nobody else has.
English
51
7
159
26.4K
SGode
SGode@Seb_Gode·
@robertomasymas @GPTWare @LottoLabs They just released a fixed NVFP4 and INT4 version and updated the weights. Generally this model seems to be hard to quantize as even some Unsloth quants are not the best. But current INT4 variant is running amazing, a lot better in code than Qwen3.6-27B for me.
English
0
0
1
15
Roberto Tomás C 🍉
Roberto Tomás C 🍉@robertomasymas·
@GPTWare @LottoLabs I've seen varied benches of it, incl Mia AI Labs' direct comp to qwen: it underperformed. The ones complaining r using nvfp4 version (but some success stories there too). I think it's early yet, people r working out kinks. Also, they appear to have done a naive nvfp4 port.🫤
English
1
0
3
42
SGode
SGode@Seb_Gode·
@dr_boese Ich hatte sogar schon ein Terminal gesehen wo man manuell 0,00€ eintippen musste und es keine Option für kein Trinkgeld gab. War allerdings nicht in Deutschland, trotzdem recht frech.
Deutsch
0
0
6
279
SGode
SGode@Seb_Gode·
@adamkbaranowski @sudoingX It simply works with --enforce-eager. Flags next to the parsers: --max-num-batched-tokens 4096 --gpu-memory-utilization 0.87 --max-model-len 262144 --enforce-eager --max-num-seqs 4 (probably can adjust this up) Sadly without enforce eager I can't make it fit, but it runs!
English
0
0
0
16
SGode
SGode@Seb_Gode·
@loktar00 Seems like the model is quite hard to quantize correctly
English
0
0
1
15
SGode
SGode@Seb_Gode·
@loktar00 What Laguna did you use? They apparently have a big issue with the NVFP4 quants which decreases capabilities and the Unsloth GGUFs also somehow aren't the best (at least some).
English
2
0
1
401
Loktar 🇺🇸
Loktar 🇺🇸@loktar00·
Laguna S-2.1 vs Qwen3.6 35b and 27b on my visual bench.. 93 prompts, thinking on, every demo has to actually run and be error free in the browser to count. The stats across the run 27B 85/93 (91%) 35B Q8 76/93 (82%) Laguna 73/93 (78%) Love having a new model to test!
English
35
11
195
17.4K
SGode
SGode@Seb_Gode·
@stevibe The gguf's seem to be pretty bad quality but the NVFP4 model is really cool.
English
0
0
1
140
stevibe
stevibe@stevibe·
Laguna S 2.1 running native NVFP4 Picked Blackwell hardware for this side-by-side: 🟢 DGX Spark | 19.44 tok/s | 162ms TTFT 🔵 RTX PRO 6000 | 108.54 tok/s | 266ms TTFT 🟣 4× RTX 5090 (TP4+EP4) | 145.81 tok/s | 521ms TTFT 117.6B MoE (8.5B active), quantized to FP4, on your own box.
English
17
17
226
19.4K
SGode
SGode@Seb_Gode·
@populartourist I'm very happy with it, performs a lot better than Qwen3.6-27B in all my tasks with the NVFP4 quant and the recommended temperature settings.
English
0
0
1
85
wd 🇵🇹🔺
wd 🇵🇹🔺@populartourist·
So far Laguna S 2.1 seems benchmaxxed and performs materially worse than Qwen3.6 35B-A3B - and nowhere near Qwen3.6 27B in terms of FE design and maths. I've tested locally and through their free inference and the results confirm. Q2 quant is poor, and Q4 is probably the bare minimum. Very inefficient to run compared to Qwen3.5 122B-A10B. Likely inference not supported properly yet - not on llama.cpp mainline yet. I'm unsure how it managed such good DeepSWE score. It was too slow to run on agentic work on my setup. Unless it shines in pure SWE work, then it might be a Devstral sibling.
English
42
7
158
19.4K
SGode
SGode@Seb_Gode·
@eric_alcaide It does really well so far! However I've some issues getting it to run with proper performance. I got 4xL4, so 96GB of VRAM which in theory should be enough to run the INT4&NVFP4 version. However I can't get it up without CUDA OOM error, except with --enforce-eager which is sad
English
0
0
0
24
SGode
SGode@Seb_Gode·
@mighty_micha @JStabentheiner @StefanIlle Ja es geht hier aber um die Self-Service-Probefahrten, also da wo du einfach einsteigst ohne mit jemanden zu quatschen. Die Fahrzeugen stehen ja teils einfach an einem Supercharger rum
Deutsch
1
0
3
47
Mighty Micha
Mighty Micha@mighty_micha·
@JStabentheiner @StefanIlle Kommt aber stark auf den Tesla-Händler an. Ich bekam für die Testfahrt einfach ne Schlüsselkarte für das Fahrzeug und es wurde eben ne Zeit ausgemacht, wann ich's zurückbringen soll...
Deutsch
1
0
2
133
Julian Stabentheiner 🔋👨‍💻🐩🍕
Am Supercharger Hilden steht ein verlassener Tesla auf der Straße. Ist ein Testwagen. Vermutlich hat da jemand noch einen Fahrerwechsel gemacht und währenddessen ist die Probefahrt abgelaufen. 😆
Julian Stabentheiner 🔋👨‍💻🐩🍕 tweet media
Deutsch
7
2
95
20.3K
SGode
SGode@Seb_Gode·
@superalesha I only manage to get it working with --enforce-eager but this somehow sucks for throughput :( Really weird how different different setups behave despite having the same VRAM
English
0
0
0
9
Alexey Fateev
Alexey Fateev@superalesha·
@Seb_Gode try this --max-num-seqs 64 --max-cudagraph-capture-size 64
English
1
0
0
36
Alexey Fateev
Alexey Fateev@superalesha·
Laguna S 2.1 runs on four used 3090s. full 262K context, no cuts. The best model for my setup. 117B MoE, INT4, vLLM. 98 tok/s single request, 503 tok/s at 64 concurrent. here is what it does on $6K of used hardware. 152 measurements.
Alexey Fateev tweet media
English
12
4
34
2.3K
Adarsh Mishra
Adarsh Mishra@ItsAdarshM·
@synthwavedd This does not has qwen 3.8max it is i guess almost same price of 3.6 flash yet way better
English
1
0
16
4.5K
leo 🐾
leo 🐾@synthwavedd·
Gemini 3.6 Flash literally got the same score on Artificial Analysis as 3.5 Flash. Worse than Meta Spark 1.1, GLM-5.2, 5.6 Luna, Sonnet 5, Grok 4.5, 5.6 Terra... yikes.
leo 🐾 tweet media
English
90
41
1.4K
184.6K
SGode
SGode@Seb_Gode·
@Day_X79 @Frank45210 @RonnyGolisch Naja Tesla pusht das Standard RWD aber auch extrem stark, ohne den Zinsaktionen usw. glaube würde das nicht so gut gehen :D
Deutsch
0
0
0
32
RadlerRampage
RadlerRampage@Day_X79·
@Frank45210 @RonnyGolisch Irre, war anfangs wirklich davon überzeugt, das das Standard RWD Modell nicht so gut laufen würde, dabei gehen die weg wie warme Semmel😅
Deutsch
1
0
2
163
Ronny ツ
Ronny ツ@RonnyGolisch·
Hui, in Grünheide ziehen die Zeiten gewaltig an! 🚚⚡️ Aktuelle Lieferzeiten fürs Model Y aus der Giga Berlin: 🔴 RWD: Okt - Nov 2026 🔴 Premium Max RWD: Jan - Feb 2027 👀 🔴 Premium AWD: Okt - Nov 2026 🔴 Performance: Sept 2026 Steht da etwa ein großes Akku-Upgrade in den Startlöchern? 🔋👀 Was ist eure Vermutung?
Ronny ツ tweet media
Deutsch
58
17
340
41.7K
SGode
SGode@Seb_Gode·
@miquelaymat @RoamingNorway Yeah they mostly do on bigger locations though. On the 8 stall ones for example I've not seen that yet and those suffer the most from slow charging speed due to only having 640kW in total for all stalls. I guess it's too expensive with own substation for those locations.
English
0
0
1
36
ma4t
ma4t@miquelaymat·
@Seb_Gode @RoamingNorway I meant AC input as well. It doesn't necessarily have to be 400 V in Europe. When Tesla installs their own substation, they can choose to output 480 V on the LV side. And in some cases, they actually do.
English
1
0
0
25
William Jarbeaux
William Jarbeaux@RoamingNorway·
Tesla’s «250 kW» charging stations (basically only when you’re alone there)
William Jarbeaux tweet media
English
6
3
58
5.8K