@VceOfReason Staat „raubt“ dir 79€? Stell dir vor, du hättest sie noch: Du würdest kettenrauchend, völlig besoffen und mit vollem Tank als rollende Brandbombe durchs Land pflügen.
Du hast 100€.
Dir bleiben nach Steuern und Abgaben durchschnittlich 47€.
Du kaufst dir eine Schachtel Zigaretten für 9€.
Davon gehen 6€ an den Staat.
Bleiben 38€.
Du kaufst anschließend eine Flasche Sekt für 12€.
Davon gehen ca. 5€ an den Staat.
Bleiben 26€.
Du kaufst von diesem Geld etwa 13 Liter Benzin.
15€ gehen an den Staat.
In diesem Szenario gehen 21€ an dich und 79€ an den Staat.
Es ist nichts anderes als Raub.
Schlimmer als die Mafia.
@spiritbuun this seems really huge improvement ,i tested it myself with latest build but i only get 45-50t/s ,on rtx 3090 windows 11 with these args jpst.it/4_8dE what should i change @spiritbuun
Now that the dust has settled, is DFlash real or hype?
Today I was able to break the 200 tok/s barrier with 27B on a single RTX 3090:
run 1: 194.6 t/s, accept=155/180
run 2: 205.1 t/s, accept=155/180
run 3: 206.3 t/s, accept=155/180
But what about real world usage? Well,
@spiritbuun Nice! My acceptance rate hovers around 0.1 . Any ideas? here is how I run it: buun-llama-cpp/build/bin/Release/llama-server.exe -m Qwen3.6-27B-Q4_K_M.gguf -md dflash-draft-3.6-q8_0.gguf --spec-type dflash -c 196608 -cd 16536 --cache-type-k turbo3_tcq --cache-type-v turbo3_tcq
The new Qwen3.6-27B now runs on Luce DFlash. Up to 2x throughput on a single RTX 3090.
Qwen3.6-27B ships the same Qwen35 architecture string and identical layer/head dims as 3.5, so the existing DFlash draft + DDTree stack loads it as-is.
Throughput is lower than on 3.5. Looking forward for the updated version from the DFlash team to implement it as well!
Repo in the first comment ⬇️
How many views do you think the Tucker X Putin interview will get in the first 24hrs?
Place your guesses in the comment section.
Winner gets a Truth Army hoodie.
@PhoenixWallet@SatCastle21 Splicing funds in seems to require 2 on-chain transactions (the original transaction from another wallet and the actual splice-in). How is this better than doing a swap-in (in case there is enough incoming liquidity) ? Could that be improved ?
@SatCastle21 Suppose that your wallet is a bucket, and your balance is the water in the bucket.
Receiving = adding more water
Spending = pouring water out
Any time the bucket needs to be resized to allow for more water, that's an on-chain operation.