Post

joseph
joseph@gitcommit90·
PSA: DGX Spark + Laguna S 2.1 NVFP4+DFlash is ~15–23 tok/s freeform. No it is not 77 tok/s. No it is not 50 tok/s everyday chat. No it is not a very fast model to run on the DGX Spark. Is it an amazing model? Remains to be seen. Am I also very excited that open weights keep landing next to frontier models? Absolutely. We can all feel excited. We should also be honest and not farm for clicks. If someone’s posting 5× that as “single session decode,” make them define the benchmark.
English
1
0
0
208
Eiso Kant
Eiso Kant@eisokant·
@gitcommit90 Our testing has us at 35-40 tok/s. Would be interesting to see what you’re using here and context length.
English
1
0
3
147
Paylaş