Escha Labs

93 posts

Escha Labs banner
Escha Labs

Escha Labs

@Eschalabs

Minimum bits, maximum intelligence.

SF Katılım Haziran 2026
69 Takip Edilen896 Takipçiler
Escha Labs
Escha Labs@Eschalabs·
@sam_defi Please do and share back what you like and don’t like.
English
1
0
1
35
Escha Labs
Escha Labs@Eschalabs·
Today we are introducing Escha-W2 quantization. A 2-bit Qwen3.6-35B-A3B model built for fast, local inference. The complete model is 12.3GB on disk—small, enough to run on a single consumer GPU — while averaging ~100% of FP8 performance across 12 benchmarks, including: MMLU-Pro: 80.9 MATH-500: 93.8 GPQA-Diamond: 77.8 LiveCodeBench v6: 62.6 BFCL tool use: 88.9 RULER 8K–128K: 89.9 Commonsense-6: 76.1 On a single RTX 4090, the model runs: 225 tok/s single-stream generation on 12.3GB on-disk model size Compressing a 35B MoE model this far without collapsing its capabilities required more than a standard quantization pass. We built an end-to-end compression system combining state-of-the-art low-bit quantization with model-aware fine tuning and recovery to preserve capabilities most vulnerable to low-bit error. Quantizing the Qwen 3.6 35B model - from the base model to the final deployable checkpoint — took approximately 10 hours to complete. Escha-W2 runs through a custom Qwen3-MoE runtime, which includes the weight loader, low-bit decoding kernels and serving integration required to execute the Escha format efficiently. The runtime currently supports SGLang and ZML deployment. vLLM and llama.cpp will be supported in the next release. No retraining from scratch. No specialized accelerator. One consumer GPU. Model download: huggingface.co/EschaLabs/Qwen… Runtime download: huggingface.co/EschaLabs/esch… Apache-2.0 model and runtime.
English
82
122
1.3K
170.4K
Escha Labs
Escha Labs@Eschalabs·
@Kajutek2006 @hara_shimpei Streaming weights from external memory is workable if you can manage the technicals of it and are ok with the added latency.
English
1
0
1
102
Azzren
Azzren@Kajutek2006·
@Eschalabs @hara_shimpei can i run it wit 8gb vram and 48gb ram? Or is it GPU vram only due to having its own runtime?
English
2
0
1
123
Escha Labs
Escha Labs@Eschalabs·
Our friends @hara_shimpei and the team at AI& independently evaluated Escha's 2bit Qwen 3.6 35B on Terminal-Bench 2.1, a challenging benchmark for agentic coding, and kindly agreed to share their findings. AI& will also test the model at longer context lengths, including 128K+, where quantization quality becomes more difficult to preserve. We’ll share those results when they are available. We’re grateful to AI& for the rigorous independent evaluation - and welcome others to share their findings!
Escha Labs tweet mediaEscha Labs tweet media
English
4
5
47
3.1K
David Bennett
David Bennett@DavidBennett__·
Exciting results too!! @Eschalabs 👏🏻 👏🏻
Escha Labs@Eschalabs

Our friends @hara_shimpei and the team at AI& independently evaluated Escha's 2bit Qwen 3.6 35B on Terminal-Bench 2.1, a challenging benchmark for agentic coding, and kindly agreed to share their findings. AI& will also test the model at longer context lengths, including 128K+, where quantization quality becomes more difficult to preserve. We’ll share those results when they are available. We’re grateful to AI& for the rigorous independent evaluation - and welcome others to share their findings!

English
1
0
5
604
the tiny corp
the tiny corp@__tinygrad__·
We have a 120 tok/s GLM-5.2 and a 42 tok/s Kimi K3 running locally on AMD boxes and it's legit a frontier setup for ~$600k. Will be interesting to see how that dollar amount varies over time.
English
69
67
1.9K
131K
Escha Labs
Escha Labs@Eschalabs·
@blankcardboard Thank you. Technically it is not "lossless" but that is a goal we are working toward to.
English
0
0
0
15
Samuel Zeng
Samuel Zeng@SamuelZengML·
Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a 4.76 Mean WER. Our 0.6B model also ranks in the Top 5. Two model sizes. One open speech stack. Thread ↓
Samuel Zeng tweet media
English
11
27
198
31.3K
Escha Labs
Escha Labs@Eschalabs·
@thinkymachines Congratulations! Will explore getting this quantized on our pipeline for the community!
English
0
0
2
232
Thinking Machines
Thinking Machines@thinkymachines·
Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available. thinkingmachines.ai/news/inkling-s… Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
English
133
405
2.9K
1.2M
kavanmce
kavanmce@kavanmce·
@Eschalabs Phenomenal work! The future of edge computing just got a bit brighter.
English
1
0
1
104
Escha Labs
Escha Labs@Eschalabs·
@_yorunoken @TsjSlgy Depends on your hardware and setup. We publish the setup that got us the toks/s on the hf page.
English
0
0
0
23
Local Ai Cherry
Local Ai Cherry@LocalAiCherry·
Here’s the latest open source local drops: • Laguna-S2.1 (118B MoE / 8B active + 1M context ,agentic beast) • Nanbeige 4.2 3B (clever looped layers, super efficient) • Solar-Open2 250B • Bonsai-27B / Ternary-Bonsai • Qwen3.6-27B + BTL-3 LoRA • OpenZero Gemma4-E4B & Qwen3-8B • Kimi K3 (2.8T open weights just dropped ,massive but the weights are finally out) When the model lives on your hardware, the data stays yours and the bill stays zero. Mutual minds, who’s already testing or shipping agents on any of these?
English
20
15
264
15.6K
Escha Labs
Escha Labs@Eschalabs·
Thank you! The pipeline is essentially PTQ-based, it's multiple stage that performs a series of optimizations in a mathematically efficient manner. And since we also get to write our own kernels, there is room to optimize for inference further. Please let us know if the results on real tasks (or benchmarks) are better/worse than you expect and share with the community the good and the bad for us to improve!
English
0
0
2
11
君子中庸
君子中庸@Chinese_XU·
@Eschalabs @jkyamog Perhaps in the future, we can focus on those models in the 100-300B range. Many dgx sparks with 128gb and RTX pro 6000 96gb will benefit from your technology eg, deepseek v4 flash
English
1
0
1
72
Escha Labs
Escha Labs@Eschalabs·
@TsjSlgy Yes, to run. Still need vram to hold all of the model weights hot (i.e. 35B). Unless you plan to stream the weights from external memory.
English
0
0
0
239
T.Tang
T.Tang@TsjSlgy·
@Eschalabs Qwen3.6 A3B 35B is already an MoE model. On my RTX 3060 12GB, it only takes about 3-4GB VRAM to run. So why not use that instead?
English
2
0
1
293