Escha Labs
93 posts

Escha Labs
@Eschalabs
Minimum bits, maximum intelligence.
SF Katılım Haziran 2026
69 Takip Edilen896 Takipçiler

Today we are introducing Escha-W2 quantization.
A 2-bit Qwen3.6-35B-A3B model built for fast, local inference. The complete model is 12.3GB on disk—small, enough to run on a single consumer GPU — while averaging ~100% of FP8 performance across 12 benchmarks, including:
MMLU-Pro: 80.9
MATH-500: 93.8
GPQA-Diamond: 77.8
LiveCodeBench v6: 62.6
BFCL tool use: 88.9
RULER 8K–128K: 89.9
Commonsense-6: 76.1
On a single RTX 4090, the model runs:
225 tok/s single-stream generation
on 12.3GB on-disk model size
Compressing a 35B MoE model this far without collapsing its capabilities required more than a standard quantization pass.
We built an end-to-end compression system combining state-of-the-art low-bit quantization with model-aware fine tuning and recovery to preserve capabilities most vulnerable to low-bit error.
Quantizing the Qwen 3.6 35B model - from the base model to the final deployable checkpoint — took approximately 10 hours to complete.
Escha-W2 runs through a custom Qwen3-MoE runtime, which includes the weight loader, low-bit decoding kernels and serving integration required to execute the Escha format efficiently. The runtime currently supports SGLang and ZML deployment. vLLM and llama.cpp will be supported in the next release.
No retraining from scratch. No specialized accelerator. One consumer GPU.
Model download: huggingface.co/EschaLabs/Qwen…
Runtime download: huggingface.co/EschaLabs/esch…
Apache-2.0 model and runtime.
English

@Kajutek2006 @hara_shimpei Streaming weights from external memory is workable if you can manage the technicals of it and are ok with the added latency.
English

@Eschalabs @hara_shimpei can i run it wit 8gb vram and 48gb ram?
Or is it GPU vram only due to having its own runtime?
English

Our friends @hara_shimpei and the team at AI& independently evaluated Escha's 2bit Qwen 3.6 35B on Terminal-Bench 2.1, a challenging benchmark for agentic coding, and kindly agreed to share their findings.
AI& will also test the model at longer context lengths, including 128K+, where quantization quality becomes more difficult to preserve. We’ll share those results when they are available.
We’re grateful to AI& for the rigorous independent evaluation - and welcome others to share their findings!


English

@DavidBennett__ Thank you! This is version one, two will be even better!
English

Exciting results too!! @Eschalabs 👏🏻 👏🏻
Escha Labs@Eschalabs
Our friends @hara_shimpei and the team at AI& independently evaluated Escha's 2bit Qwen 3.6 35B on Terminal-Bench 2.1, a challenging benchmark for agentic coding, and kindly agreed to share their findings. AI& will also test the model at longer context lengths, including 128K+, where quantization quality becomes more difficult to preserve. We’ll share those results when they are available. We’re grateful to AI& for the rigorous independent evaluation - and welcome others to share their findings!
English

@blankcardboard Thank you. Technically it is not "lossless" but that is a goal we are working toward to.
English

@Eschalabs First quantization with zero loss is very impressive.
English

@SamuelZengML @huggingface Congratulations, will explore quantizing this!
English

Open-source ASR just took the top spot.
Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a 4.76 Mean WER.
Our 0.6B model also ranks in the Top 5.
Two model sizes. One open speech stack. Thread ↓

English

@thinkymachines Congratulations! Will explore getting this quantized on our pipeline for the community!
English

Today, we are releasing Inkling-Small.
Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.
thinkingmachines.ai/news/inkling-s…
Fine-tune it on Tinker today, or chat with it in text, image, and audio on Tinker Playground.
English

@Eschalabs Phenomenal work! The future of edge computing just got a bit brighter.
English

@_yorunoken @TsjSlgy Depends on your hardware and setup. We publish the setup that got us the toks/s on the hf page.
English

@LocalAiCherry We are new, please give us a look too:
huggingface.co/EschaLabs/Qwen…
English

Here’s the latest open source local drops:
• Laguna-S2.1 (118B MoE / 8B active + 1M context ,agentic beast)
• Nanbeige 4.2 3B (clever looped layers, super efficient)
• Solar-Open2 250B
• Bonsai-27B / Ternary-Bonsai
• Qwen3.6-27B + BTL-3 LoRA
• OpenZero Gemma4-E4B & Qwen3-8B
• Kimi K3 (2.8T open weights just dropped ,massive but the weights are finally out)
When the model lives on your hardware, the data stays yours and the bill stays zero.
Mutual minds, who’s already testing or shipping agents on any of these?
English

@NuminousVoidX @AstraiaAI We plan to release MLX first, then llama.cpp soon after, hopefully all within the next 2-3 weeks
English

@Eschalabs @AstraiaAI Any ETA for the gguf? Non Nvidia user here 😬
English

Thank you! The pipeline is essentially PTQ-based, it's multiple stage that performs a series of optimizations in a mathematically efficient manner. And since we also get to write our own kernels, there is room to optimize for inference further.
Please let us know if the results on real tasks (or benchmarks) are better/worse than you expect and share with the community the good and the bad for us to improve!
English

@Eschalabs @jetha @miguelsaavedra I tested this with an RTX 5070Ti, and the speed was quite good. I'm very curious about the quantization method!
English

@Chinese_XU @jkyamog Yes, we plan to release bigger models and sub 2bit ones too :)
English

@Eschalabs @jkyamog Perhaps in the future, we can focus on those models in the 100-300B range. Many dgx sparks with 128gb and RTX pro 6000 96gb will benefit from your technology
eg, deepseek v4 flash
English

@TsjSlgy Yes, to run. Still need vram to hold all of the model weights hot (i.e. 35B). Unless you plan to stream the weights from external memory.
English

@Eschalabs Qwen3.6 A3B 35B is already an MoE model. On my RTX 3060 12GB, it only takes about 3-4GB VRAM to run. So why not use that instead?
English

@Eschalabs Will your techniques work with Qwen3.6-27B?! That would be incredible!
English





