T.Tang retweetledi
T.Tang
668 posts


@Eschalabs Qwen3.6 A3B 35B is already an MoE model. On my RTX 3060 12GB, it only takes about 3-4GB VRAM to run. So why not use that instead?
English

Today we are introducing Escha-W2 quantization.
A 2-bit Qwen3.6-35B-A3B model built for fast, local inference. The complete model is 12.3GB on disk—small, enough to run on a single consumer GPU — while averaging ~100% of FP8 performance across 12 benchmarks, including:
MMLU-Pro: 80.9
MATH-500: 93.8
GPQA-Diamond: 77.8
LiveCodeBench v6: 62.6
BFCL tool use: 88.9
RULER 8K–128K: 89.9
Commonsense-6: 76.1
On a single RTX 4090, the model runs:
225 tok/s single-stream generation
on 12.3GB on-disk model size
Compressing a 35B MoE model this far without collapsing its capabilities required more than a standard quantization pass.
We built an end-to-end compression system combining state-of-the-art low-bit quantization with model-aware fine tuning and recovery to preserve capabilities most vulnerable to low-bit error.
Quantizing the Qwen 3.6 35B model - from the base model to the final deployable checkpoint — took approximately 10 hours to complete.
Escha-W2 runs through a custom Qwen3-MoE runtime, which includes the weight loader, low-bit decoding kernels and serving integration required to execute the Escha format efficiently. The runtime currently supports SGLang and ZML deployment. vLLM and llama.cpp will be supported in the next release.
No retraining from scratch. No specialized accelerator. One consumer GPU.
Model download: huggingface.co/EschaLabs/Qwen…
Runtime download: huggingface.co/EschaLabs/esch…
Apache-2.0 model and runtime.
English
T.Tang retweetledi
T.Tang retweetledi

As usual, Weibo is flooded with the same gloating over the Kumamoto earthquake — “deserved,” “karma,” open celebration of Japanese suffering.
What these accounts never grasp is the simplest law of nature:
What you throw out often comes back harder.
The “karma” they so eagerly wish upon their neighbour often circle back as a boomerang.
Controversial question:
When the next major disaster hits China, will these same people still call it “karma” — or will they suddenly demand the world show them the mercy they currently refuse to show anyone else?
Chinese netizens keep wishing ill on Japan.
Those wishes have a habit of returning.




English

▼概要
画像生成モデル「Krea 2」ベースの量子化派生モデル「krea 2 int4 convrot」のアップデート版「v1.0 turbo/raw」が公開。ComfyUIで動作する。
▼ポイント
・Krea 2 Raw/TurboをConvRot INT4(4bit)形式に量子化
・SM 8.0+(Ampere以降)のGPUを対象としRTX 3060でテスト済
・RawはLoRA学習や多様な出力、Turboは8ステップの高速推論に対応
▼使い所
LoRA学習や出力の多様性を求めるならRaw、高速生成ならTurboを選択。Rawで学習したLoRAをTurboに適用した生成も可能。
#krea2int4convrot #ComfyUI
URLはリプ⬇️

日本語

@ktsuki_novel それなら『足跡』がどのように結成され、メンバーを集めていったのかを書いてみるのはどうでしょうか?
例えばアークや他のギルドメンバーがなぜ加入したのか、なぜ脱退しなかったのか……など。
日本語
T.Tang retweetledi
T.Tang retweetledi
T.Tang retweetledi
T.Tang retweetledi

最近在看那种 AI 生成的 AV,非常有趣。
我觉得在 AI 视频生成这个赛道,最先取代的应该是 AV。
经常看 AV 的都知道,拍 AV 的演员基本演技都很差,不管是日本的还是中文的麻豆传媒,拍出来的片几乎都是一眼假,大部分只是展示性爱,没有太多的演技可言。
所以在演技这方面,AI 要超越电影电视剧的正经科班演员,可能还需要一段时间,但是在 AV 这一块,已经早早实现超越,并且遥遥领先。
而且 AI 生成的视频,没有任何题材限制,只要你脑子想得到,什么牛鬼蛇神能给你生成出来,这简直太棒了,因为人类的性癖几乎都是天马行空的,大部分现实是拍不出来的,就算能拍出来,也需要极高的拍摄成本。
但是对于 AI,你让它生成一个 3p 的 AV,和生成一个 10000p 的 AV,消耗的都是同等的 Token。
中文
T.Tang retweetledi

@RpsAgainstTrump Thanks for your enlightening comment. Now we know who else needs to be blocked.
English




















