
while everyone's fine-tuning the same 3 open-source architectures, this Korean lab said "nah" and built a 314 BILLION parameter model completely from scratch.
Motif-3-Beta:
🧠 ~314B total params, only ~13B active per token (MoE sparsity go brrr)
📏 256K context, natively — not "extended," not duct-taped
⚡ 384 experts, 8 fire per token + 1 shared expert always on
🌐 fully multilingual
and they invented their OWN attention mechanism (Grouped Differential Latent Attention) and their OWN activation function (Grouped PolyNorm) because apparently reusing existing architectures was too easy
weights are fully open, no gatekeeping, no access request
this is still the BETA. final version drops soon. imagine what that looks like
huggingface.co/Motif-Technolo…
English













