Kyle Kastner

9.5K posts

Kyle Kastner

Kyle Kastner

@kastnerkyle

computers and music are fun

Out of the city, Massachusetts Katılım Ocak 2011
3.3K Takip Edilen3.3K Takipçiler
Kyle Kastner retweetledi
Chuning Zhu
Chuning Zhu@chuning_zhu·
Can we get robots to think before they act? In Latent Memory Palace, we achieve this through a classic principle: variational inference. (1/8)
English
4
34
148
28K
Kyle Kastner retweetledi
Kyle Kastner retweetledi
Bhavin Jawade
Bhavin Jawade@BhavinJawade·
𝗢𝗻-𝗽𝗼𝗹𝗶𝗰𝘆 𝗱𝗶𝘀𝘁𝗶𝗹𝗹𝗮𝘁𝗶𝗼𝗻 𝗶𝘀𝗻'𝘁 𝗮 𝗳𝗿𝗲𝗲-𝗹𝘂𝗻𝗰𝗵 On-policy distillation has become a default post-training tool in many open-source frontier model training recipes. Recent releases lean on it heavily: DeepSeek v4, MiMO, and Nemotron-Cascade-2 use MOPD, and GLM 5.x uses on-policy cross-stage self-distillation. It provides RL's on-policy nature reducing exposure bias, while providing token level supervision like SFT. But OPD and OPSD have their own failure modes. In this post I discuss a few of them: 1. 𝗘𝗮𝗿𝗹𝘆 𝗺𝗶𝘀𝘁𝗮𝗸𝗲𝘀 𝗮𝗿𝗲 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗮𝗹𝗹𝘆 𝘂𝗻𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗮𝗯𝗹𝗲. When the student samples a rollout and takes an early wrong turn, the per-token KL computed along that frozen rollout cannot pull it back onto a correct path. TRD proves that this failure is built into the objective rather than being a matter of noisy gradients. Even with a perfect teacher, the gradient obtained from token-level KL on the student's own rollout agrees with the ideal corrective gradient at exactly one point, the token where the student first diverged, and disagrees everywhere after it. Every later supervision target is therefore anchored to a context that the student should never have entered. Because reweighting or clipping only rescales the magnitude of each token's gradient, and here the terms point in the wrong direction, no per-token adjustment can recover the correct update. TRD's proposed fix is to distill along a teacher-refined trajectory rather than the raw student rollout, which restores a target the student can actually follow. 2. 𝗔 𝘀𝘁𝗿𝗼𝗻𝗴𝗲𝗿 𝘁𝗲𝗮𝗰𝗵𝗲𝗿 𝗰𝗮𝗻 𝗯𝗲 𝗮 𝘄𝗼𝗿𝘀𝗲 𝘁𝗲𝗮𝗰𝗵𝗲𝗿. On-policy distillation can only teach the student at states the student itself visits, and the usable signal at each of those states lives in the overlap between the student's and teacher's next-token distributions. Rethinking OPD shows that a bigger, higher-scoring teacher can fail to move a student while a weaker one succeeds, because if the teacher's token distribution places its mass on tokens the student rarely produces, the overlap is small and almost nothing transfers, no matter how capable the teacher is in absolute terms. What actually predicts success is early top-k thinking-pattern overlap. In runs that work, the shared top-k tokens carry 97 to 99% of the probability mass and the overlap ratio climbs steadily during training, whereas a run that starts with low overlap never recovers it. A teacher trained on the same recipe as the student also converges toward the student's own distribution, so its higher benchmark score does not correspond to any new knowledge it can transfer. The practical rule is to pick teachers by distributional closeness to the student, not by leaderboard rank. 3. 𝗣𝗿𝗶𝘃𝗶𝗹𝗲𝗴𝗲𝗱-𝗶𝗻𝗳𝗼𝗿𝗺𝗮𝘁𝗶𝗼𝗻-𝗰𝗼𝗻𝗱𝗶𝘁𝗶𝗼𝗻𝗲𝗱 𝗢𝗣𝗦𝗗 𝗰𝗮𝗻 𝗳𝗮𝗶𝗹 𝘁𝗼 𝘁𝗿𝗮𝗻𝘀𝗳𝗲𝗿. In OPSD you distill a teacher that was conditioned on privileged information, such as the gold answer, into a student that will never have it. The Many Faces of OPD shows what goes wrong when that information is instance-specific. The student cannot recover the teacher's per-instance reasoning, since it never sees the answer, so it instead learns a single answer-free policy that effectively averages the teacher's behavior across all problems, and that averaged policy is too generic to solve any particular one. The signature is initial gains followed by collapse: rollouts grow long, fill with hedging tokens, and accuracy craters toward zero. The approach works only when the privileged information is a shared rule that applies across all instances, such as a system prompt or an alignment preference, and not when it is a per-problem answer. 4. 𝗧𝗵𝗶𝗻𝗸𝗶𝗻𝗴 𝗰𝗼𝗹𝗹𝗮𝗽𝘀𝗲: 𝗱𝗲𝗻𝘀𝗲 𝘀𝘂𝗽𝗲𝗿𝘃𝗶𝘀𝗶𝗼𝗻 𝘀𝘂𝗽𝗽𝗿𝗲𝘀𝘀𝗲𝘀 𝘁𝗵𝗲 𝗺𝗼𝗱𝗲𝗹'𝘀 𝗼𝘄𝗻 𝗱𝗲𝗹𝗶𝗯𝗲𝗿𝗮𝘁𝗶𝗼𝗻. A teacher conditioned on the answer has no reason to hesitate, backtrack, or explore, so its per-token targets quietly push down the student's deliberation tokens. Diagnosing and Mitigating Thinking Collapse names this phenomenon thinking collapse: over training, the student's native reasoning behavior erodes as the exploratory tokens that carry it, words like wait, maybe, and alternatively, become progressively less frequent. The mechanism is local rather than global. The damage concentrates at high-entropy decision forks, the branch points where the student is genuinely uncertain and would normally deliberate. Exactly there, the student's top-1 token is often an exploratory marker while the answer-conditioned teacher's top-1 token is not, so the mismatch produces a strong gradient that suppresses the very tokens that make reasoning work. The result is a model whose native reasoning behavior is measurably suppressed, and downstream reasoning accuracy falls in step with it.
Bhavin Jawade tweet media
English
8
57
469
34.7K
Kyle Kastner retweetledi
Matt Henderson
Matt Henderson@matthen2·
does your audio-input LLM secretly know how to generate speech? yes but it sounds like a scary demon! deepdream but for speech: starting with noise audio, I tried optimizing the probability of a desired transcription, while keeping the model frozen. volume up! 🎧
English
50
84
1.2K
82.3K
Kyle Kastner retweetledi
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
Register tokens were introduced to fix artifacts in Vision Transformers (ViTs) (arxiv.org/abs/2309.16588) In Diffusion Transformers (DiTs), we don't observe such outliers, so registers shouldn't help. But they do... And especially for pixel-space DiTs. Simply adding a few empty tokens can improve FID from 3.52 to 2.69.
Nikita Starodubcev tweet media
English
5
23
132
8.4K
Kyle Kastner retweetledi
Axel Darmouni
Axel Darmouni@ADarmouni·
arxiv.org/pdf/2607.12395 Wonderful feat from Ant Group and Gaoling School of Artifical Intelligence, which managed to scale RL up to a 1T model What they do is a 4 sequence training RL pipeline: 1- Token-based RL using clipped importance sampling policy gradient: the importance sampling is computed from current policy vs rollout policy, and high clip is applied but not low clipped ; with a KL divergence penalty. Worth noting they do token-level in purpose, not normalizing in output length to make sure that the base model actually goes through longer chains of thoughts. 2- Most surprising part here is that yes, they use self-distillation to now teach the model shorter ways of solving the problems it had to face in RL step 1. It’s not however OPSD: what they do is take the shortest traces among multiple rollouts, ask the model to simplify them by removing circular logic, and then perform SFT on said traces. 3- Then, they do sample-level RL, this time with the output length normalization and without the KL. 4- Afterwards, a last RL step is being done, splitting the corpus into 3 difficulties and adapting the model to adapt its thought processes to the problem at hand (ie creating reasoning low, mid and high modes). Additional notes: - Reward is format + correct output. If the problem is more complex, an LLM as a judge is used. Format is - Main model in BF16, but attention softmax and lmhead in FP32 strictly to eliminate training inference mismatch. - Context parallelism is optimized with respect to their hybrid architecture. All the setup+hyperparameters of the experience is noted within the work. Model is competitive with SoTA on all math benchmarks without having been SFT. Funnily enough, they also test distilling from this model over 100k samples and achieve better results on Qwen-2.5 and Llama-3.3 than Deepseek-R1 did with 800k. They also do a lot of ablations to validate their training choices, and also perform their process on their Flash model (10 times less parameters) for worse but still good results. Would love to see how the model performs on non-math reasoning benchmarks, but it’s already quite a feat that they made online RL work in a 1T model!
Axel Darmouni tweet mediaAxel Darmouni tweet media
English
1
12
71
4.9K
Kyle Kastner retweetledi
Xiuyu Li
Xiuyu Li@sheriyuo·
OPD normally copies everything the teacher prefers, including habits inherited before reasoning tuning. OPD² instead uses the token-level logprob change between a reasoning teacher and its own base checkpoint, so the reward isolates what post-training actually added. On-Policy Delta Distillation Paper: arxiv.org/abs/2607.15161
Xiuyu Li tweet media
English
8
32
222
41K
Kyle Kastner retweetledi
Soumith Chintala
Soumith Chintala@soumithchintala·
Excited for our first general model Inkling -- open weights, 975B, natively multimodal (text, image, audio). Available on Tinker, HuggingFace and partners. It is yours to personalize and use openly. It is yours.
Thinking Machines@thinkymachines

Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

English
66
149
2.9K
236.5K
Kyle Kastner retweetledi
Thinking Machines
Thinking Machines@thinkymachines·
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
English
565
2K
14.9K
8M
Kyle Kastner retweetledi
Tinker
Tinker@tinkerapi·
Inkling is our first open model from @thinkymachines and is now available on Tinker! Check out these quotes from Tinker customers on their experience with Inkling: @_Mantic_AI: "Not only does Inkling outperform Kimi K2.6 on our forecasting evals, it does so with half the output tokens." @trajectorylabs: "We’ve been impressed by how sharp and efficient the model is. Its reasoning is concise, its tool calling is consistently strong, and it holds up well on complex, long-horizon agentic tasks. It feels like a meaningful unlock for what teams can build with open-source models designed for customization." @lightningrodai: "We came away impressed by the model’s underlying reasoning ability. It’s thoughtful, original, and refreshingly unsycophantic.”
Thinking Machines@thinkymachines

Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. thinkingmachines.ai/news/introduci… Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

English
9
14
81
11.3K
Kyle Kastner retweetledi
Huan-ang Gao
Huan-ang Gao@c7wc7w·
RLVR is powerful, but repeating it for every larger target model is expensive: each target must generate its own rollouts and rediscover useful learning signals from sparse outcome rewards. Can RL on a small, weaker model improve a stronger student—even when the student already outperforms the small model after RL? We found that it can—but not by distilling the weak model itself. Today, we’re excited to share Direct-OPD, joint work w/ @Shiyuan040223, @c7wc7w, @Ahydchh, @Han_lin_Wu, @zhilong_zhang26, Zheng Jiang, @HBX_hbx, Wei-Ying Ma, @yaqinzhang, @haozhou_ai, developed at SIA-Lab @hello_gensi, a joint lab of Tsinghua AIR and ByteDance Seed. bytedtsinghua-sia.github.io/Direct-OPD/ Our alternative is simple: 1. Run RL on a small model, where exploration and rollouts are cheaper. 2. Treat the model’s pre- and post-RL checkpoints as a teacher pair, whose difference captures the direction learned through RL. 3. On-policy distill this policy shift—what RL changed—using the stronger student’s own rollouts. In one setting on AIME24: [*] 1.5B teacher pair: pre-RL and post-RL checkpoints (Post-RL teacher score: 51.3) [*] 7B student before transfer: 56.7 [*] 7B student + vanilla on-policy distillation: ~50 [*] 7B student + Direct-OPD: 63.1 (+6.4) The 7B student already starts stronger than the post-RL teacher. Distilling the teacher itself makes the student worse. But distilling what the teacher pair learned through RL improves it further. In other words, the reusable outcome of an RL run is not only the final checkpoint—it can also be the policy shift encoded by the pre- and post-RL checkpoint pair.
English
2
6
25
4.1K
Kyle Kastner retweetledi
Kyle Kastner retweetledi
Amiri Hayes
Amiri Hayes@amirihayes_·
What if attention were code? We show that many attention heads in transformer LMs can be replaced by human-readable Python programs. Swap them in and the model barely notices. See our experiments here: Explaining Attention with Program Synthesis [arxiv.org/abs/2606.19317]
English
28
114
994
159.3K
Kyle Kastner retweetledi
Nathan Barry
Nathan Barry@nathanrs·
New post! A drawback of diffusion LLMs is that bidirectional attention causes keys and values to drift across steps, breaking KV caching. But generation quality is robust to slight KV drift, and much work has gone into maximizing stale KV reuse without causing quality drops (1/3)
English
1
20
169
10.3K
Kyle Kastner retweetledi
Chelsea Finn
Chelsea Finn@chelseabfinn·
Can we translate a rough sense for what to do + VLA prior into successful behavior? Flow reversal steering: 1) runs the VLA's flow ODE backwards to back out noise that's closest to coarse traj 2) runs flow forwards to get closest good behavior Paper: flow-reversal-steering.github.io
Andy Tang@tangerinecoder

Generalist robot policies learn many useful skills, but struggle to select good behaviors for new tasks. To solve this, we introduce Flow Reversal Steering (FRS), a method to refine coarse semantic guidance into precise, in-distribution motions. flow-reversal-steering.github.io 1/N

English
6
26
221
31.9K
Kyle Kastner retweetledi
Nando de Freitas
Nando de Freitas@NandoDF·
A model trained for next-token prediction is forced to build compressed representations of latent structure in text. Ilya Sutskever correctly refers to this phenomenon as understanding. Here, a model trained for next-step sensor prediction, with a robot that has proprioception and touch sensors but no vision, is forced to build compressed representations of latent structure in the physical world. The robot becomes aware of the shape of external objects. That is, it understands the physical properties of the external world that enable it to make better next-step sensor predictions. This research was previously done by a diverse team of expert engineers at DeepMind over a month - including stars like @notmisha and @yuvaltassa. Remarkably, this reproduction with a completely different robot took only a few hours to implement using Codex. The automatic creation of physical environments by AI will likely lead to huge advances in areas of science and engineering that use physical simulators or twin models. The paper and notebook are available at ❤️∀ love4all.ai
English
16
28
266
44K
Kyle Kastner retweetledi
Yuchen Jin
Yuchen Jin@Yuchenj_UW·
DeepSeek is the GOAT. 🐳 They just published DSpark, a new speculative decoding method that boosts throughput by 51% to 400%. They also open-sourced DeepSpec, the training framework behind it. This is the real open AI.
Yuchen Jin tweet media
English
104
456
3.7K
369.8K