Nikita Starodubcev

13 posts

Nikita Starodubcev banner
Nikita Starodubcev

Nikita Starodubcev

@NikStarodub

research @Yandex, ex intern @Adobe

Saint Petersburg, Russia Katılım Haziran 2026
45 Takip Edilen38 Takipçiler
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
the visualization looks good, thanks for sharing. SD3 DOES have outliers, but they are located in the text sequence, not the image sequence, see the image below We're talking specifically about outliers in PATCH/IMAGE tokens here, same kind that show up in ViTs. And we didn't find outliers in patch tokens
Nikita Starodubcev tweet media
English
1
0
1
46
Jonas
Jonas@LoosJonas·
@NikStarodub in my experience DiTs (e.g. SD3) *can* have outliers, although much less pronounced than in ViTs usually.
English
1
0
1
60
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
Register tokens were introduced to fix artifacts in Vision Transformers (ViTs) (arxiv.org/abs/2309.16588) In Diffusion Transformers (DiTs), we don't observe such outliers, so registers shouldn't help. But they do... And especially for pixel-space DiTs. Simply adding a few empty tokens can improve FID from 3.52 to 2.69.
Nikita Starodubcev tweet media
English
5
20
132
8.5K
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
Sounds good, thx for sharing, would be happy to discuss more. Where are those outliers located? One thing worth flagging: here we specifically analyze IMAGE diffusion models, which are architecturally and training-wise closer to ViTs. Outliers in video DMs are actually pretty well explored already and your world-model setup sounds closer to that regime than to ours. We cover this in the related work.
English
0
0
1
105
rama
rama@RamaAdrien·
@NikStarodub "in DiTs we don't observe such outliers" --> in our recent @kyutai_labs x @gen_intuition collaboration on world modeling (MIRA), we observed such outliers in DiTs and registers helped.Gated attention also helped. (IIRC others have also observed that on other DiTs)
English
1
0
1
140
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
Yup, saw that one First, the authors mostly focus on the DINOv2 encoder WITHOUT registers, which already has outliers, so of course a diffusion model built on such a "corrupted" latent space inherits those outliers. That framing is impractical though since in practice we don't use DINOv2 without registers only WITH them. If you look at Figure 4 the difference between models with and without registers is significant in terms of outliers. In our work, we analyze the original case WITH registers. Second, they don't show per-token norm plots, and their verification is only visual, on a single image. Attention maps alone can be pretty noisy. These artifacts should really be tied to high norms in the image patches, not just attention maps (see Figure 2 in this paper: arxiv.org/pdf/2506.08010…). But the authors don't show that connection here
English
0
0
2
37
Yitong chen
Yitong chen@1997yrrr·
@NikStarodub RAE also is a class-conditional models, not MM-DiT. The author didn't show the JiT/SiT internal token nomr heatmap, but they indeed see the the improvement across them
Yitong chen tweet mediaYitong chen tweet media
English
3
0
0
42
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
@1997yrrr BUT once you append register tokens, high-norm outliers do emerge, but inside the registers themselves (Fig. 2 in the paper 👀). So the model seems to "want" outlier slots, it just can't afford them in patch tokens since they're all tied to the loss
English
0
0
0
83
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
@1997yrrr our setup is different: class-conditional models where the transformer only processes image tokens, just like a ViT. And there we consistently find no outliers among patch tokens.
English
1
0
1
81
Nikita Starodubcev
Nikita Starodubcev@NikStarodub·
We put this to use. Registers improve image quality and add a useful signal, so why not amplify it? Register Guidance uses the model's prediction without registers as the negative direction, just as CFG uses the unconditional prediction. It combines naturally with CFG and consistently sharpens details and structure
Nikita Starodubcev tweet media
English
1
0
3
693