Kunhee Kim (@kunhee_k) - Twitter Profili | Zamantika Mersobahis Locabet

To fix this, Directional Textual Inversion (DTI) constrains the norm and optimizes on the hypersphere via Riemannian SGD. Cast as MAP with a vMF prior, DTI drastically improves text fidelity and unlocks smooth conceptual interpolation (slerp).

English

Kunhee Kim@kunhee_k·3d

Why does standard TI fail? We trace it to embedding norm inflation. In pre-norm Transformers, massive magnitudes cause Positional Attenuation (washing out contextual prompts) and Residual Stagnation (freezing token direction as residual updates become negligible).

English

125

Kunhee Kim@kunhee_k·3d

Excited to present Directional Textual Inversion (DTI) at #ICLR 2026! Textual Inversion (TI) often fails on complex prompts. We fix this by optimizing strictly on the hypersphere. Work by our team at @KAIST_AI . arXiv: arxiv.org/abs/2512.13672 Project: kunheek.github.io/dti

English

2.6K

Kunhee Kim retweetledi

AK@_akhaliq·13 Eyl

TextBoost Towards One-Shot Personalization of Text-to-Image Models via Fine-tuning Text Encoder discuss: huggingface.co/papers/2409.08… Recent breakthroughs in text-to-image models have opened up promising research avenues in personalized image generation, enabling users to create diverse images of a specific subject using natural language prompts. However, existing methods often suffer from performance degradation when given only a single reference image. They tend to overfit the input, producing highly similar outputs regardless of the text prompt. This paper addresses the challenge of one-shot personalization by mitigating overfitting, enabling the creation of controllable images through text prompts. Specifically, we propose a selective fine-tuning strategy that focuses on the text encoder. Furthermore, we introduce three key techniques to enhance personalization performance: (1) augmentation tokens to encourage feature disentanglement and alleviate overfitting, (2) a knowledge-preservation loss to reduce language drift and promote generalizability across diverse prompts, and (3) SNR-weighted sampling for efficient training. Extensive experiments demonstrate that our approach efficiently generates high-quality, diverse images using only a single reference image while significantly reducing memory and storage requirements.

English

117

15.1K

Kunhee Kim

Keşfet