
Backtracking Improves Generation Safety
Introduces a method to improve adversarial robustness by teaching the model to use a special reset token on unsafe generations and regenerate a response from scratch.
arxiv.org/abs/2409.14586

English
ML Safety Daily
378 posts

@topofmlsafety
ML safety papers as they are released. Course: https://t.co/l0e0Y2i3AU Newsletter: https://t.co/8Y1kh2D7K6 Main Twitter: https://t.co/AXoYPryldd











































