Post

Boris Cherny
Boris Cherny@bcherny·
Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully. And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon. #page=73" target="_blank" rel="nofollow noopener">www-cdn.anthropic.com/c5fbac3f0b1280…
Boris Cherny tweet media
Claude@claudeai

On several coding and knowledge work evaluations, Opus 5 is the new state-of-the-art:

English
347
476
6.2K
659.8K
OneHop
OneHop@OneHopAI·
@bcherny Prompt-injection resistance deserves more attention than leaderboard headlines. A model that holds up in messy tool-using workflows is what lets teams expand autonomy responsibly. We care about that practical bar at OneHop.
English
0
0
0
8
Paylaş