penlu
692 posts

penlu
@penlume
human computer interface. naturally occurring feature of your environment



We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria that reflect user experience. One such verifier is a scope criterion, which penalizes modifications to the codebase beyond what the task requires. We observe this effect across all frontier models we evaluated, though it is most pronounced in Opus 5: at higher reasoning efforts, the model shows a stronger tendency to refactor code unprompted.


I asked people which companies have the highest density of talented people they know: 1st. Cognition - 9 votes Equal 1st. Anthropic - 9 votes 2nd. Modal - 5 votes 3rd. OpenAI - 4 votes 4th. Standard Intelligence - 3 votes 4th. Cursor - 3 votes Two votes each: - Ramp - Flapping Airplanes - DeepMind - Long Lake - Applied Compute One vote each: - SpaceXAI - SpaceX (treated separately, one vote was for the AI lab subsidiary and one was for the rocket team) - American Terawatt - Mechanize - Olix - Fluidstack - Chai Discovery - Sail Research - Etched - Core Automation - Specter - Clay - Applied Intuition - Sierra - Hivemind - Bitrig - Retro - Thinking Machines - Decagon - Precigenetics - Pangram - Reflect - Thrive Holdings - Adaption

Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale

Introducing SWE-1.7, the most capable model we’ve trained yet. It scores within a few points of the strongest frontier models at a fraction of the cost, and is now available at 1000 tok/s. RL is not hitting its limit: after refining our recipe, we keep seeing gains as we scale


















