
Dev Chheda
541 posts

Dev Chheda
@devmchheda
building @cognition




We've received several questions about the Opus 5 FrontierCode results, where scores decline as reasoning effort increases. In fact, the behavior is expected under the benchmark design. FrontierCode evaluates merge-ability rather than correctness alone, incorporating criteria that reflect user experience. One such verifier is a scope criterion, which penalizes modifications to the codebase beyond what the task requires. We observe this effect across all frontier models we evaluated, though it is most pronounced in Opus 5: at higher reasoning efforts, the model shows a stronger tendency to refactor code unprompted.


Introducing Cursor Router, our intelligent model router that selects the right model for the task at hand. Router delivers frontier-quality results at 60% lower cost.











Introducing Devin Outposts: run Devin on any machine. Your Mac mini, a GPU box in your lab, a VM inside your private network, or a Kubernetes cluster next to your internal services.



Two years ago, we set out to make software self-driving. Today, TierZero is joining @Cognition to finish the job: make agents write code and keep it running. @yunpark93 and I couldn't be more excited for what's next! 🚀

Introducing ask-web: Rox’s in-house web search agent. ask-web sits on the cost-per-accuracy pareto frontier of the hyper-parameter grid when compared to frontier labs and commercial search agent providers. The agent delivers 91.3% accuracy at 1.03 cents per query on real production prompts. It has been running in production for more than 6 months with continuous evals. Inference partners: @togethercompute, @baseten, @modal Commercial Search vendors benchmarked: @perplexity_ai, @ExaAILabs, @p0. Frontier Search vendors benchmarked: @OpenAI, @AnthropicAI Exa, OpenAI and Anthropic excel on accuracy. Parallel and Perplexity are cost-efficient. Here’s the breakdown:


