Anna Mitchell
2.3K posts

Anna Mitchell
@annarmitchell
@cognition balancing exit and voice 🗣



I asked people which companies have the highest density of talented people they know: 1st. Cognition - 9 votes Equal 1st. Anthropic - 9 votes 2nd. Modal - 5 votes 3rd. OpenAI - 4 votes 4th. Standard Intelligence - 3 votes 4th. Cursor - 3 votes Two votes each: - Ramp - Flapping Airplanes - DeepMind - Long Lake - Applied Compute One vote each: - SpaceXAI - SpaceX (treated separately, one vote was for the AI lab subsidiary and one was for the rocket team) - American Terawatt - Mechanize - Olix - Fluidstack - Chai Discovery - Sail Research - Etched - Core Automation - Specter - Clay - Applied Intuition - Sierra - Hivemind - Bitrig - Retro - Thinking Machines - Decagon - Precigenetics - Pangram - Reflect - Thrive Holdings - Adaption

Much of the AI narrative revolves around architecture research, compute scaling & buying human data. Very few understand the deep technical discipline of "data research," which has driven many of the gains in model capabilities. @vvkgopalan sits down with @madiator & @AlexGDimakis, founders of @bespokelabsai, a company defining the category of data research. (00:00) Introducing Mahesh and Alex (03:40) Teaching a 1000 person ML class at Berkeley / changes in CS education in the age of AI (06:10) Open source data for frontier model training, OpenThoughts & training models for data curation (12:00) Why multiple-answer yields more performant models, benefits of keeping unverifiable examples in datasets (17:20) RL scaling laws / introduction to Bespoke (19:45) "Environment-driven Development" (28:35) New capabilities in frontier models & coding agents (34:30) Market outlook for RL environments (42:30) Exciting recent research and how we arrived at today's state of RL (50:00) Bespoke's Series A




Frankly I think this is the reason Devin is having such a comeback Nobody is really doubting the productivity gains of AI, and I would guess that companies would still be willing to pay the exponential if they must... But token spend is scaled and open source is now really good. It makes sense we are now spending energy to curb the runaway train Extreme high-growth startups are only now thinking about token spend, but this has been an enterprise (read: Publicly Traded Company) concern since day 1 Want to understand how Cognition so quickly grabbed all the big banks and giant Fortune 100 enterprises as customers? Aligned incentives is the answer. 1. Being an independent company Because we are not a model lab with $100B+ raised and $1T+ of data center commitments, we don't need to "catch up" by selling increasingly more expensive tokens Nor do we need to push a specific model family to make margins. Our only calculus is - "Is this the best model for the job?" - "Can we make the user more productive?" - "Can we save the user money?" (increasingly) This comes in the form of post-training research (building cheap + specifically tuned coding models) + new coding evals (FrontierCode benchmarks) + model routing (a lot behind-the-scenes of Devin's cloud harness). You should be skeptical of an Italian restaurant pushing the expensive market price specials. Just like you should be skeptical of a model lab pushing the newest most expensive model 2. Enterprise cost controls As a pre-requisite to selling enterprise contracts to the biggest companies in the world, you need really good spend controls. These banks and big conglomerates have been token-sensitive since day 1. They saw the writing on the exponential. For this reason, Devin has the most complete & robust spend controls of any coding agent on the market. The boring stuff of orgs, users, scopes, limits. But it matters. 3. AI Productivity alignment Cognition has an "AI Productivity Guarantee" That means if Devin delivers less engineering value than you’re paying for, Cognition will fund your usage until it does, up to $10 million. This is the tip of the iceberg and the one thing about Cognition that has been most novel to me since joining. Everything (and I mean everything) in our GTM motion is oriented around ROI. Every conversation is rooted in the actual engineering tickets we are taking off the backlog. I can only imagine what it would be like if instead conversations were rooted in "how can we entice users to burn through tokens"

Introducing Devin Security Swarm A more cost effective and accurate way to find security vulnerabilities in complex codebases, based on a new architecture: Agentic MapReduce.

Conventional model routing sucks. It passes benchmarks but fails to write code you'd actually merge. Introducing Devin Fusion, a new hybrid-model harness for agentic coding. In testing, it reduces the cost of Fable-level intelligence by 35% and still feels good to use.















