
Shopify's LLMs beat frontier models on a range of tasks at a fraction of the cost. The reason: we put systems in place that enable them to improve themselves, learning from a range of commerce tasks every day. We're presenting our Model Optimization Flywheel at @ICMLconf: a continuous pipeline that turns Shopify's product expertise into robust evals, mines low-scoring conversations, critiques them, repairs them, and feeds them back into the model. Then we compress the prompts without losing quality, so we can make it faster and cheaper. We present an example of the flywheel working at scale: our GraphQL agent. Serving cost dropped from $27M to $1M annualized (−96%). We compressed our system prompt 4× and still beat frontier models on quality. @Drewch and @cmazzaanthony will share concrete recipes, quality-cost-latency trade-offs, and a blueprint you can actually build from. 📅 Monday, July 6 · 11:30am–12:30pm KST 📍 COEX, Hall D1 Link in thread. 👇



