Cody Mazza-Anthony

17 posts

Cody Mazza-Anthony banner
Cody Mazza-Anthony

Cody Mazza-Anthony

@cmazzaanthony

Machine Learning @shopify. Former quant.

Katılım Kasım 2020
178 Takip Edilen80 Takipçiler
Cody Mazza-Anthony retweetledi
Shopify Engineering
Shopify Engineering@ShopifyEng·
Starts in 3 hours at COEX Hall D1 (11:30am KST) — great lunch break talk. @Drewch and @cmazzaanthony show how we cut GraphQL agent serving costs 96% while still beating frontier models on quality. 👊 Looking forward to a great onsite discussion!
Shopify Engineering@ShopifyEng

Shopify's LLMs beat frontier models on a range of tasks at a fraction of the cost. The reason: we put systems in place that enable them to improve themselves, learning from a range of commerce tasks every day. We're presenting our Model Optimization Flywheel at @ICMLconf: a continuous pipeline that turns Shopify's product expertise into robust evals, mines low-scoring conversations, critiques them, repairs them, and feeds them back into the model. Then we compress the prompts without losing quality, so we can make it faster and cheaper. We present an example of the flywheel working at scale: our GraphQL agent. Serving cost dropped from $27M to $1M annualized (−96%). We compressed our system prompt 4× and still beat frontier models on quality. @Drewch and @cmazzaanthony will share concrete recipes, quality-cost-latency trade-offs, and a blueprint you can actually build from. 📅 Monday, July 6 · 11:30am–12:30pm KST 📍 COEX, Hall D1 Link in thread. 👇

English
6
4
24
10.4K
Cody Mazza-Anthony retweetledi
Shopify Engineering
Shopify Engineering@ShopifyEng·
Shopify's LLMs beat frontier models on a range of tasks at a fraction of the cost. The reason: we put systems in place that enable them to improve themselves, learning from a range of commerce tasks every day. We're presenting our Model Optimization Flywheel at @ICMLconf: a continuous pipeline that turns Shopify's product expertise into robust evals, mines low-scoring conversations, critiques them, repairs them, and feeds them back into the model. Then we compress the prompts without losing quality, so we can make it faster and cheaper. We present an example of the flywheel working at scale: our GraphQL agent. Serving cost dropped from $27M to $1M annualized (−96%). We compressed our system prompt 4× and still beat frontier models on quality. @Drewch and @cmazzaanthony will share concrete recipes, quality-cost-latency trade-offs, and a blueprint you can actually build from. 📅 Monday, July 6 · 11:30am–12:30pm KST 📍 COEX, Hall D1 Link in thread. 👇
English
8
14
292
54.6K
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
I rebuilt my personal website (codymazzaanthony.com) from scratch with no frameworks or build tools. Just HTML, CSS, and the simplest possible markdown rendering. Turns out "as little as possible" is pretty liberating.
English
0
0
2
268
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
Just discovered this amazing repository of machine learning algorithm implementations! Check it out 👉🏼 buff.ly/42GNjPy. The Gaussian mixture model and linear model implementations were particularly enjoyable. #AI #ML
English
0
0
0
297
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
Game changer for inference of open source models! 🚀
Jay Scambler@JayScambler

Groq is serving the fastest responses I've ever seen. We're talking almost 500 T/s! I did some research on how they're able to do it. Turns out they developed their own hardware that utilize LPUs instead of GPUs. Here's the skinny: Groq created a novel processing unit known as the Tensor Streaming Processor (TSP) which they categorize as a Linear Processor Unit (LPU). Unlike traditional GPUs that are parallel processors with hundreds of cores designed for graphics rendering, LPUs are architected to deliver deterministic performance for AI computations. The LPU's architecture is a departure from the SIMD (Single Instruction, Multiple Data) model used by GPUs and favor a more streamlined approach that eliminate the need for complex scheduling hardware. This design allows every clock cycle to be utilized effectively, ensuring consistent latency and throughput. For developers, this means that performance can be precisely predicted and optimized which is critical in real-time AI applications. Energy efficiency is another area where LPUs shine. By reducing the overhead of managing multiple threads and avoiding the underutilization of cores, LPUs can deliver more computations per watt. Groq's innovative chip design allows multiple TSPs to be linked together without the traditional bottlenecks found in GPU clusters making them extremely scalable. This enables linear scaling of performance as more LPUs are added simplifying the hardware requirements for large-scale AI models and making it easier for developers to scale their applications without rearchitecting their systems. So what does this all mean? LPUs could provide a massive improvement compared to GPUs for serving AI applications in the future! If anything it will be great to have alternative high performing hardware since A100s and H100s are so in demand

English
0
0
0
422
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
"Competence is often less of a problem than confidence. An underrated aspect of doing anything hard is believing in yourself. Action creates both confidence and momentum ..." A Tiny Thought found in the @farnamstreet Weekly Newsletter fs.blog/brain-food/jan…
English
0
0
1
68
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
I struggle to write system prompts. I prefer a more TDD approach to writing system prompts. Prompts Royale makes it easy (promptsroyale.com). It takes your test cases and generates multiple system prompts, and chooses the best one. Game-changing tool! #promptengineering
English
0
0
2
63
Cody Mazza-Anthony
Cody Mazza-Anthony@cmazzaanthony·
Just watched A Hackers' Guide to Language Models (youtube.com/watch?app=desk…). I love the practical approach to understanding LLMs, it reminded me of Andrej Karpathy's Zero to Hero series.
YouTube video
YouTube
English
0
0
1
63
Cody Mazza-Anthony retweetledi
tobi lutke
tobi lutke@tobi·
It's amazing that you can start a fresh python notebook, and this code is all you need to run hardware accelerated, streamed, inference on your mac. It even downloads the model.
tobi lutke tweet media
English
7
25
244
44.1K