Sinan Ozdemir

516 posts

Sinan Ozdemir banner
Sinan Ozdemir

Sinan Ozdemir

@Prof_OZ

AI Education @FireworksAI_HQ 🔥🎆 NLP + Gen AI Expert / LLM whisperer 🤖 AI Author 📚 Founder @Aikylie (acquired) 📈 Fellow @YCombinator

San Francisco, CA Katılım Mart 2012
1.6K Takip Edilen1.8K Takipçiler
Fireworks
Fireworks@FireworksAI_HQ·
@Prof_OZ The way you found an emoji for "distill" is why we hired you.
English
1
0
2
178
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
1. Control your own data 🔒 2. Adapt (fine-tune!) models 🛠️ 3. Deploy them where you need to 🚀 4. Fortify your evals 💪 5. Distill from the frontier ⚗️ 6. Own the value you create and win instead of giving it away🏆 I cannot wait for the day where we don't need to even say this out loud, it's just the norm
Jensen Huang@JensenHuang

For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models. images.nvidia.com/pdf/Open-Weigh…

English
1
0
5
434
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
Yesss this is excellent stuff! More examples on the cookbook are coming all the time for us (and our agents) to use github.com/fw-ai/cookbook
Terry Chen@227Terry

Almost every agent product has an intent-routing layer: an LLM that looks at a task and decides what should handle it. For a proof of concept, a prompt and an OpenAI API key may be enough. But once you are routing millions of production requests, cost and latency start to matter. You could send every decision to the best frontier model and expect it to work. In our test, GPT-5.5 reached 91.7% routing accuracy with 1.19-second P50 latency, costing $2.40 per 1K requests. To see how open source performs, I asked my coding agent to “vibe tune” a router using Fireworks’ Training Agent skill. It selected managed LoRA SFT and trained Qwen3-8B on 600 labeled examples. Fine-tuning improved Qwen’s routing accuracy from 80.0% to 93.3%, with 0.38-second P50 latency. On this small holdout, its quality was effectively comparable to GPT-5.5, while latency was roughly three times lower. The economics depend on scale. One dedicated H200 costs $5,040 per month, breaking even with GPT around 2.1M monthly routing requests. At 10M requests, tuned Qwen is about 4.8 times cheaper. Below break-even, GPT remains cheaper. Closed APIs are often the fastest way to get started. But when production volume makes cost, latency, and control important, custom-trained open-weight models become compelling.

English
0
0
2
168
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
"Which is better: Kimi K3 or Fable?" is the wrong question. Go multi-model. We at @FireworksAI_HQ tested K3 and Fable on ~1,000 tasks. Route each task to the right model and you land above both with up to 50x cost savings. fireworks.ai/blog/kimik3-fa…
English
0
0
5
163
Fireworks
Fireworks@FireworksAI_HQ·
@Prof_OZ Make sure to put out the cookies and coffee we bought, Sinan!
English
1
0
1
379
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
@Vicxai @FireworksAI_HQ It's coming up! Demos will be me walking through code live and you will have access to that code from our Github. Also we will have office hours the week after for any follow up questions and another webinar the following month and the month after that :)
English
0
0
1
29
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
I'm launching a live series at @FireworksAI_HQ to help you ship on inference + training Every month: ⚡Webinar - live demos + what's new at Fireworks ⚡2 office hours - bring any AI/Fireworks question ⚡Discord AMA First webinar runs next Thursday 👇 fireworks.ai/event/fine-tun…
Sinan Ozdemir tweet media
English
0
1
4
4.9K
Sinan Ozdemir retweetledi
Dmytro Dzhulgakov
Dmytro Dzhulgakov@dzhulgakov·
do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers give 1/5 cached discount for GLM-5.2 we at @FireworksAI_HQ dropped it to 1/10, matching GPT/Claude that's -40% typical savings, have fun!
English
17
9
130
13.6K
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
My own talk track on top of this excellent thread. Let me know your questions!
Dmytro Dzhulgakov@dzhulgakov

DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production system Let's understand it with 10 ideas, starting from the very basics 🧵

English
0
0
0
149
Sinan Ozdemir retweetledi
Dmytro Dzhulgakov
Dmytro Dzhulgakov@dzhulgakov·
DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production system Let's understand it with 10 ideas, starting from the very basics 🧵
Dmytro Dzhulgakov tweet mediaDmytro Dzhulgakov tweet media
English
13
119
970
306.8K
Sinan Ozdemir retweetledi
Fireworks
Fireworks@FireworksAI_HQ·
In a joint Fireworks and @Faros_AI evaluation of 211 real engineering tasks, Claude Code + GLM-5.2 beat both Claude Code + Opus 4.8 and Codex + GPT-5.5: - Judge score: 0.568 vs. 0.521 and 0.466 - Time per task: 321s vs. 775s and 392s - Cost per task: $0.92 vs. $1.76 and $2.06 Most importantly, Faros tested the models on its own repositories and work, not just public benchmarks. Model choice should be based on where a model clears your bar, on your work, at a cost that makes sense.
Faros@Faros_AI

With our partner @fireworksai_hq, we ran an evaluation to find out if open models have gotten strong enough to handle real software engineering work. Our result: Claude Code + GLM-5.2 matched Kimi on quality while running faster and cheaper. More details: faros.ai/blog/open-mode…

English
9
14
162
17.4K
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
@thegenioo @FireworksAI_HQ Agreed! What I'm showing is that now I can turn off that custom model in favor of the native GLM 5.2 model which only just got added recently
English
0
0
1
19
Hamza
Hamza@thegenioo·
@Prof_OZ @FireworksAI_HQ yup they will with Cursor Sub but what you are showing is custom model added via API and that was always available from day one
English
1
0
0
22
Sinan Ozdemir
Sinan Ozdemir@Prof_OZ·
Catch me live talking about optimizing token use on the model side for the Infer() Summit, starting soon! I'll be talking about when reasoning doesn't add anything, how to set up a routing flywheel and what it gets you, and how training jobs like distillation can make a huge impact. luma.com/cluvtqhu?tk=Jq…
English
0
0
0
63