Rez Karim

24K posts

Rez Karim banner
Rez Karim

Rez Karim

@rezkhere

Giving your AI agents superpowers at @unsora_ai. I talk AI video, ads & prompt engineering.

Try Seedance 2.0 → Katılım Ocak 2023
1K Takip Edilen38.1K Takipçiler
Sabitlenmiş Tweet
Rez Karim
Rez Karim@rezkhere·
Remember when AI couldn't draw a hand? Seven fingers, knuckles pointing backwards. And the AI spaghetti videos. That was three years ago. Images are done now. Video is close enough that you scrolled past AI ads this week and clocked exactly zero of them. Code writes itself and there are like 40 coding agents. AI voice spent that entire stretch sounding like the lady voice in a 2014 GPS. Flat, evenly spaced and every sentence landing with the same weight, like it's reading from a phone book. Here's why it stayed broken. Bad images are funny. You screenshot the seven fingers, it goes viral for being bad, someone fixes it. Bad audio is just boring. It doesn't fail spectacularly, so it never got that pressure. The bigger problem was the scoring. The whole industry graded AI voices on whether you could make out the words. So the models learned to over pronounce everything, hitting every syllable like a newsreader. Perfectly clear but robotic. Everyone was chasing a score that had nothing to do with sounding human. Meanwhile a small open-source team was doing something harder. Their lead researcher, an ex-NVIDIA engineer, went all in on an approach the rest of the field had written off. Two years early. No funding announcements or launch tour. He just put the whole thing on GitHub for free. It's sitting at 50,000+ stars now. Then they ran the test everyone else avoided. For 10 days they piped real users through their model and every big competitor with the listener never told which was which. Thousands of real people, real scripts. Whichever voice you actually preferred, they logged it. Theirs came out on top. It beat ElevenLabs about 6 times out of 10, head to head. It beat OpenAI's voice model 8 times out of 10. The gap was widest on the breathing, the pauses, the little hesitations, which is exactly the stuff that makes a voice sound like a person instead of a machine reading. They ran on real users rather than a lab, which is more than most of these claims can say. That's Fish Audio. This week they shipped S2.1 Pro: - Clone anyone's voice from 15 seconds of audio - Fast enough to hold a live conversation - 83 languages, one model - Type [whisper] or [sigh] mid-sentence and it does it - Around 70% cheaper than ElevenLabs - Free to download and run yourself Voice was the last thing on the list. Around 20 people with a free repo got there before other billion dollar companies did.
English
12
26
42
8.1K
Rez Karim
Rez Karim@rezkhere·
The gap between AI voice and everything else was the quietest failure in the field. No headlines, no hype, just a decade of models that read text instead of speaking it. That's the gap S2.1 Pro just closed.
Rez Karim@rezkhere

Remember when AI couldn't draw a hand? Seven fingers, knuckles pointing backwards. And the AI spaghetti videos. That was three years ago. Images are done now. Video is close enough that you scrolled past AI ads this week and clocked exactly zero of them. Code writes itself and there are like 40 coding agents. AI voice spent that entire stretch sounding like the lady voice in a 2014 GPS. Flat, evenly spaced and every sentence landing with the same weight, like it's reading from a phone book. Here's why it stayed broken. Bad images are funny. You screenshot the seven fingers, it goes viral for being bad, someone fixes it. Bad audio is just boring. It doesn't fail spectacularly, so it never got that pressure. The bigger problem was the scoring. The whole industry graded AI voices on whether you could make out the words. So the models learned to over pronounce everything, hitting every syllable like a newsreader. Perfectly clear but robotic. Everyone was chasing a score that had nothing to do with sounding human. Meanwhile a small open-source team was doing something harder. Their lead researcher, an ex-NVIDIA engineer, went all in on an approach the rest of the field had written off. Two years early. No funding announcements or launch tour. He just put the whole thing on GitHub for free. It's sitting at 50,000+ stars now. Then they ran the test everyone else avoided. For 10 days they piped real users through their model and every big competitor with the listener never told which was which. Thousands of real people, real scripts. Whichever voice you actually preferred, they logged it. Theirs came out on top. It beat ElevenLabs about 6 times out of 10, head to head. It beat OpenAI's voice model 8 times out of 10. The gap was widest on the breathing, the pauses, the little hesitations, which is exactly the stuff that makes a voice sound like a person instead of a machine reading. They ran on real users rather than a lab, which is more than most of these claims can say. That's Fish Audio. This week they shipped S2.1 Pro: - Clone anyone's voice from 15 seconds of audio - Fast enough to hold a live conversation - 83 languages, one model - Type [whisper] or [sigh] mid-sentence and it does it - Around 70% cheaper than ElevenLabs - Free to download and run yourself Voice was the last thing on the list. Around 20 people with a free repo got there before other billion dollar companies did.

English
2
6
7
4.5K
Rez Karim retweetledi
Rez Karim
Rez Karim@rezkhere·
Remember when AI couldn't draw a hand? Seven fingers, knuckles pointing backwards. And the AI spaghetti videos. That was three years ago. Images are done now. Video is close enough that you scrolled past AI ads this week and clocked exactly zero of them. Code writes itself and there are like 40 coding agents. AI voice spent that entire stretch sounding like the lady voice in a 2014 GPS. Flat, evenly spaced and every sentence landing with the same weight, like it's reading from a phone book. Here's why it stayed broken. Bad images are funny. You screenshot the seven fingers, it goes viral for being bad, someone fixes it. Bad audio is just boring. It doesn't fail spectacularly, so it never got that pressure. The bigger problem was the scoring. The whole industry graded AI voices on whether you could make out the words. So the models learned to over pronounce everything, hitting every syllable like a newsreader. Perfectly clear but robotic. Everyone was chasing a score that had nothing to do with sounding human. Meanwhile a small open-source team was doing something harder. Their lead researcher, an ex-NVIDIA engineer, went all in on an approach the rest of the field had written off. Two years early. No funding announcements or launch tour. He just put the whole thing on GitHub for free. It's sitting at 50,000+ stars now. Then they ran the test everyone else avoided. For 10 days they piped real users through their model and every big competitor with the listener never told which was which. Thousands of real people, real scripts. Whichever voice you actually preferred, they logged it. Theirs came out on top. It beat ElevenLabs about 6 times out of 10, head to head. It beat OpenAI's voice model 8 times out of 10. The gap was widest on the breathing, the pauses, the little hesitations, which is exactly the stuff that makes a voice sound like a person instead of a machine reading. They ran on real users rather than a lab, which is more than most of these claims can say. That's Fish Audio. This week they shipped S2.1 Pro: - Clone anyone's voice from 15 seconds of audio - Fast enough to hold a live conversation - 83 languages, one model - Type [whisper] or [sigh] mid-sentence and it does it - Around 70% cheaper than ElevenLabs - Free to download and run yourself Voice was the last thing on the list. Around 20 people with a free repo got there before other billion dollar companies did.
English
12
26
42
8.1K
Arsalan
Arsalan@AIwithArsalan·
🚨 RIP ChatGPT. RIP Canva. RIP PowerPoint! One prompt can now generate documents, presentations, videos, images, audio, and research. Here’s what surprised me most 👇
English
23
2
27
7.9K
Rez Karim
Rez Karim@rezkhere·
@nrqa__ Templafy is a great one for making slides.
English
0
0
1
27
Nelly;
Nelly;@nrqa__·
using AI to make a work deck: 2023: generate slides, open in PowerPoint, rebuild the whole thing anyway 2024: prettier slides, same rebuild 2026: Templafy's PowerPoint Agent ships a real editable .pptx, charts and brand colors intact, nothing to redo the rebuild tax is finally gone: templafy.com/ai-powerpoint-…
Templafy@Templafy

Creating PowerPoint presentations with Templafy's free AI Agent just got even better! 🗨️Chat history 🎨Custom brand colors 📊Editable charts Create presentations—fully editable in PowerPoint: templafy.com/experience-tem…

English
4
1
19
5.1K
Nawi
Nawi@Aria_Nawi·
I'm speechless. Fish Audio just beat an $11B voice AI company at its own game.They ran a 10-day blind test with 5,000+ real users. Fish Audio won 66% of the time. Nearly 2x the next closest model. In Chinese, it wasn't even close. They scored 8x higher than ElevenLabs. And today they're shipping something even better: S2.1 Pro. You can now clone any voice from 15 seconds of audio in 83+ languages with emotion control on every single word. Watch the video to hear the different yourself:
English
32
278
311
115.3K
Mushfiq Sajib
Mushfiq Sajib@heysajib·
Most AI slide tools give you a pretty deck you cannot actually edit. Templafy's free agent hands you a real PowerPoint file, brand colors intact, charts still editable. That last part is the difference between a demo and something you can actually ship.
Templafy@Templafy

Creating PowerPoint presentations with Templafy's free AI Agent just got even better! 🗨️Chat history 🎨Custom brand colors 📊Editable charts Create presentations—fully editable in PowerPoint: templafy.com/experience-tem…

English
28
48
118
61K
EyeingAI
EyeingAI@EyeingAI·
This is probably the first time.. I’ve seen a slide update focus on what happens after generation. You can come back to the same chat, keep refining the deck, match your own brand colours and edit the charts properly in PowerPoint. 👇
Templafy@Templafy

Creating PowerPoint presentations with Templafy's free AI Agent just got even better! 🗨️Chat history 🎨Custom brand colors 📊Editable charts Create presentations—fully editable in PowerPoint: templafy.com/experience-tem…

English
6
1
17
6.5K
Rez Karim
Rez Karim@rezkhere·
You don't understand... A small team just shipped something the big labs couldn't, at a fraction of the price. S2.1 Pro clones from 5 seconds of audio, gives you word-level control over emotion and intonation. Already runs in production at HeyGen, LiveKit, Retell and Sanas. And they are open source. Wild stuff
Fish Audio@FishAudio

Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro. >It can clone a voice from 5 seconds of audio >2x faster than Cartesia & 1/6th the cost of Eleven Labs >most expressive model with word level control over emotion, intonation, pacing etc We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production. If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free. Book a demo: s.fish.audio/tmapke To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.

English
3
11
27
6.2K
Fish Audio
Fish Audio@FishAudio·
Today we’ve raised $52M Seed and we are announcing the public launch of S2.1 Pro. >It can clone a voice from 5 seconds of audio >2x faster than Cartesia & 1/6th the cost of Eleven Labs >most expressive model with word level control over emotion, intonation, pacing etc We support frontier AI companies including HeyGen, LiveKit, Retell, Sanas, and OpenArt all run our model in production. If you're a business and we can't cut your voice AI costs by 50%, we'll give you 1 year of Fish Audio for free. Book a demo: s.fish.audio/tmapke To celebrate our first birthday, we'll give you 1 month of S2.1 Pro for free. Like, retweet, and comment “Fish” to get it.
English
4.6K
6.4K
12.9K
5.5M
Rez Karim
Rez Karim@rezkhere·
@FishAudio the 5 second clone is a little too good imo. every dev paying $100/M characters needs to see this
English
0
0
0
102
Rez Karim
Rez Karim@rezkhere·
Lighting used to cost me the shot twice over. You'd get a performance you loved with a key light coming from the wrong side. And the only fix was regenerating. So you'd keep the bad light or keep rolling and lose the moment. Relight in Topview Film Studio separates those two problems. You keep the performance and move the light. Sounds small but it has changed which takes I want to keep, which is most of editing.
TopviewAI@TopviewAIhq

AI can generate beautiful clips. But filmmaking needs direction. Introducing Topview Film Studio: the cinematic generation suite for AI filmmakers. 🎬 Direct micro-expressions. Set the camera. Block the scene in 3D. Refine portraits to remove the AI look. Topview Film Studio brings these film-grade controls into one cinematic canvas. Stop prompting clips. Start directing shots.

English
8
7
25
2K
TopviewAI
TopviewAI@TopviewAIhq·
AI can generate beautiful clips. But filmmaking needs direction. Introducing Topview Film Studio: the cinematic generation suite for AI filmmakers. 🎬 Direct micro-expressions. Set the camera. Block the scene in 3D. Refine portraits to remove the AI look. Topview Film Studio brings these film-grade controls into one cinematic canvas. Stop prompting clips. Start directing shots.
English
75
65
297
686.9K
Gina Acosta
Gina Acosta@ginacostag_·
Claude Fable 5 is powerful. Your prompts decide how powerful. Here are 12 tips covering: → memory, pacing, and verification, no guesswork, built for how people actually work in 2026. [Bookmark for later]
Gina Acosta tweet media
Gina Acosta@ginacostag_

A vendor claimed their API was 40x faster than the competition. The number was real. The comparison wasn't. I write about AI infrastructure and build the explainer graphics people actually bookmark. My whole job is checking the number before it goes on a chart. Every launch ships a benchmark now. Almost none ship the methodology. To know if a claim is honest, you have to dig: → what was actually measured → on what hardware, at what payload → whether competitors were configured fairly, or just quoted off their own marketing pages That's an afternoon. Per claim. So for a while, I skipped the technical comparisons. Not worth the time. Which meant the best posts went to people just restating the press release. Now an AI employee does that legwork. He pulls: - the published numbers - the methodology (if one exists) - the changelog - what competitors claim in their own docs Then lays it all side by side with the gaps marked. Last week: a latency claim of 62ms vs. a field quoted at 244ms–2.6s. The 62ms was real. Reproducible. But: → 62ms was measured on cached queries → competitor numbers were cold starts, lifted straight from marketing pages Like for like? The real gap was ~3x. Not 40x. Still a good number. Just not the headline. That mismatch became the graphic. He assembles. I verify every source. Then it ships. The bottleneck in technical content was never the design. It was never having the time to check if the claim was true. Tell me I'm wrong. Hire @viktor__com for your team. $100 in credits included, no card. Full link in first comment.

English
3
3
14
12.4K
Gina Acosta
Gina Acosta@ginacostag_·
A vendor claimed their API was 40x faster than the competition. The number was real. The comparison wasn't. I write about AI infrastructure and build the explainer graphics people actually bookmark. My whole job is checking the number before it goes on a chart. Every launch ships a benchmark now. Almost none ship the methodology. To know if a claim is honest, you have to dig: → what was actually measured → on what hardware, at what payload → whether competitors were configured fairly, or just quoted off their own marketing pages That's an afternoon. Per claim. So for a while, I skipped the technical comparisons. Not worth the time. Which meant the best posts went to people just restating the press release. Now an AI employee does that legwork. He pulls: - the published numbers - the methodology (if one exists) - the changelog - what competitors claim in their own docs Then lays it all side by side with the gaps marked. Last week: a latency claim of 62ms vs. a field quoted at 244ms–2.6s. The 62ms was real. Reproducible. But: → 62ms was measured on cached queries → competitor numbers were cold starts, lifted straight from marketing pages Like for like? The real gap was ~3x. Not 40x. Still a good number. Just not the headline. That mismatch became the graphic. He assembles. I verify every source. Then it ships. The bottleneck in technical content was never the design. It was never having the time to check if the claim was true. Tell me I'm wrong. Hire @viktor__com for your team. $100 in credits included, no card. Full link in first comment.
Gina Acosta tweet media
English
9
3
18
22.4K
Markandey Sharma
Markandey Sharma@TechByMarkandey·
The part I like most is that Ben actually went to the GPU company. Not just “we optimized infra” in a blog post. The same founder who goes to Paris to buy cheese went to see the racks keeping Polsia alive. $1.2M/month AI bill after scaling to 5,000 paying users. So he rented GPUs, moved work to open-source models, and cut it to ~$100k. The agents are the cool part. The based part is going straight to the source and fixing the economics.
Ben Cera@Bencera

China is winning the AI war. It saved me from bankruptcy. My AI bill was 1M+/month. I fixed it with open source models. Presenting aisloP episode 3: "The Bill".

English
16
17
89
43K
Rez Karim
Rez Karim@rezkhere·
@Bencera makes sense. people still underestimate the complexity imo
English
0
0
1
55
Ben Cera
Ben Cera@Bencera·
@rezkhere It’s actually very hard to get oss models right in production.
English
1
1
2
486
Ben Cera
Ben Cera@Bencera·
China is winning the AI war. It saved me from bankruptcy. My AI bill was 1M+/month. I fixed it with open source models. Presenting aisloP episode 3: "The Bill".
English
120
149
1.5K
1.1M
Rez Karim
Rez Karim@rezkhere·
> be Ben > build AI agent company that helps raise its own $30M round > people go crazy > 500 → 5,000 paying users very fast > Anthropic bill hits $1.2M/month > friend checks the numbers > “bro this is trending to multi-million/month” > “yeah yeah I know, I just raised money, it’s fine” > it was not fine > rent GPUs > move work to open-source models > spend 2 months fixing the cost stack > June bill drops to ~$100k > whole time, the solution was renting hardware like it’s 2011
Ben Cera@Bencera

China is winning the AI war. It saved me from bankruptcy. My AI bill was 1M+/month. I fixed it with open source models. Presenting aisloP episode 3: "The Bill".

English
5
14
32
4.9K