Carter Huffman

43 posts

Carter Huffman banner
Carter Huffman

Carter Huffman

@whuffman

@MIT alum | Formerly @NASA @NASAJPL | Current CTO & Co-founder of @Modulate_ai

Katılım Ağustos 2015
74 Takip Edilen23 Takipçiler
Carter Huffman
Carter Huffman@whuffman·
One of the most effective ways to build robust AI systems is to have them challenge each other 🤼‍♂️ It's tempting to think of generation and detection as two separate machine learning problems. In reality, they're part of the same feedback loop 🔁 👉 Every breakthrough in synthetic voice generation creates new edge cases for detection. 👉 Every improvement in detection exposes new limitations in generation. Advancing one without the other leaves you with an incomplete understanding of the problem. That's one of the reasons we've invested heavily in both at @modulate_ai. The two systems continually push each other forward, helping us build more resilient voice AI as the technology evolves. I'm curious - do you think the future of AI will be shaped more by specialized models solving individual problems, or by systems that continuously learn from and improve one another? 👀
English
0
0
0
10
Carter Huffman
Carter Huffman@whuffman·
Mandarin just got a new best-in-class transcription model. It's ours. Modulate Transcribe posted the lowest Utterance Error Rate of any provider on Sierra's μ-Bench for Mandarin (zh-CN): 27.4%. Google Chirp-3 came in at 29.5%. ElevenLabs Scribe v2 at 37.1%. We beat both 🏅 And this wasn't the easy locale to win. Sierra's own study named Mandarin the hardest in the benchmark - tone, segmentation, code-switching, the works. Most models take their biggest accuracy hit right here. We picked the hardest fight on the board and won it. Quick context on why the number matters: UER is not WER. It measures the share of utterances with at least one significant, meaning-changing error - the kind that actually breaks a voice agent or a downstream workflow. You can look fine on WER and still fail here. We didn't. Then there's speed. Our P95 batch completion latency was 658ms. Chirp-3's was 1,196ms. That's 45% lower latency, on top of the lower error rate. Not a trade-off. Both. Best accuracy on the leaderboard. Faster than Chirp-3. Hardest locale on the leaderboard. Proud of the team for this! 🙌 Don't take my word for it. Benchmark us on your own audio: 1,000 free credits: bit.ly/4b6kS35
Carter Huffman tweet media
English
1
1
2
37
Carter Huffman
Carter Huffman@whuffman·
I've always believed that curiosity compounds ☄️ The questions you obsess over early in life rarely stay confined to one field - they shape how you approach every problem that comes after. Whether you're studying the universe, writing code, or building AI, the process is surprisingly similar: ask better questions, understand the fundamentals, and keep pulling on the thread until things click. I've found that mindset has been far more valuable than just expertise in any one domain. What's a topic you got obsessed with years ago that still influences how you think today? 👀
English
0
0
1
11
Carter Huffman
Carter Huffman@whuffman·
Every startup has a moment where the question shifts. It stops being "Can we build this?" and becomes "How do we make this useful for as many people as possible?" That transition changes how you prioritize, how you make decisions, and ultimately what success looks like. In this clip, I talk about how that shift happened for me at @modulate_ai - and how my perspective on building AI has evolved over the years.
English
0
0
2
17
Carter Huffman
Carter Huffman@whuffman·
Modulate Transcribe just ranked #1 out of 84 models on @huggingface Open ASR Leaderboard. Trained on 500M+ hours of real-world, noisy audio - and still costs $0.03/hour. Incredibly proud of my team at @modulate_ai! We're outperforming players with 10x our resources, on a benchmark anyone can go check for themselves. Nothing beats building the best product and letting the numbers speak. Check out the leaderboard here: bit.ly/4b6yWK3
Carter Huffman tweet media
English
1
1
5
166
Carter Huffman
Carter Huffman@whuffman·
We launched AI Music Detection today. Here's what it actually took to build it. The naive version of this problem is a binary classifier: AI or not.  That's not hard.  The hard version is what the music industry actually needs:  Something that can tell you where in a track AI appears, whether it's the vocals or the instrumentation, and do it reliably across genres as diverse as compressed hip-hop vocals and dense electronic production. We solved it with an ensemble of three models running together.  - One identifies frames containing music, speech, or both.  - A second evaluates AI-generated vocals.  - A third evaluates AI-generated instrumentals.  The result is segment-level probabilities across the full clip - not a single score you can't act on. We built this on the same audio-native foundation we developed for real-world voice intelligence in online gaming: noisy, live, multilingual, emotionally variable audio that breaks lab-condition models.  That background shaped how we approached music - treating detection as an audio signal problem, not a metadata problem. 95% precision across 76 genres in internal testing against leading models including @Suno 5.5. The arms race is real — generative quality is improving fast and detection has to keep up.  We're continuing to improve across genres, production styles, and new generation platforms. But this is a practical, scalable signal platforms can use today. If you're building in the music space and want to dig into the technical approach, I'd love to talk. Test the API here. First 1000 credits on us
English
0
0
3
54
Carter Huffman
Carter Huffman@whuffman·
Every company processing voice has the same blind spot 👀 Good PII tools exist for text. For audio? Transcribe first, filter the transcript, call it done. The recording stays intact. In storage. Unredacted. That's the actual exposure - and it's what every existing tool leaves untouched. Nobody solved it properly because of where they started. If your architecture began with text, you can't retrofit real audio redaction. You'd need to understand spoken content in the stream, before it hits storage. That requires being native to audio. Most PII vendors aren't. Velma by @modulate_ai is. Same foundation we've built for years - models that understand audio directly, not as a stepping stone to text. That's what made this possible. A few things worth flagging in what we shipped: 94 entity types isn't a vanity number. Sensitive information doesn't follow format rules. A diagnosis doesn't announce itself with 16 digits. Getting those right forces genuine conversation understanding - and that makes the whole system more accurate across the board. Security question answers are almost entirely undetected across the market. Mother's maiden name. First pet. Childhood best friend. The exact vectors social engineering is built around - sitting in call recordings, everywhere, unredacted. If you already know what's in your voice data and want to do something about it before a regulation forces your hand - test our API. First 1000 credits are on us (Link below)
English
1
0
2
42
Carter Huffman
Carter Huffman@whuffman·
Deepfake fraud attempts surged 1,300% between 2024 and 2025 📈 The reason? Synthetic voice is cheap to generate. And most bank fraud stacks are architecturally blind to it. They're built around transaction data - what happened after the call. But the attack has moved. It's happening during the call. The urgency script. The spoofed voice. The social engineering that gets an agent to bypass verification before a single flag gets raised. A transcript won't catch that ❌ Rules won't catch that ❌ You need models that understand voice the way voice actually works - tone, cadence, stress, acoustic anomalies. Not just words. We wrote a practical breakdown of what real-time voice AI detection looks like in banking: where it plugs in, what it scores, and how it shifts intervention from post-wire to mid-call. (Link in caption) If you're building in this space, our API is open.
English
1
0
2
48
Carter Huffman
Carter Huffman@whuffman·
@modulate_ai So excited to share our rich conversation understanding model with everyone - please give it a try!
English
0
0
1
249
Carter Huffman retweetledi
Modulate
Modulate@modulate_ai·
The world's first audio-native AI model is now available as an API. Velma listens and understands like a human — emotions, tone, intent, rhythm, vocal stress. Already analyzed 550M+ hours of conversation for Fortune 500s. Now open to developers. 🧵
English
52
54
264
7.4M
Carter Huffman
Carter Huffman@whuffman·
Words are the least interesting part of a conversation. Tone. Hesitation. What someone almost said. That's where the meaning lives. Voice AI that only reads transcripts will never get there. At @modulate_ai, we built something that does 💪👀
English
0
0
1
8
Carter Huffman
Carter Huffman@whuffman·
If I had to spend the next 5 years on just one problem: It’d be getting different AI systems to actually work together - in sync, toward a single goal, without breaking accuracy or efficiency. Feels like we’re still early on that. At @modulate_ai, this is a big part of how we think about building - not just better models, but better systems. My answer’s set - curious about yours: If you could only work on one problem for the next 5 years, what would it be?
English
0
0
0
12
Carter Huffman
Carter Huffman@whuffman·
The intersection of audio + AI is moving fast - and some of the best people building it are right here in Boston. @modulate_ai is joining Yamaha Music Innovations Fund next Thursday night to bring them all together. Harmonix, Suno, and more under one roof for live demos, real talk, and good conversations. If this space is on your radar, come hang. Would love to see you there! bit.ly/4tTLzi3 #BosTechWeek
Carter Huffman tweet media
English
0
0
0
46
Carter Huffman
Carter Huffman@whuffman·
I think about this a lot: We get to build technology that impacts millions of people every day. That’s a rare kind of responsibility - and it doesn’t feel abstract. It shows up in the small decisions, the tradeoffs, the standards we hold ourselves to. At @modulate_ai, that’s the bar: maximize the good we can do, and be really intentional about where and if we get it wrong. Grateful for the team, and for the chance to work on problems that actually matter❗🙌 Curious how others think about this - how do you balance moving fast with getting things right?
English
0
0
0
18
Carter Huffman
Carter Huffman@whuffman·
The most important thing we’re working on at @modulate_ai right now: Moving beyond narrow detection - toward systems that truly can understand risk in context. Real-world threats don’t show up in neat categories. They overlap, evolve, and look different in every conversation. So instead of building for one slice at a time, we’re focused on a unified layer of voice intelligence: one that interprets what’s actually happening, not just flags keywords. That shift is where this becomes genuinely useful👏
English
1
0
0
8