dsa

8.7K posts

dsa

dsa

@dsa

builder @livekit, @x, @23andme, yc07 investor @sundayrobotics, @coreautoai

future Katılım Ağustos 2007
247 Takip Edilen10.3K Takipçiler
Sabitlenmiş Tweet
dsa
dsa@dsa·
Today is a day I’ll never forget. I grew up in Cupertino. My dad was a tech founder in the 80s/90s. I was in YC S07. LiveKit is my 5th company. The first 4 didn’t work out. I’ve had a lot of advantages — it still took 20 years to get here. Founders: keep taking shots.
LiveKit@livekit

We learn to speak before we learn to read. Voice is the most natural interface we have. We just raised a $100M to make building voice AI as easy as a web app.

English
85
39
938
101.8K
dsa retweetledi
LemonSlice
LemonSlice@LemonSliceAI·
The YC Office Hours Simulator was a huge success!  These were physical booths at Startup School where attendees could talk to AI avatars of the @ycombinator Partners. We interviewed over 100 attendees about their experience. Here's a recap 👇️ 1️⃣ The #1 complaint = people wished they could talk longer! Most conversations were 10min (the enforced time limit). There was a line of people waiting to talk to the avatars for the entire two-day conference 2️⃣ Some people wished the avatars were meaner or harsher. This was interesting! 3️⃣ Many people took notes (pen & paper, @meetgranola, or typing furiously) and said the conversation was genuinely helpful. Several folks waited in line for a 2nd or 3rd conversation :)  4️⃣ People said the avatar aspect felt different than a chatbot. The "face-to-face" interaction helped them feel a stronger connection, pay more attention, and open up more. A few people even said they forgot it was AI!  5️⃣ The magic moment = when people realized the avatars knew who they were and what they were working on (based on their conference application) 🪄 The "brains" for Office Hours Simulator was trained on hundreds of hours of real office hours. It truly simulated how each Partner talks and the types of questions they like to ask.  Built by the one-and-only @snowmaker and YC software team (so impressive) 👏 Powered by: @lemonslice (live avatars), @elevenlabs @ElevenLabsDevs (voice AI), @livekit (webRTC infra), @deepgramai (voice AI), @ployai (animations), Hamming AI (testing), and many other great companies.
LemonSlice tweet media
English
13
4
60
27.2K
dsa
dsa@dsa·
Mom: what do you think of Claude Tag? Me: haven't gotten a chance to use it. I just heard @jack has a new thing in this direction, actually. Mom: yeah Buzz, i know. 💀
English
1
0
10
827
dsa retweetledi
LiveKit
LiveKit@livekit·
LiveKit Unity SDK 2.0.0 is out of Developer Preview. Bring real-time voice, video, and AI into your Unity game. Build voice chat, live streaming, robot teleoperation, and voice AI agents as NPCs that listen, respond, and react. Give it a try today and let us know what you think.
English
2
5
52
4.8K
dsa retweetledi
Google Gemma
Google Gemma@googlegemma·
Every millisecond counts for voice agents. Introducing Gemma 4 31B on LiveKit Inference! Optimized for real-time voice agents: - 354ms time to first audio - 192ms time to first token - Beats GPT-4.1 in agentic tool use (31B achieves 76.9% on tau2bench) Fast, capable, and efficient!
English
62
186
2.2K
220.6K
dsa retweetledi
LiveKit
LiveKit@livekit·
Until now, it was impossible to synchronize video frames with other sensor data such as LiDAR, IMU, joint states, etc. streamed over WebRTC with microsecond level precision. Introducing Video Frame Metadata in LiveKit. Stamp every frame with your ground-truth timestamp, frame ID, and metadata. livekit.com/blog/frame-met…
English
1
8
52
4.1K
dsa retweetledi
Catalin Voss
Catalin Voss@CatalinVoss·
Everyone's asking what AI can do for their job. We asked if it could teach a 6-year-old. Introducing Ello 2.0: reading, math & more for ages 4-9. Free tier for all. If we want AI to force-multiply humanity, let's start with teaching our youngest.
English
66
49
278
166.4K
dsa
dsa@dsa·
LLMs that are great for coding agents aren't great for voice agents. So today we're launching our first hosted model: Gemma 4, optimized for time-to-first-token without sacrificing task completion rates. Fable/Opus/GPT 5.X optimize for intelligence over speed. But a voice agent handling a customer support call doesn't need PhD-level intelligence. It needs to be smart enough, and it needs to be fast. We verified this across thousands of real-world simulations before shipping.
LiveKit@livekit

Frontier models keep getting smarter…and slower. This is bad for voice agents, where every millisecond before the first word counts. So we optimized for speed. Introducing Gemma 4 31B on LiveKit Inference: 🗣️ 381ms to first sentence — 2x+ faster than the next model ✅ 88% task completion (see the criteria in reply) 💰 $1.20 / 1M output tokens Fast enough for real conversation, smart enough for effective agents.

English
4
1
19
4K
dsa retweetledi
This Week in AI
This Week in AI@ThisWeeknAI·
Apple hasn't fixed this problem since 2011. Using Siri is slower and more painful than using your browser. The UX isn't clear. Voice AI does the job that Siri was supposed to do: Tackle your tasks and have human-like interactions. Siri still feels like a transaction. @dsa @livekit
English
1
2
5
1.3K
dsa
dsa@dsa·
@mitchellh They copied all they could follow, but they couldn't copy my mind, And I left them sweating and stealing a year and a half behind. — Rudyard Kipling
English
1
7
151
8.8K
dsa retweetledi
LiveKit
LiveKit@livekit·
Voice AI has a benchmarking problem. Everyone claims their end-of-turn model is the best, but you couldn't actually compare them. Datasets are proprietary, methods are opaque, and there is no shared ground truth. That changes today. We hit this while developing Turn Detector v1, so we open-sourced eot-bench. 5,000+ real user conversation turns across 14 languages, an evaluation harness that measures the real production tradeoff between latency and false cutoffs, and a live public leaderboard. This should become the default way we evaluate turn detection models.
LiveKit tweet media
English
19
11
110
5.4K
dsa
dsa@dsa·
@tonyzzhao First they ignore you, then they laugh at you, then they fight you, then you win. — Gandhi
English
0
0
3
437
Tony Zhao
Tony Zhao@tonyzzhao·
Reviews of the original ALOHA paper (2023): "It is very hard to find any strength in a paper that forgets about 50 year of development of robotics... the authors are referred to any good book of control theory or robotics to integrate their background... the solution proposed is just damaging the ongoing discussion on low cost-high performance systems."
John Schulman@johnschulman2

PPO: rejected from NIPS 2017

English
16
26
473
68.9K
dsa retweetledi
LiveKit
LiveKit@livekit·
We shipped LiveKit Turn Detector v1. Instead of reading transcripts, it listens to speech directly, combining semantic and acoustic cues into one end-of-turn prediction. The result: high accuracy, low latency—the best model we tested across 14 languages. Available on LiveKit Cloud.
English
16
18
162
11.7K
dsa retweetledi
This Week in AI
This Week in AI@ThisWeeknAI·
"If you had a great tutor, you wouldn't sit down and say please write my paper for me." - @stevenbjohnson The first generation raised on ChatGPT booed AI at graduation. Three founders unpack why, plus Apple's brand new Siri, NotebookLM's biggest update, and whether anyone "wins" the agent race. This week @Jason covered these topics and more with: • Jeffrey Quesnelle (@theemozilla) building Hermes Agent at Nous Research @NousResearch • Steven Johnson (@stevenbjohnson) Editorial Director of NotebookLM @NotebookLM and @GoogleLabs • Russ d'Sa (@dsa) building the voice infra behind ChatGPT, Tesla & Grok at LiveKit @livekit 7 moments worth your time: 1:40 Hermes Agent and the open source race 7:32 Why AI got booed at graduation 10:48 LiveKit powering ChatGPT, Tesla, Grok 27:50 The massive Siri update 52:49 "Functional AGI, unevenly distributed" 57:24 Brittle to brilliant: agents that now just work 1:02:00 Tokenmaxxing and the $1M a year engineer
English
8
13
47
71.1K
dsa retweetledi
LiveKit
LiveKit@livekit·
Operating a robot over the internet means camera frames and joint state arrive at different times, so your observations drift and training data gets misaligned. LiveKit Portal fuses them back together with the same code, whether the robot's in the next room or another continent.
English
1
5
40
4.8K
dsa retweetledi
LiveKit
LiveKit@livekit·
Introducing the LiveKit C++ SDK. Realtime audio, video, and data tracks for C++ apps, with the same low-latency transport our other clients use. Built for the C++ stacks behind robotics, autonomous vehicles, and high-performance media pipelines. livekit.com/blog/livekit-c…
English
2
5
31
1.5K
Snibby
Snibby@ItsSnibby·
@dsa @livekit real-time systems for physical AI have always been bottlenecked by transport layers, so moving beyond audio/video semantics into raw sensor streams feels like an important infra step. @dsa follow back? let’s keep the loop
English
1
0
0
73
dsa
dsa@dsa·
Today @livekit launched Data Tracks. Physical AI and robotics applications need low-latency, realtime transport for data beyond just audio and video. Data tracks let you transmit binary frames from any source: IMUs, LiDAR, RGBD cameras, control systems with no codec overhead and the same low-latency semantics as media. They support full end-to-end encryption and every frame includes a timestamp, so you can easily align data from different sensors. Excited to see what folks build with this! youtube.com/watch?v=Ju9Jz0…
YouTube video
YouTube
English
2
2
12
6.6K
dsa retweetledi
Binh Pham
Binh Pham@pham_blnh·
we won the embodied ai hackathon at @southpkcommons last week this is how we did it 🧵
English
18
5
131
10.5K