kwindla

6.5K posts

kwindla banner
kwindla

kwindla

@kwindla

Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai

San Francisco, CA Katılım Eylül 2008
3.9K Takip Edilen14.8K Takipçiler
kwindla
kwindla@kwindla·
The LangChain team has built some nice LangSmith tracing/observability integrations with voice AI orchestration frameworks and the speech-to-speech APIs from OpenAI and Google. LangChain pioneered a lot of early and influential tooling for text agents. I often talk to product teams that built their text agents with LangChain and are interested in leveraging that knowledge and code for voice agents. It's great to see LangSmith expanding to support voice as a first-class modality.
LangChain@LangChain

Voice agents are exploding. Don’t let them be a black box in production. Today, we’re launching LangSmith tracing for 4 voice frameworks: 🎙️ @pipecat_ai 🎙️ @livekit 🎙️ @OpenAI Realtime 🎙️ @GeminiApp Live (Google ADK) Learn more: langchain.com/blog/trace-voi…

English
2
6
28
8K
kwindla
kwindla@kwindla·
@leonho I like this extension of the analogy!
English
0
0
0
43
Leon Ho
Leon Ho@leonho·
@kwindla The thread analogy holds all the way down. If the loop is the new thread, the shared context is the new shared memory, and that's always where the hard bugs lived. Two subagents writing it at once is a race, and non-blocking is exactly what makes it fast and hard to reproduce.
English
1
0
2
44
kwindla
kwindla@kwindla·
Lots of new stuff in recent Pipecat releases. There are 245 items in the CHANGELOG.md entries for the three releases so far this summer. Pipecat is the most widely used voice and video AI framework, and supports a very large range of use cases. So opinions will differ about what's most important in each release! To me, the three biggest new, new things in Pipecat recently are: 1. Subagents 2. Folding Pipecat Flows into core 3. A new behavior evals framework Subagents. We built Pipecat to power realtime, human-in-the-loop AI software. The first-generation voice agents and voice assistants that showed people what's possible with LLMs+voice are just the beginning. Most of our customers' production deploys today, and all of our experiments and side projects, now use multiple LLMs and multiple inference loops. The LLM is the new function. The inference loop is the new thread. The new `TaskManager` and `Workers` abstractions in Pipecat provide support for starting, stopping, and managing these subagent "threads." Subagents share LLM context using a message bus. The bus can be local or can serve multiple agents running on different machines. Delegating work to subagents keeps the conversational voice loop fast, relatively simple, and non-blocking. Common uses of subagents include managing communication with legacy enterprise systems, implementing guardrails, and dynamically generating on-screen user interfaces. Pipecat Flows. One of my mantras is that everything is a state machine. Many production voice agents are explicitly designed as state machines. Conversation phases with well-defined transitions between them. Think about the different "jobs to be done" during a healthcare appointment pre-visit call: authentication, time/date confirmation, confirmation of reason for the visit, pre-existing conditions and current medications, etc. At each transition point, we swap out the system instruction, generate a new tools list, and do a (non-blocking) contextually aware summarization of the context history. Pipecat Flows is a set of helper modules for implementing these flexible but robust conversational state machines. When people are having trouble getting their complex voice agents to score well on "task success" evals, the single most powerful lever is Pipecat Flows. New evals support. We've always worked closely with evals and devops platforms, building strong support in Pipecat for all of the widely used tools in this space. And we strongly recommend that every enterprise that is deploying voice agents at scale explore the evals/tooling landscape and pick a vendor to work with. But early in the development process, you get huge leverage from simple automation of smoke tests, end-to-end tests, and basic behavior testing. If you have a lot of time in the saddle building both voice agents and working with coding agents, you know how to prompt these automations into existence. The new Pipecat evals framework makes this easy for new developers, too, and makes this workflow trivially "visible" to coding agents. I'll leave you with a little bit of history, because I've been talking about this a lot, lately. A bunch of us who work on Pipecat have been building this "new new" human-in-the-loop stuff together for a long time, across two startups. We worked on the user interfaces in Minority Report and Iron Man. LLMs+voice are giving us the final, missing pieces we need to build the things we've been dreaming about and working on since ~2005 or so. Here's a demo video from 2012. Pre-LLMs. Pre-Pipecat. But, for us, the same long project!
English
17
2
56
4.7K
kwindla
kwindla@kwindla·
Yeah, that’s what it felt like to me, at the time. We could build amazing things but not at a price point that allowed mass adoption. Our average contract value was about $1m. We didn’t want to compromise on functionality. We were committed to our multi-user, multi-screen, multi-device, gestural, spatial, full environment, fully networked vision. But now we can build all of that with commodity hardware. And LLMs and deep learning in general make a bunch of new things possible even beyond that original vision. I gave a talk about this at AI Engineer World’s Fair earlier this month. The history of computing, and how what we can do next with LLMs fits into that history. I think the video of the talk will be up soon. (cc @swyx)
English
0
0
0
44
electronic Max
@kwindla @medialab @PARCinc honestly i think oblong was too far ahead of its time. when people say 'what will interacting with AI feel like' it certainly won't be prompting - i will point them to oblong & things like MS Live Labs photosynth, another one too early–
English
1
0
0
39
kwindla
kwindla@kwindla·
@emax @medialab @PARCinc We had so, so much fun building the Oblong stuff. It was also an incredibly valuable experience for me to climb up the learning curve of selling complex, new tech to enterprise customers and then supporting deployments at scale.
English
1
0
1
42
electronic Max
@kwindla @medialab extremely cool :) i was blown away by oblong's demos, so cool. absolute respect. they were more beautiful & human than our full body tracking ubicomp work at the AI lab ('Intelligent Room') after I left the @medialab - or acoustic tracking at xerox @PARCinc
English
1
0
0
26
kwindla
kwindla@kwindla·
@emax @medialab I founded Oblong with John. I was at the Media Lab from ‘96 to ‘99 (Mitch’s group). John is doing his own thing. But a bunch of former Oblong people work on Pipecat!
English
1
0
1
81
electronic Max
@kwindla wait, you were at oblong? with john underkoffler? is he now at pipecat? send my regards. i was at the @medialab with him. :)
English
1
0
1
98
kwindla
kwindla@kwindla·
@boardyai (And by you, I mean you, a most excellent automation.)
English
1
0
0
119
kwindla
kwindla@kwindla·
@boardyai Of course! Would love to. Will DM you.
English
2
0
2
170
Bin Yang
Bin Yang@binyangderek·
@kwindla @chenwanch1 Latency & intelligence of s2s models are catching up very soon (with systems like GPT-Live, and more to come later this year), the flexibility part (generalised interactivity behavior) seems very under-explored.
English
1
0
0
67
William Chen
William Chen@chenwanch1·
I love how every time an audio LLM is released, there are always posts about how it will finally “break the status quo” of “ASR+LLM+TTS” cascades. Are influencers outdated about speech? Or are E2E models just not having the expected industry impact? My bet is on the latter 🤪
English
6
1
25
3.2K
kwindla
kwindla@kwindla·
> The observability point carries a lot of weight. It's vastly easier to build evals for pipelined system design. 100% agree, and this is hard for people to wrap their heads around until they are supporting and debugging voice agents in production. Devops and managing AI systems at scale is still very new.
English
1
0
1
83
kwindla
kwindla@kwindla·
@itzik009 Agreed that multi-turn, voice use cases are very hard for today's models. We weren't happy with any of the existing benchmarks, so we developed benchmarks specifically to test multi-turn, conversational performance. github.com/kwindla/aiewf-…
English
1
0
1
54
Timmy Ghiurau
Timmy Ghiurau@itzik009·
@kwindla Voice is the hardest memory test. NTU and NUS frontier evals found 17 to 19% degradation within a single 30 minute conversation, and voice sessions run long. Retrieval is solved. The controller, what gets written and trusted, is the open problem.
English
1
0
0
58
kwindla
kwindla@kwindla·
Hacking on fully local realtime voice and video AI is hugely fun. Highly recommend. Building voice agents teaches you a lot about the hardest design, orchestration, and infra challenges: subagents, multi-model data pipelining, sharing and optimizing context, memory and continual learning, optimizing TTFT and throughput situationally, observability.
Daniel | software engineer@0xfelaback

I got inspired seeing @kwindla post on his local voice agent and decided to see if I could set up mine fully local using Pipecat also. ended up scaffolding a RAG voice agent pipeline running entirely on a 2020 m1 macBook with 16gb RAM. Read more below; 🧵

English
5
3
28
2.9K
kwindla
kwindla@kwindla·
New realtime transcription model from the team at @cartesia. A very large number (and wide variety) of production voice agents use Cartesia text-to-speech models. The company's voice models score very high in both human and quantitative testing. Now there's a similarly high-scoring speech-to-text model from Cartesia, too.
Cartesia@cartesia

For voice agents, STT has to nail three things - accuracy, turn detection, and latency. If any one falls short, the experience breaks down: the agent misunderstands, interrupts, or just feels slow. We built Ink-2 to lead on all three. Here’s how it stacks up against other providers: link in comments.

English
2
6
79
11.7K
kwindla
kwindla@kwindla·
I feel the same gap in my skill/tooling right now that @vr000m highlights. My general approach is plan, refine the plan, iteratively implement the steps in the plan. But during implementation, it's harder to see and understand the whole picture than it feels like it should be. This feels like a combination of not yet being fully fluent with these new tools, and of gaps in the current generation of agentic coding harnesses. I'd like the harnesses to give me more visualization artifacts, progress tracking in a standard format that I can get used to relying on, and a better blend of high- and low-level status updates.
Varun Singh@vr000m

Over the weekend, flying SF-HEL, I had Fable execute a dev-plan. The implementation and verifier steps use cheaper models. only the review models run Opus on xhigh. The plan was vetted and modularizes the game-collector from the game alerts repo On its own, it found an issue and iterated on a fix. Fable delegated implementation and verification to subagents, then reset to re-review, found a new related issue, fixed it, and repeated. This ran 15 review iterations before it whittled the problems down. This iteration blew away half my fable budget and by Sunday 48h after the reset: 20% usage left, 4 days to go. Good thing, codex reset yesterday, so no productivity loss. On the topic of Fable, Two things I want to improve. First, I need a way for the model to step back and look at feedback holistically instead of chipping at it iteration by iteration. Second, I need to re-review the original plan. There was likely ambiguity that only surfaced during implementation. Normally that would come back to me via the AskUserQuestion tool, but with no human in the loop, Fable assumed it was a straightforward call and proceeded. Both exciting and a clear signal that my plan docs and implementation/review skills need to get sharper. PR: github.com/vr000m/gameale… Skills: github.com/vr000m/skein

English
1
0
4
1.7K
Danielle Fong 🔆
Danielle Fong 🔆@DanielleFong·
yay raised more money the journey of lightsaber powerplant continues
English
31
5
360
12.1K
kwindla
kwindla@kwindla·
@jheitzeb Joe, that means so much to me. I really, really admire how you've built the AI Tinkerers community to be so inclusive, global, and incredibly fun and interesting.
English
0
0
5
216
kwindla
kwindla@kwindla·
I just gave a talk about this at the AI Engineer World's Fair. Like a lot of people, I'm making things with LLMs that feel completely, mind-expandingly new, and at the same time revisiting work from earlier in my career and implementing things we dreamed of but couldn't quite do at the time. When the Iron Man movie your screenshot is from was in pre-production, Robert Downey Jr. came to our office to practice using the gestural interfaces. (The movie UI was based on our work.) Here's our demo video from 2012.
English
5
1
37
3.6K
will brown
will brown@willccbb·
there will soon come a time, perhaps next year, when voice interaction + multimodal reasoning models are good enough and fast enough that you can actually just program like this
will brown tweet media
English
57
23
617
31.5K