Turn emails into revenue with Brew. No credit card, free credits to try.
livekit.io · newsletter
Explore this email design and adapt it to your own brand. Review the copy, links, and offer before sending.
Hey LiveKit devs,
We’ve got this month’s roundup for you, let’s get into it.
Product Changelog
Answering Machine Detection (AMD)
LiveKit Agents can now figure out who (or what) actually picks up an outbound call. Answering machine detection (AMD) classifies recipients as human, voicemail, IVR, or unavailable within the first second of the call so your agent can adapt accordingly: start the conversation, leave a voicemail, navigate the IVR menu, or hang up and retry.
AMD is available in Python 1.5.9+ and TypeScript 1.4.2+ and surfaces classifications in Agent Console and Agent Observability as distinct events. Learn more in our blog or watch Shayne’s walkthrough to see it in action.
Answering Machine Detection with LiveKit Agents
Custom voices in LiveKit Inference
You can now clone a voice on LiveKit Inference across multiple TTS providers. A custom voice gives your agent character that fits your brand and the specific job it’s doing. Record or upload an audio sample and get a custom voice ID to use with your agents.
If your TTS request fails mid-call, LiveKit Inference automatically falls back to the same voice on another provider so the call can continue with minimal disruptions instead of switching to a different default voice. Voice cloning is available today on all paid plans on LiveKit Cloud. Read our announcement blog and watch Jesse’s video to learn more.
Clone Your Voice Once, Deploy to Multiple TTS Providers
Structured data collection in LiveKit Agents
Need your agent to capture lead details, appointment info, or any other structured output during a call? Define a schema and your LiveKit agent will collect that data over the course of the conversation, prompting for missing fields, validating answers, and emitting a clean JSON record at the end.
Structured data collection is available in code using Tasks and TaskGroups, or in your browser using Data collection mode in Agent Builder. Watch Jesse’s walkthrough to see an example and read the blog for more details.
Build a Voice Agent That Captures Structured Data with LiveKit
Content & Community Highlights
Add agent guardrails with the observer pattern
Most voice agent guardrails live in the system prompt, where they compete with everything else the LLM has to track. Our latest deep dive walks through using the observer pattern in LiveKit Agents to run guardrails as a separate process, catching prompt injection, off-topic responses, and policy violations in real time without slowing down your main loop. Check out Shayne’s video for the full breakdown:
LiveKit Agent Observer Guardrails
Run a full voice pipeline on xAI
With xAI's new STT model, you can now run an entire voice agent pipeline on xAI (STT, Grok, and TTS) through LiveKit Inference with a single API key. Watch Jesse's video on why a cascaded pipeline still wins over a realtime model when you need control, debuggability, and full visibility at every stage, plus a live demo of Grok's built-in tools.
Build a Full xAI Voice Agent with Grok, STT (NEW), and TTS
Prompting Gemini 3.1 Flash TTS
Gemini 3.1 Flash TTS is one of the most prompt-controllable TTS models available today. With the right instructions, you can shape pacing, tone, emotion, and even multi-speaker delivery. Read our prompting guide and watch Shayne’s video for techniques that work especially well for voice agents.
The Gemini TTS Hack Nobody's Talking About
ai-coustics Quail Voice Focus 2.1
We've upgraded the ai-coustics plugin to Quail Voice Focus 2.1, which improves speaker isolation in noisy environments and tightens STT accuracy on real-world conversations. The realtime audio enhancement model helps filters out background noise and competing voices so your STT only hears the person actually talking to your agent. If you're already using the plugin, you'll get the new model automatically.
Inworld Realtime TTS-2
Inworld Realtime TTS-2 is now available through LiveKit Inference, Agent Builder, and the Agents SDK plugins. The new version brings improved latency and more expressive voice control, allowing your agents to sound more natural by mirroring the tone and pacing of prior turns in the conversation.
Add a face to your voice agent with LiveAvatar
The LiveAvatar plugin from HeyGen lets you add a realtime avatar to a LiveKit voice agent without rebuilding your conversation loop. Watch our demo to see how it comes together in a few lines of code.
How to Build an AI Avatar Agent with HeyGen
LiveKit in the Wild
Building voice agents with Gemini Live
Jesse was featured on the Google for Developers channel to walk through building a voice agent with the Gemini Live API and LiveKit. He covers project setup, native audio, server-side and client-side tool calling, Google Search integration, and live multilingual switching.
Building LiveKit Agents with Gemini Live API
Open source voice agents with Simplismart
Missed last month's webinar with Chris and the Simplismart team on building voice agents with open source models? Watch the recording to learn how to stand up a complete voice agent end-to-end using Whisper v3 Turbo for STT, Gemma 3 4B as the LLM, and Qwen for TTS.
InferSights: Build Real-Time AI Voice Agents with LiveKit + Simplismart
Voice AI Speakeasy in NYC
Earlier this month, we co-hosted the Voice AI Speakeasy with Cartesia at Madame George during AI Agent Week in NYC. It was great to see over 100 people show up for well-poured drinks and candid conversation. Thanks to everyone who joined, and stay tuned for more of these throughout the year.
LiveKit Dev Roundtable
📅 Thursday, May 28 at 1:00 PM PT (4:00 PM ET / 9:00 PM BST)
📍 Virtual (Google Meet)
Following the success of our first session, the LiveKit Dev Roundtable is back, this time scheduled at a US-friendly time. The first half hour covers updates from our DevRel team, including AMD and voice cloning, followed by an open community discussion. Come prepared with a feature you're building, a problem you're stuck on, or just listen in.
Microsoft Build
📅 Tuesday, June 2 – Wednesday, June 3
📍 Fort Mason Center, San Francisco
Catch us at Microsoft Build, Microsoft's AI-first developer festival in San Francisco. Stop by booth G221 to chat with the team and check out some cool voice agent demos in action. We'll be showing off the latest agent and platform capabilities with our DevRel and engineers on hand to talk through your use case.
Happy hacking!