Turn emails into revenue with Brew. No credit card, free credits to try.
livekit.io · newsletter
Explore this email design and adapt it to your own brand. Review the copy, links, and offer before sending.
Hey LiveKit devs,
We’ve got this month’s roundup for you, let’s get into it.
Product Changelog
A new audio-based LiveKit Turn Detector
We shipped v1 of our LiveKit Turn Detector. Instead of reading the transcript, it listens to the audio directly, using cues like intonation and pacing that aren’t captured in text. That extra signal helps your agent decide when the user has ended their turn.
Because it works on the audio, the model handles the natural pauses that text-based detection tends to misread, like someone trailing off mid-thought or pausing to gather what they want to say next. It runs inside your AgentSession alongside your existing STT and VAD with no other pipeline changes, works across languages, and is available now in both Agents SDKs (Python 1.6.1 and TypeScript 1.4.7).
It's also the most accurate model we tested. On our open benchmark, v1 leads the latency-versus-interruption tradeoff at every budget: at 600ms it hits a 4.5% false-cutoff rate and stays on top across 14 languages. We're also shipping v1-mini, a smaller model with the same architecture tuned for fast CPU inference, plus eot-bench, an open benchmark suite and datasets so you can evaluate end-of-turn models yourself.
Watch Jesse’s tutorial or read the blog for the full results and an interactive report.
The End of Annoying AI Interruptions? LiveKit Turn Detector v1 Tested
Add a voice agent to any website
Add a voice agent to any website with a single code snippet. The Agent Embed Widget is a lightweight interface that runs as a floating button or an inline component on any web page or mobile app, and every agent you deploy on LiveKit Cloud ships with a hosted widget ready to go.
Configure the widget in the LiveKit Cloud dashboard. You can enable options like camera, screen share, and text chat, as well as branding elements like accent color, icon, and theme. Read the blog, watch the tutorial, or use our web embed starter app to customize it further.
Add a Voice Agent to Any Website with One Script Tag
Control any robot from the cloud
LiveKit Portal lets you run an agent in the cloud that remotely controls a robot with no specialized inference hardware on the machine itself. The policy runs wherever the GPUs are, the robot operates where you need it to, and Portal handles the realtime link between them, whether the robot's in the next room or on another continent.
Under the hood, Portal reconstructs clean, time-aligned observations from camera and joint-state streams that would otherwise drift apart over the network, and adds multi-operator sessions for data collection, clean policy-to-human handoff, and live metrics down to true observation-to-action latency.
Portal is available today with pip install livekit-portal. Read more in the blog, watch Binh's demo video, and build your own human-in-the-loop training pipeline with our tutorial.
Can Humans Teach Robots to Think? Real-Time AI Control
LiveKit C++ SDK
Version 1.0 of our C++ SDK is now available, bringing LiveKit to native desktop, server, robotics, and embedded applications. Built on the same Rust core as our other clients, it accepts raw audio and video frames directly from your application, so you can connect resource-constrained hardware and devices to LiveKit rooms.
The easiest way to get started is with a prebuilt release and the CMake helper in our example collection. Check out the blog and see the C++ quickstart for the full installation flow.
Content & Community Highlights
Multilingual speech-to-text with NVIDIA Nemotron 3.5 ASR
NVIDIA's Nemotron 3.5 ASR is a 600M-parameter streaming speech recognition model that transcribes 40 language-locales and is small enough to run on a laptop. Shayne put it inside a LiveKit voice agent and built a local teleprompter that scrolls under your voice in whatever language you're reading.
Read his blog and install the NVIDIA Riva STT plugin to use it in your own pipeline.
Where to apply noise cancellation in your voice pipeline
Real microphones pick up fans, music, and the person talking three feet away, and all of it degrades transcription and turn detection. Our latest guide covers how noise cancellation works in LiveKit, where to apply it (in the agent, the client, or the SIP trunk), and how to choose between background noise suppression and voice isolation across the Krisp and ai-coustics models.
Build a RAG agent with MongoDB Atlas Vector Search
Most voice agents can’t remember you. Our new tutorial shows how to give your voice agent access to a private knowledge base using MongoDB Atlas Vector Search. Use vector search as a tool the LLM can call mid-conversation to retrieve relevant context and ground its answers.
Watch Jesse's tutorial for a step-by-step walkthrough, or dive deeper into our docs on external data and RAG for the patterns behind it.
Give Your LiveKit Voice Agent Memory with MongoDB Atlas Vector Search
Bring your LangChain agents to LiveKit
The LangChain integration lets you bring graph-based LangChain agents into a LiveKit voice agent as the LLM, whether that's a LangGraph workflow you've compiled yourself, a prebuilt agent from create_agent, or a deep agent. Wrap it with theLLMAdapter and the rest of your agent, tools, graph structure, and prompts stay exactly as they are while LiveKit drives it with voice instead of text.
Read the blog for a complete, runnable example, and watch Darryn’s video for a walkthrough.
Give Your LangChain Agent a Voice
Expressive avatars with Runway Characters
The best avatars feel engaged while your users are talking, not just when they're speaking. Runway Characters make eye contact, small head movements, and expressions that respond to the conversation as it happens, all generated live from a single reference image. Watch how the avatars behave in Jesse's video, and add one to your agent with the new LiveKit Agents plugin.
Real-Time AI Avatar Voice Agent with LiveKit and Runway Characters
LiveKit in the Wild
Realtime translation with Gemini 3.5 Live Translate
We put together a demo for Google Deepmind that shows live, multilingual translation built on the Gemini Live API and LiveKit. Gemini 3.5 Live Translate streams video and audio natively within its realtime protocol, making it a strong fit for translation experiences that need to keep up with a live conversation. Watch Jesse explain how he built it on X, see a live demo in Google's launch video, and check out the full open-source example to build your own.
Introducing Gemini 3.5 Live Translate
Customer spotlight
We published three new stories on how teams ship voice and realtime AI on LiveKit. SAP is partnering with LiveKit to bring realtime voice to Joule, its AI copilot, putting natural, spoken interactions in front of enterprise users at scale. telli scaled to 5 million calls with enterprise customers like Sky by pairing LiveKit with the ai-coustics Quail Voice Focus model to keep short utterances and turn-taking reliable on real-world calls. And Casper Studios built the Stranger Things × Doritos interactive telethon on LiveKit Cloud, running a multi-agent system that handled 400,000+ live sessions last fall.
Doritos Telethon For Hawkin
Contact Center Week
📅 Monday, June 22 – Thursday, June 25
📍 Caesars Forum, Las Vegas