$ cat /etc/cookies.conf
We use cookies to understand how people use this site.
Analytics cookies help us improve your experience.
They are off by default. Nothing tracks you until you say so.
$ select cookie_preferences
Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Turn detection is a major challenge for conversational voice AI. I’ll talk about training a new open source, open code, open data “smart turn detection model.” I’ll touch on model architectures for audio processing tasks, the latest developments in this space from the big labs (OpenAI Realtime API and Gemini Multimodal Live API) and startups, creating data sets, evaluating training runs, deploying a small model like this at scale in production, and how anyone can get involved.
Wav2Vec2 transformer performs 14-language conversational audio turn endpoint prediction using PyTorch.
Open-source native audio ML model for human-like conversational turn detection.
Loading recent emails...