Gemini Live audio
Summary
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new speech-to-speech models (AI systems that convert spoken words into responses) similar to OpenAI's GPT-Live family. A web UI was created that lets users select a model and voice, enter an optional system prompt (instructions given to the AI before interaction), and have voice conversations through a browser, including the ability to interrupt the model while speaking. The implementation connects to Google's WebSocket endpoint (a two-way communication channel between a browser and server) and uses Web Audio API (a tool for capturing and playing back sound in browsers) for audio capture and playback.
Classification
Affected Vendors
Related Issues
Original source: https://simonwillison.net/2026/Sep/15/gemini-live/
First tracked: September 15, 2026 at 08:01 PM
Classified by LLM (prompt v3) · confidence: 85%