Google DeepMind has released Gemini 3.8 Live and 3.8 Live Extended Thinking, two new audio models accessible via the Gemini API and Google AI Studio. Designed for voice agents, the models support background API calls, concurrent visual processing, and over 97 languages. The Extended Thinking variant currently leads the Artificial Analysis Speech-to-Speech Leaderboard with an 82.6 percent score.
Google is pricing the input at $0.005 per minute and output at $0.018 per minute for audio, which calculates to roughly $1.38 per hour of voice conversation. This undercuts OpenAI’s competing GPT-Live-1 model, which costs at least $3.00 per hour based on its $0.05 per minute rate.
While Google offers a cost advantage, OpenAI’s model maintains an edge in conversational quality due to full-duplex capabilities that enable simultaneous listening and speaking. Developers can access sample applications for Google’s new models on GitHub.
Why it matters
Dramatically lowers operational costs for developers building real-time voice and multimodal AI applications.
Pushes competitive pressure on OpenAI by topping speech leaderboards while undercutting API pricing.
Enables complex voice agents capable of executing background tool calls during live user interactions.
Source: the-decoder.com



