Gemini Live API Development Skill
Overview
The Live API enables low-latency, real-time voice and video interactions with Gemini over WebSockets. It processes continuous streams of audio, video, or text to deliver immediate, human-like spoken responses and background reasoning.
Key capabilities:
- Bidirectional audio streaming — real-time mic-to-speaker conversations
- Background reasoning (extended thinking) — multi-step background reasoning with spoken conversational fillers
- Live streaming transcription — real-time speech-to-text with interim and finalized streams
- Video streaming — send camera/screen frames alongside audio
- Text input/output — send and receive text within a live session
- Audio transcriptions — get text transcripts of both input and output audio
- Voice Activity Detection (VAD) — automatic server VAD, client-side Hybrid VAD, and manual Push-to-Talk
- Asynchronous function calling — non-blocking tool execution while audio continues streaming
- Full-session client content — inject and update conversation turns mid-stream
- Session management — context compression, session resumption, GoAway signals
- Ephemeral tokens — secure client-side authentication
[!NOTE] The Live API connects directly via WebSockets. For WebRTC support or simplified integration, use a partner integration.
Models
Current Models (Use These)
gemini-3.8-live— Default option for most low-latency voice agent experiences and real-time dialogue without reasoning delays. Supports interleaved reasoning, asynchronous function calling by default (behavior: NON_BLOCKING), and full-session client content updates.gemini-3.8-live-extended-thinking— High-reasoning audio-to-audio model recommended when higher background reasoning is required during live interactions. Processes background reasoning and async tool calls (behavior: NON_BLOCKINGrequired) while streaming continuous spoken conversational fillers; lifecycle managed viainteraction_status(IN_PROGRESSvsIDLE).gemini-3.5-transcribe-live— Real-time streaming speech-to-text with interim hypotheses, finalized transcripts, smart formatting, and Hybrid VAD.gemini-3.5-live-translate-preview— Real-time speech-to-speech streaming translation across 70+ languages.
[!WARNING] Legacy Models (
gemini-3.1-flash-live-preview,gemini-2.5-flash-native-audio-*,gemini-live-2.5-flash-preview,gemini-2.0-flash-live-001): Readreferences/migration.mdfor breaking protocol changes (behavior: "NON_BLOCKING",thinking_level,interaction_status,send_client_content).
SDKs
- Python:
google-genai>=2.3.0—pip install -U google-genai - JavaScript/TypeScript:
@google/genai>=2.3.0—npm install @google/genai
[!WARNING] Legacy SDKs
google-generativeai(Python) and@google/generative-ai(JS) are deprecated. Never use them.
Partner Integrations
To streamline real-time audio/video app development, use a third-party integration supporting the Gemini Live API over WebRTC or WebSockets:
- LiveKit — Use the Gemini Live API with LiveKit Agents.
- Pipecat by Daily — Create a real-time AI chatbot using Gemini Live and Pipecat.
- Fishjam by Software Mansion — Create live video and audio streaming applications with Fishjam.
- Vision Agents by Stream — Build real-time voice and video AI applications with Vision Agents.
- Voximplant — Connect inbound and outbound calls to Live API with Voximplant.
- Firebase AI SDK — Get started with the Gemini Live API using Firebase AI Logic.
Audio Formats
- Input: Raw PCM, little-endian, 16-bit, mono. 16kHz native (will resample others). MIME type:
audio/pcm;rate=16000 - Output: Raw PCM, little-endian, 16-bit, mono. 24kHz sample rate.
[!IMPORTANT] Use
send_realtime_input/sendRealtimeInputfor all real-time streaming user input (audio, video, and text). On Gemini 3.8 models,send_client_content/sendClientContentis supported across the full session lifecycle with explicit roles (userormodel) to inject conversation context (turn_complete=trueunconditionally interrupts active generation).
[!WARNING] Do not use
mediainsendRealtimeInput. Use the specific keys:audiofor audio data,videofor images/video frames, andtextfor text input.
Quick Start
Authentication
Python
pythonfrom google import genai client = genai.Client(api_key="YOUR_API_KEY")
JavaScript
jsimport { GoogleGenAI } from '@google/genai'; const ai = new GoogleGenAI({ apiKey: 'YOUR_API_KEY' });
Connecting to the Live API
Python
pythonfrom google.genai import types config = types.LiveConnectConfig( response_modalities=[types.Modality.AUDIO], system_instruction=types.Content( parts=[types.Part(text="You are a helpful assistant.")] ) ) async with client.aio.live.connect(model="gemini-3.8-live", config=config) as session: pass # Session is active
JavaScript
jsconst session = await ai.live.connect({ model: 'gemini-3.8-live', config: { responseModalities: ['audio'], systemInstruction: { parts: [{ text: 'You are a helpful assistant.' }] } }, callbacks: { onopen: () => console.log('Connected'), onmessage: (response) => console.log('Message:', response), onerror: (error) => console.error('Error:', error), onclose: () => console.log('Closed') } });
Sending Text
Python
pythonawait session.send_realtime_input(text="Hello, how are you?")
JavaScript
jssession.sendRealtimeInput({ text: 'Hello, how are you?' });
Sending Audio
Python
pythonawait session.send_realtime_input( audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000") )
JavaScript
jssession.sendRealtimeInput({ audio: { data: chunk.toString('base64'), mimeType: 'audio/pcm;rate=16000' } });
Sending Video
Python
python# frame: raw JPEG-encoded bytes await session.send_realtime_input( video=types.Blob(data=frame, mime_type="image/jpeg") )
JavaScript
jssession.sendRealtimeInput({ video: { data: frame.toString('base64'), mimeType: 'image/jpeg' } });
Receiving Audio and Text
[!IMPORTANT] A single server event can contain multiple content parts simultaneously (e.g., audio chunks and transcript). Always process all parts in each event to avoid missing content.
Python
pythonasync for response in session.receive(): content = response.server_content if content: # Audio — process ALL parts in each event if content.model_turn: for part in content.model_turn.parts: if part.inline_data: audio_data = part.inline_data.data # Transcription if content.input_transcription: print(f"User: {content.input_transcription.text}") if content.output_transcription: print(f"Gemini: {content.output_transcription.text}") # Interruption if content.interrupted is True: pass # Stop playback, clear audio queue
JavaScript
js// Inside the onmessage callback const content = response.serverContent; if (content?.modelTurn?.parts) { for (const part of content.modelTurn.parts) { if (part.inlineData) { const audioData = part.inlineData.data; // Base64 encoded } } } if (content?.inputTranscription) console.log('User:', content.inputTranscription.text); if (content?.outputTranscription) console.log('Gemini:', content.outputTranscription.text); if (content?.interrupted) { /* Stop playback, clear audio queue */ }
Background Reasoning (Extended Thinking)
Use gemini-3.8-live-extended-thinking when your voice agent must evaluate complex data, plan multiple steps, or handle long-running tools. The model speaks natural conversational fillers (e.g. "Checking flight options now...") while executing asynchronous tools in the background.
Key requirements:
- Thinking config: Set
thinking_config=types.ThinkingConfig(thinking_level="low")("minimal"|"low"|"medium"|"high"). - Non-blocking tools: All function declarations must set
behavior="NON_BLOCKING". Synchronous blocking mode is not supported and returns an error. - Lifecycle tracking (
interaction_status): Do not rely onturn_complete=Truealone to detect turn completion. Monitormessage.interaction_status(Python) /message.interactionStatus(JS):"IN_PROGRESS": Server is reasoning, speaking conversational fillers, or waiting for async tool responses."IDLE": Server has completed all background reasoning and tool calls; session is ready for user input.
See references/migration.md and the Thinking in Live API Guide for complete Python and JavaScript implementation examples.
Live Translation (Gemini Live Translate)
The Live API supports real-time, low-latency streaming translation of speech (audio) across 70+ languages. For full details on options and capabilities, see the Live Translate Guide.
Model
gemini-3.5-live-translate-preview— The recommended translation model for all Live Translate use cases.
Configuration (TranslationConfig)
To enable translation, specify a TranslationConfig object inside your live session setup:
- Python SDK: Configure the connection using
translation_configonLiveConnectConfig:pythonconfig = types.LiveConnectConfig( response_modalities=[types.Modality.AUDIO], translation_config=types.TranslationConfig( target_language_code="es", # Target language code (e.g. es, fr, pl) echo_target_language=True, ), input_audio_transcription=types.AudioTranscriptionConfig(), output_audio_transcription=types.AudioTranscriptionConfig(), ) - Raw WebSockets: Place
translationConfiginsidegenerationConfig:json{ "setup": { "model": "models/gemini-3.5-live-translate-preview", "generationConfig": { "responseModalities": ["AUDIO"], "translationConfig": { "targetLanguageCode": "es", "echoTargetLanguage": true } } } }
Live Streaming Transcription (Gemini Live Transcribe)
The Live API supports real-time streaming speech-to-text over WebSockets with low-latency interim hypotheses, finalized transcripts, and Hybrid VAD. For full details, see the Live Transcription Guide and Colab Cookbook.
Model
gemini-3.5-transcribe-live
Modes
smart: cleans up filler words, resolves inline self-corrections, and structures formatting.verbatim(default): exact word-for-word transcript.
Python
pythonconfig = types.LiveConnectConfig( response_modalities=["TEXT"], input_audio_transcription=types.AudioTranscriptionConfig(), ) async with client.aio.live.connect(model="gemini-3.5-transcribe-live", config=config) as session: # Stream audio await session.send_realtime_input(audio=types.Blob(data=chunk, mime_type="audio/pcm;rate=16000")) # Hybrid VAD: notify turn end on client-detected silence for zero latency await session.send_realtime_input(audio_stream_end=True)
JavaScript
javascriptconst session = await ai.live.connect({ model: 'gemini-3.5-transcribe-live', config: { responseModalities: ['text'], inputAudioTranscription: { mode: 'smart' } }, callbacks: { onmessage: (msg) => { if (msg.serverContent?.interimInputTranscription) { console.log('Interim:', msg.serverContent.interimInputTranscription.text); } if (msg.serverContent?.inputTranscription) { console.log('Final:', msg.serverContent.inputTranscription.text); } } } }); session.sendRealtimeInput({ audio: { data: chunkBase64, mimeType: 'audio/pcm;rate=16000' } }); session.sendRealtimeInput({ audioStreamEnd: true }); // Hybrid VAD
Raw WebSockets
json{ "setup": { "model": "models/gemini-3.5-transcribe-live", "generationConfig": { "responseModalities": ["TEXT"], "speechConfig": { "voiceConfig": {} } }, "inputAudioTranscription": { "mode": "smart" } } }
Limitations
- Response modality — Only
TEXTorAUDIOper session, not both. Native audio models output audio (response_modalities=["AUDIO"]); enableoutput_audio_transcriptionif you need text transcripts. - Audio-only session — 15 min without compression
- Audio+video session — 2 min without compression
- Connection lifetime — ~10 min (use session resumption)
- Context window — 128k input tokens / 64k output tokens
- Code execution / URL context — Not supported
Upgrading & Migration
For step-by-step migration checklists and protocol deltas when upgrading from gemini-3.1-flash-live-preview, gemini-2.5-flash-native-audio-*, or gemini-2.0-flash-live-001 to Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, read references/migration.md.
Best Practices
- Use headphones when testing mic audio to prevent echo/self-interruption
- Enable context window compression for sessions longer than 15 minutes
- Implement session resumption to handle connection resets gracefully
- Use ephemeral tokens for client-side deployments — never expose API keys in browsers
- Use
send_realtime_inputfor real-time user input (audio, video, text). Usesend_client_contentwith explicituser/modelroles to inject context turns mid-stream - Send
audioStreamEnd/audio_stream_end(Hybrid VAD) when the mic is paused or user finishes speaking - Clear audio playback queues on interruption signals (
interrupted: true) - Process all parts in each server event — events can contain multiple content parts
- Monitor
interaction_status(IN_PROGRESSvsIDLE) when usinggemini-3.8-live-extended-thinkingrather than relying onturn_completealone
Documentation Lookup
When MCP is Installed (Preferred)
If the search_docs tool (from the Google MCP server) is available, use it as your only documentation source:
- Call
search_docswith your query - Read the returned documentation
- Trust MCP results as source of truth for API details — they are always up-to-date.
[!IMPORTANT] When MCP tools are present, never fetch URLs manually. MCP provides up-to-date, indexed documentation that is more accurate and token-efficient than URL fetching.
When MCP is NOT Installed (Fallback Only)
If no MCP documentation tools are available, fetch from the official docs index:
llms.txt URL: https://ai.google.dev/gemini-api/docs/llms.txt
This index contains links to all documentation pages in .md.txt format. Use web fetch tools to:
- Fetch
llms.txtto discover available documentation pages - Fetch specific pages (e.g.,
https://ai.google.dev/gemini-api/docs/live-session.md.txt)
Key Documentation Pages
[!IMPORTANT] Those are not all the documentation pages. Use the
llms.txtindex to discover available documentation pages
- Live API Overview — getting started, raw WebSocket usage
- Thinking in Live API — background reasoning, conversational fillers, interaction_status, non-blocking tools
- Model Card: Gemini 3.8 Live — default low-latency voice agent model & migration guide
- Model Card: Gemini 3.8 Live Extended Thinking — high-reasoning voice model & upgrading guide
- Live Transcription — real-time speech-to-text, interim hypotheses, smart formatting, and Hybrid VAD
- Live Translate — configuration options and capabilities for translation
- Live API Capabilities Guide — voice config, transcription config, VAD configuration, media resolution
- Live API Tool Use — function calling (sync and async), Google Search grounding
- Session Management — context window compression, session resumption, GoAway signals
- Ephemeral Tokens — secure client-side authentication for browser/mobile
- WebSockets API Reference — raw WebSocket protocol details
- Migration & Upgrading Guide — step-by-step checklists and code examples for Gemini 3.8 Live and Extended Thinking
Supported Languages
The Live API supports 70 languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Hindi, Arabic, Russian, and many more. Native audio models automatically detect and switch languages.

