Current environment: gemini-3.8-flash-tts greeting audio is verified in all three demos.
This recipe uses the published Agora Agent Kit SDK v2.11.0.
The pipeline is GeminiSTT → Gemini gemini-3.6-flash → Gemini 3.8 Flash TTS.
All three stages use server-only GOOGLE_API_KEY. AgentSession automatically uses
agora-feature: gemini-live and the preview endpoint for start and subsequent
session operations. Keep the retained session for stopping the agent.
Set these optional values in server/.env.local:
GEMINI_TTS_MODEL=gemini-3.8-flash-tts
GEMINI_TTS_VOICE=Puck
GEMINI_TTS_STYLE="warm and reassuring"Use GEMINI_TTS_MODEL=gemini-3.8-flash-tts and restart the backend after configuration changes.
SDK dependencies are pinned to v2.11.0; no sibling SDK checkout is required.
Gemini TTS remains a preview provider within the released SDK.
Build and run a real-time voice agent with GeminiSTT, Gemini 3.6, and Gemini TTS preview.
This project includes a Next.js web client and a Python FastAPI backend. The browser connects to Agora RTC and RTM, while the backend creates and manages the Conversational AI agent session.
Microphone -> GeminiSTT -> Gemini 3.6 LLM -> Gemini TTS -> Browser
Gemini ASR, Gemini LLM, and Gemini TTS use the same Google API key. Gemini TTS uses the same server-only Google API key through the preview SDK.
- Python 3.10 or newer
- Bun
- Agora CLI
- An Agora project with Conversational AI access enabled
- A Google API key with access to Gemini
git clone https://github.com/AgoraIO-Community/Gemini-Agora-Voice-Agents-Python.git
cd Gemini-Agora-Voice-Agents-PythonSkip installation if agora is already available on your path.
curl -fsSL https://raw.githubusercontent.com/AgoraIO/cli/main/install.sh | sh -s -- --add-to-path
agora login
agora project use <project-id-or-name>Install dependencies and create server/.env.local from the example:
bun run setupWrite the Agora App ID and App Certificate from the selected Agora project:
agora project env write server/.env.localOpen server/.env.local and add your Google API key:
GOOGLE_API_KEY=your_google_api_keyKeep server/.env.local private. It is ignored by Git and is never sent to the browser.
bun run doctor:local
bun run devOpen http://localhost:3000 and select Start conversation.
| Service | URL |
|---|---|
| Web client | http://localhost:3000 |
| FastAPI backend | http://localhost:8000 |
| FastAPI docs | http://localhost:8000/docs |
The backend reads configuration from server/.env.local.
| Variable | Required | Description |
|---|---|---|
AGORA_APP_ID |
Yes | Agora project App ID. |
AGORA_APP_CERTIFICATE |
Yes | Server-only certificate used to create RTC and RTM tokens. |
GOOGLE_API_KEY |
Yes | Google API key used by GeminiSTT, Gemini 3.6, and Gemini TTS. |
AGENT_GREETING |
No | Overrides the default opening message. |
GEMINI_TTS_MODEL |
No | TTS model; defaults to gemini-3.8-flash-tts. |
GEMINI_TTS_VOICE |
No | Default Gemini TTS voice; defaults to Puck. |
GEMINI_TTS_STYLE |
No | Optional description of the speaking style. |
PORT |
No | FastAPI port. Defaults to 8000. |
The source template is server/.env.example.
Choose from all 30 Gemini voices before starting; Puck is the default. The selected
voice applies to that session. API callers that omit ttsVoice use
GEMINI_TTS_VOICE or Puck. End the conversation to choose another voice.
The prompt describes the Gemini ASR/LLM/TTS pipeline and selected voice and model. It permits occasional performance cues. The transcript view hides known cues in agent messages, including incomplete streamed cues; raw transcript events and TTS input stay unchanged. The Show cues toggle displays cues in the transcript. Streaming cue interpretation by the preview TTS has not been verified.
The agent introduces itself as Gemini in the prompt and default greeting.
- The browser requests a channel, UID, and RTC/RTM token from FastAPI.
- The browser joins Agora RTC and RTM and publishes microphone audio.
- FastAPI starts a Conversational AI agent in the channel.
- GeminiSTT transcribes the user, Gemini 3.6 generates the response, and Gemini TTS produces the agent audio.
- Transcript, agent state, and pipeline metrics are delivered to the browser over RTM.
The browser uses stable /api/* paths. In local development, Next.js rewrites those requests to FastAPI through AGENT_BACKEND_URL.
bun run setup # Install dependencies and create server/.env.local
bun run doctor:local # Check local prerequisites and environment
bun run dev # Run FastAPI and Next.js together
bun run verify # Run the verification suiteRun bun run doctor:local and confirm that AGORA_APP_ID, AGORA_APP_CERTIFICATE, and GOOGLE_API_KEY are non-empty. Also confirm that the selected Agora project has Conversational AI access enabled.
Confirm that FastAPI is listening on port 8000. When running the web client separately, set:
cd web
AGENT_BACKEND_URL=http://localhost:8000 bun run devCheck the browser console for the agent connection and confirm that microphone permission was granted. The pipeline panel displays the latest Gemini ASR, Gemini LLM, and Gemini TTS latency metrics when those events arrive.
web/ Next.js voice conversation UI
server/ Python FastAPI and Agora Agent Server SDK integration
server/.env.example Safe configuration template
ARCHITECTURE.md Runtime architecture and request flow
docs/ai/RECIPE.md Implementation recipe and design constraints
Released under the MIT License.