Real-time conversational AI with a 3D avatar that speaks back — lip-synced and streamed end-to-end.
talk2avatar is a full-stack web app that lets you have a real-time conversation with a 3D VRM avatar. You type or speak, an LLM responds with streaming tokens, each sentence is synthesized into speech with viseme timing data, and the avatar's lips move in sync — all pipelined with sub-second time-to-first-audio.
Voice / Text --> LLM (streaming) --> TTS + Visemes --> 3D Avatar Lip Sync
input Groq / Cerebras Kokoro 82M VRM + Three.js
/ OpenRouter (HeadTTS) (React Three Fiber)
- Streaming LLM — Token-by-token responses via Groq, Cerebras, or OpenRouter (switchable at runtime)
- Sentence-level TTS pipelining — Sentences are split from the token stream and synthesized in parallel so speech starts before the full response arrives
- Real-time lip sync — Oculus-standard viseme weights applied per-frame to VRM blend shapes via
expressionManager - Dual TTS engines — Server-side ONNX Runtime for speed, client-side WASM fallback for resilience, browser
speechSynthesisas last resort - 3D avatar viewer — VRM model loading with idle blink animation, cinematic camera reveal, and orbit controls
- Speech input — Browser-native speech-to-text via Web Speech API
- Responsive glass UI — Glass-morphism chat panel, provider tabs, TTS loading progress, animated status indicator
| Layer | Technology |
|---|---|
| Framework | Next.js 16 (App Router, Node.js runtime) |
| UI | React 19, Tailwind CSS 4 |
| 3D Rendering | Three.js r182, React Three Fiber, @pixiv/three-vrm |
| LLM | Vercel AI SDK with OpenAI-compatible providers (Groq, Cerebras, OpenRouter) |
| TTS | HeadTTS — Kokoro 82M ONNX model, server + WASM dual-engine |
| State | Zustand |
| Speech Input | Web Speech API |
| Deployment | Vercel |
- Node.js 20+ (LTS recommended for stable native ONNX Runtime bindings)
- API key for at least one LLM provider: Groq, Cerebras, or OpenRouter
- Modern desktop browser with microphone support
git clone https://github.com/pradhankukiran/talk2avatar.git
cd talk2avatar
npm installcp .env.example .env.localAdd your API keys in .env.local:
GROQ_API_KEY=gsk_...
CEREBRAS_API_KEY=csk-...
OPENROUTER_API_KEY=sk-or-...npm run devOpen http://localhost:3000. Click the microphone or type a message to start chatting.
First run may take longer while the TTS worker initializes and downloads voice model assets.
src/
app/
api/
chat/route.ts # Streaming LLM endpoint (Groq / Cerebras / OpenRouter)
tts/route.ts # Server-side TTS synthesis endpoint
page.tsx # Main page layout
layout.tsx # Root layout with fonts
components/
AvatarViewer.tsx # Three.js canvas, camera reveal, orbit controls
VrmModel.tsx # VRM loading, idle blink, per-frame lip sync
ChatPanel.tsx # Message list, provider tabs, status
ChatInput.tsx # Text input + microphone button
hooks/
useChat.ts # Full pipeline: LLM stream -> sentence split -> TTS -> audio queue
useTTS.ts # Server / client WASM / browser TTS fallback chain
useLipSync.ts # Per-frame viseme weight interpolation
useSpeechRecognition.ts # Web Speech API wrapper
lib/
audio-queue.ts # FIFO audio playback with binary-search viseme tracking
sentence-splitter.ts # Streaming token accumulator with sentence boundary detection
viseme-mapping.ts # Oculus viseme -> VRM/RPM blend shape weight tables
headtts-client.ts # Browser-side HeadTTS singleton (WASM engine)
audio-context.ts # Shared AudioContext singleton
server/
headtts-worker.ts # Node.js worker thread managing HeadTTS ONNX inference
stores/
app-store.ts # Zustand global state
types/
index.ts # Shared type definitions
headtts.d.ts # HeadTTS type declarations
| Variable | Required | Description |
|---|---|---|
GROQ_API_KEY |
* | Groq API key |
CEREBRAS_API_KEY |
* | Cerebras API key |
OPENROUTER_API_KEY |
* | OpenRouter API key |
HEADTTS_DEBUG |
No | 1 to enable server-side TTS debug logs |
NEXT_PUBLIC_TTS_DEBUG |
No | 1 to enable client-side TTS debug logs |
* At least one LLM provider key is required.
- Server-side TTS runs on CPU via
/api/ttsusing ONNX Runtime - Set
HEADTTS_DEBUG=1for detailed server worker and device timing logs - Set
NEXT_PUBLIC_TTS_DEBUG=1for client pipeline and TTS timing in browser console - If you see
Module did not self-registerorImporting modules failed, inspect the precedingHeadTTS Workererror line for the root cause (module resolution vs native addon)
The TTS route runs on the Node.js serverless runtime. ONNX Runtime and HeadTTS assets are bundled via outputFileTracingIncludes in next.config.ts.