Skip to content

Repository files navigation

talk2avatar

Real-time conversational AI with a 3D avatar that speaks back — lip-synced and streamed end-to-end.

Live Demo Next.js 16 React 19 Three.js TypeScript Tailwind CSS 4 License


Overview

talk2avatar is a full-stack web app that lets you have a real-time conversation with a 3D VRM avatar. You type or speak, an LLM responds with streaming tokens, each sentence is synthesized into speech with viseme timing data, and the avatar's lips move in sync — all pipelined with sub-second time-to-first-audio.

Pipeline

Voice / Text  -->  LLM (streaming)  -->  TTS + Visemes  -->  3D Avatar Lip Sync
   input            Groq / Cerebras       Kokoro 82M          VRM + Three.js
                    / OpenRouter          (HeadTTS)           (React Three Fiber)

Features

  • Streaming LLM — Token-by-token responses via Groq, Cerebras, or OpenRouter (switchable at runtime)
  • Sentence-level TTS pipelining — Sentences are split from the token stream and synthesized in parallel so speech starts before the full response arrives
  • Real-time lip sync — Oculus-standard viseme weights applied per-frame to VRM blend shapes via expressionManager
  • Dual TTS engines — Server-side ONNX Runtime for speed, client-side WASM fallback for resilience, browser speechSynthesis as last resort
  • 3D avatar viewer — VRM model loading with idle blink animation, cinematic camera reveal, and orbit controls
  • Speech input — Browser-native speech-to-text via Web Speech API
  • Responsive glass UI — Glass-morphism chat panel, provider tabs, TTS loading progress, animated status indicator

Tech Stack

Layer Technology
Framework Next.js 16 (App Router, Node.js runtime)
UI React 19, Tailwind CSS 4
3D Rendering Three.js r182, React Three Fiber, @pixiv/three-vrm
LLM Vercel AI SDK with OpenAI-compatible providers (Groq, Cerebras, OpenRouter)
TTS HeadTTS — Kokoro 82M ONNX model, server + WASM dual-engine
State Zustand
Speech Input Web Speech API
Deployment Vercel

Getting Started

Prerequisites

  • Node.js 20+ (LTS recommended for stable native ONNX Runtime bindings)
  • API key for at least one LLM provider: Groq, Cerebras, or OpenRouter
  • Modern desktop browser with microphone support

Install

git clone https://github.com/pradhankukiran/talk2avatar.git
cd talk2avatar
npm install

Configure

cp .env.example .env.local

Add your API keys in .env.local:

GROQ_API_KEY=gsk_...
CEREBRAS_API_KEY=csk-...
OPENROUTER_API_KEY=sk-or-...

Run

npm run dev

Open http://localhost:3000. Click the microphone or type a message to start chatting.

First run may take longer while the TTS worker initializes and downloads voice model assets.

Project Structure

src/
  app/
    api/
      chat/route.ts          # Streaming LLM endpoint (Groq / Cerebras / OpenRouter)
      tts/route.ts            # Server-side TTS synthesis endpoint
    page.tsx                  # Main page layout
    layout.tsx                # Root layout with fonts
  components/
    AvatarViewer.tsx          # Three.js canvas, camera reveal, orbit controls
    VrmModel.tsx              # VRM loading, idle blink, per-frame lip sync
    ChatPanel.tsx             # Message list, provider tabs, status
    ChatInput.tsx             # Text input + microphone button
  hooks/
    useChat.ts                # Full pipeline: LLM stream -> sentence split -> TTS -> audio queue
    useTTS.ts                 # Server / client WASM / browser TTS fallback chain
    useLipSync.ts             # Per-frame viseme weight interpolation
    useSpeechRecognition.ts   # Web Speech API wrapper
  lib/
    audio-queue.ts            # FIFO audio playback with binary-search viseme tracking
    sentence-splitter.ts      # Streaming token accumulator with sentence boundary detection
    viseme-mapping.ts         # Oculus viseme -> VRM/RPM blend shape weight tables
    headtts-client.ts         # Browser-side HeadTTS singleton (WASM engine)
    audio-context.ts          # Shared AudioContext singleton
    server/
      headtts-worker.ts       # Node.js worker thread managing HeadTTS ONNX inference
  stores/
    app-store.ts              # Zustand global state
  types/
    index.ts                  # Shared type definitions
    headtts.d.ts              # HeadTTS type declarations

Environment Variables

Variable Required Description
GROQ_API_KEY * Groq API key
CEREBRAS_API_KEY * Cerebras API key
OPENROUTER_API_KEY * OpenRouter API key
HEADTTS_DEBUG No 1 to enable server-side TTS debug logs
NEXT_PUBLIC_TTS_DEBUG No 1 to enable client-side TTS debug logs

* At least one LLM provider key is required.

Debugging TTS

  • Server-side TTS runs on CPU via /api/tts using ONNX Runtime
  • Set HEADTTS_DEBUG=1 for detailed server worker and device timing logs
  • Set NEXT_PUBLIC_TTS_DEBUG=1 for client pipeline and TTS timing in browser console
  • If you see Module did not self-register or Importing modules failed, inspect the preceding HeadTTS Worker error line for the root cause (module resolution vs native addon)

Deploy

Deploy with Vercel

The TTS route runs on the Node.js serverless runtime. ONNX Runtime and HeadTTS assets are bundled via outputFileTracingIncludes in next.config.ts.

License

MIT

About

Real-time conversational AI with a 3D avatar that speaks back — lip-synced and streamed end-to-end.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages