A complete real-time voice AI assistant with virtual avatar integration using LiveKit, OpenAI Realtime API, MediaPipe face detection, and Tavus virtual avatars.
Features
✅ Real-time Voice Communication - OpenAI Realtime API integration
✅ Virtual Avatar - Live interactive Tavus avatar as room participant
✅ Face Detection - MediaPipe-powered face detection with triggers
✅ Secure Token Generation - Local Flask server for LiveKit tokens
✅ Web Frontend - Complete browser-based interface
✅ Auto-reconnection - Robust connection handling
Architecture
┌─────────────────┐ ┌──────────────────┐ ┌─────────────────┐
│ Web Frontend │ │ Token Server │ │ Avatar Agent │
│ │ │ │ │ │
│ • Face Detection│◄──►│ • JWT Generation │ │ • Tavus Avatar │
│ • Voice Control │ │ • CORS Support │ │ • OpenAI Voice │
│ • Avatar Display│ │ • Secure Auth │ │ • LiveKit Room │
└─────────────────┘ └──────────────────┘ └─────────────────┘
│ │ │
└───────────────────────┼───────────────────────┘
│
┌──────────────────┐
│ LiveKit Cloud │
│ │
│ • Room Management│
│ • WebRTC Streams │
│ • Participant Sync│
└──────────────────┘
Setup
1. Environment Configuration
Create a .env file with your credentials:
# Tavus API
TAVUS_API_KEY=your_tavus_api_key
TAVUS_REPLICA_ID=your_replica_id
TAVUS_PERSONA_ID=your_persona_id
# LiveKit Cloud
LIVEKIT_URL=wss://your-project.livekit.cloud
LIVEKIT_API_KEY=your_livekit_api_key
LIVEKIT_API_SECRET=your_livekit_api_secret
# OpenAI
OPENAI_API_KEY=your_openai_api_key
2. Install Dependencies
pip install -r req.txt
3. Create Tavus Persona
Run the persona creation script to set up a LiveKit-compatible Tavus persona:
python create_tavus_persona.py
This will:
- Create a new Tavus persona with
pipeline_mode: "echo"andtransport_type: "livekit" - Add the persona ID to your
.envfile - List your existing personas
Usage
1. Start the Token Server
python token_server.py
The server will run on http://localhost:5000 and provide secure JWT tokens for LiveKit authentication.
2. Start the Avatar Agent
python run_avatar_agent.py
Or directly:
python livekit_tavus_agent.py
This will:
- Connect to your LiveKit room
- Create a live Tavus avatar as a room participant
- Handle voice interactions and face detection messages
3. Open the Frontend
Open livekit-frontend/index.html in your browser.
The interface provides:
- Face Detection Panel - Enable webcam and face detection
- Avatar Panel - View the live Tavus avatar video stream
- Voice Assistant Panel - Connect and control voice interaction
- Messages Panel - View real-time logs and transcriptions
4. Interact with the System
- Enable Face Detection - Click “Enable Face Detection” to start monitoring
- Connect to Voice Agent - Click “Connect” to join the LiveKit room
- Talk to Avatar - Speak naturally; the avatar will respond with voice and video
- Face Detection Triggers - When new faces are detected, a greeting message is sent to the avatar
File Structure
├── livekit_tavus_agent.py # Main avatar agent with Tavus integration
├── tavus_avatar_agent.py # Legacy avatar agent (custom implementation)
├── token_server.py # Flask server for secure token generation
├── create_tavus_persona.py # Script to create LiveKit-compatible personas
├── run_avatar_agent.py # Launcher script with environment checks
├── req.txt # Python dependencies
├── .env # Environment variables (create this)
├── livekit-frontend/
│ ├── index.html # Main frontend interface
│ ├── script.js # JavaScript with LiveKit + MediaPipe integration
│ ├── style.css # Responsive CSS styling
│ └── config.js # Frontend configuration
└── README.md # This file
How It Works
Avatar Integration
The system uses the official LiveKit Tavus plugin to create a real-time interactive avatar:
- Persona Creation - A Tavus persona is created with
pipeline_mode: "echo"for real-time interaction - Avatar Session -
tavus.AvatarSessioncreates a live avatar participant in the LiveKit room - Video Stream - The avatar appears as a video participant that other users can see
- Voice Interaction - The avatar responds to voice input with synchronized lip movement
Face Detection Workflow
- MediaPipe Integration - Browser-based face detection using Google’s MediaPipe
- Real-time Processing - Optimized for 12 FPS with bounding box visualization
- Event Triggers - New face detection sends “new face entered the room” message
- Avatar Response - Avatar receives the message and provides a personalized greeting
Security
- JWT Tokens - Secure LiveKit authentication with 15-minute expiry
- Local Token Server - No hardcoded credentials in frontend
- CORS Protection - Proper cross-origin resource sharing configuration
- Environment Variables - All sensitive data stored in
.envfile
Troubleshooting
Common Issues
Avatar not appearing:
- Check that
TAVUS_PERSONA_IDis set in.env - Verify the persona has
pipeline_mode: "echo"andtransport_type: "livekit" - Ensure the avatar agent is running before connecting frontend
Token generation fails:
- Verify LiveKit credentials in
.env - Check that token server is running on port 5000
- Ensure no firewall blocking localhost connections
Face detection not working:
- Allow camera permissions in browser
- Check browser console for MediaPipe errors
- Verify
blaze_face_short_range.tflitemodel is accessible
Audio issues:
- Click anywhere on the page to enable audio playback
- Check microphone permissions
- Verify audio devices are working
Debug Mode
Enable verbose logging by setting environment variable:
export LIVEKIT_LOG_LEVEL=debug
python livekit_tavus_agent.py
API References
License
This project is for demonstration purposes. Please ensure you comply with the terms of service for all integrated APIs (LiveKit, Tavus, OpenAI, MediaPipe).