All projects

live-media-tavus

Realtime voice assistant with a live Tavus virtual avatar in the room: OpenAI Realtime, MediaPipe face detection triggers, LiveKit tokens.

A complete real-time voice AI assistant with virtual avatar integration using LiveKit, OpenAI Realtime API, MediaPipe face detection, and Tavus virtual avatars.

Features

✅ Real-time Voice Communication - OpenAI Realtime API integration
✅ Virtual Avatar - Live interactive Tavus avatar as room participant
✅ Face Detection - MediaPipe-powered face detection with triggers
✅ Secure Token Generation - Local Flask server for LiveKit tokens
✅ Web Frontend - Complete browser-based interface
✅ Auto-reconnection - Robust connection handling

Architecture

┌─────────────────┐    ┌──────────────────┐    ┌─────────────────┐
│   Web Frontend  │    │  Token Server    │    │  Avatar Agent   │
│                 │    │                  │    │                 │
│ • Face Detection│◄──►│ • JWT Generation │    │ • Tavus Avatar  │
│ • Voice Control │    │ • CORS Support   │    │ • OpenAI Voice  │
│ • Avatar Display│    │ • Secure Auth    │    │ • LiveKit Room  │
└─────────────────┘    └──────────────────┘    └─────────────────┘
         │                       │                       │
         └───────────────────────┼───────────────────────┘
                                 │
                    ┌──────────────────┐
                    │   LiveKit Cloud  │
                    │                  │
                    │ • Room Management│
                    │ • WebRTC Streams │
                    │ • Participant Sync│
                    └──────────────────┘

Setup

1. Environment Configuration

Create a .env file with your credentials:

# Tavus API
TAVUS_API_KEY=your_tavus_api_key
TAVUS_REPLICA_ID=your_replica_id
TAVUS_PERSONA_ID=your_persona_id

# LiveKit Cloud
LIVEKIT_URL=wss://your-project.livekit.cloud
LIVEKIT_API_KEY=your_livekit_api_key
LIVEKIT_API_SECRET=your_livekit_api_secret

# OpenAI
OPENAI_API_KEY=your_openai_api_key

2. Install Dependencies

pip install -r req.txt

3. Create Tavus Persona

Run the persona creation script to set up a LiveKit-compatible Tavus persona:

python create_tavus_persona.py

This will:

  • Create a new Tavus persona with pipeline_mode: "echo" and transport_type: "livekit"
  • Add the persona ID to your .env file
  • List your existing personas

Usage

1. Start the Token Server

python token_server.py

The server will run on http://localhost:5000 and provide secure JWT tokens for LiveKit authentication.

2. Start the Avatar Agent

python run_avatar_agent.py

Or directly:

python livekit_tavus_agent.py

This will:

  • Connect to your LiveKit room
  • Create a live Tavus avatar as a room participant
  • Handle voice interactions and face detection messages

3. Open the Frontend

Open livekit-frontend/index.html in your browser.

The interface provides:

  • Face Detection Panel - Enable webcam and face detection
  • Avatar Panel - View the live Tavus avatar video stream
  • Voice Assistant Panel - Connect and control voice interaction
  • Messages Panel - View real-time logs and transcriptions

4. Interact with the System

  1. Enable Face Detection - Click “Enable Face Detection” to start monitoring
  2. Connect to Voice Agent - Click “Connect” to join the LiveKit room
  3. Talk to Avatar - Speak naturally; the avatar will respond with voice and video
  4. Face Detection Triggers - When new faces are detected, a greeting message is sent to the avatar

File Structure

├── livekit_tavus_agent.py      # Main avatar agent with Tavus integration
├── tavus_avatar_agent.py       # Legacy avatar agent (custom implementation)
├── token_server.py             # Flask server for secure token generation
├── create_tavus_persona.py     # Script to create LiveKit-compatible personas
├── run_avatar_agent.py         # Launcher script with environment checks
├── req.txt                     # Python dependencies
├── .env                        # Environment variables (create this)
├── livekit-frontend/
│   ├── index.html              # Main frontend interface
│   ├── script.js               # JavaScript with LiveKit + MediaPipe integration
│   ├── style.css               # Responsive CSS styling
│   └── config.js               # Frontend configuration
└── README.md                   # This file

How It Works

Avatar Integration

The system uses the official LiveKit Tavus plugin to create a real-time interactive avatar:

  1. Persona Creation - A Tavus persona is created with pipeline_mode: "echo" for real-time interaction
  2. Avatar Session - tavus.AvatarSession creates a live avatar participant in the LiveKit room
  3. Video Stream - The avatar appears as a video participant that other users can see
  4. Voice Interaction - The avatar responds to voice input with synchronized lip movement

Face Detection Workflow

  1. MediaPipe Integration - Browser-based face detection using Google’s MediaPipe
  2. Real-time Processing - Optimized for 12 FPS with bounding box visualization
  3. Event Triggers - New face detection sends “new face entered the room” message
  4. Avatar Response - Avatar receives the message and provides a personalized greeting

Security

  • JWT Tokens - Secure LiveKit authentication with 15-minute expiry
  • Local Token Server - No hardcoded credentials in frontend
  • CORS Protection - Proper cross-origin resource sharing configuration
  • Environment Variables - All sensitive data stored in .env file

Troubleshooting

Common Issues

Avatar not appearing:

  • Check that TAVUS_PERSONA_ID is set in .env
  • Verify the persona has pipeline_mode: "echo" and transport_type: "livekit"
  • Ensure the avatar agent is running before connecting frontend

Token generation fails:

  • Verify LiveKit credentials in .env
  • Check that token server is running on port 5000
  • Ensure no firewall blocking localhost connections

Face detection not working:

  • Allow camera permissions in browser
  • Check browser console for MediaPipe errors
  • Verify blaze_face_short_range.tflite model is accessible

Audio issues:

  • Click anywhere on the page to enable audio playback
  • Check microphone permissions
  • Verify audio devices are working

Debug Mode

Enable verbose logging by setting environment variable:

export LIVEKIT_LOG_LEVEL=debug
python livekit_tavus_agent.py

API References

License

This project is for demonstration purposes. Please ensure you comply with the terms of service for all integrated APIs (LiveKit, Tavus, OpenAI, MediaPipe).