# Video Role-Play Pipeline — Roadmap

Reference architecture target (paraphrased from the management diagram):

```
Mic + Avatar UI
   │
   ▼
Socket.IO  →  FastAPI  →  Whisper STT  ─┐
                       →  GPT-5         ─┤
                       →  Emotion Engine ┤
                       │                 │
                       ▼                 │
                    XTTS-v2 TTS  ◄───────┘
                       │
                       ▼
              Avatar Animation Core
              ├── MuseTalk (realtime lips)
              ├── LivePortrait (head/eyes)
              └── SadTalker (expressions)
                       │
                       ▼
                 Listening Engine
                       │
                       ▼
                  Wav2Lip Refinement
                       │
                       ▼
              FFmpeg GPU Renderer (CUDA / TensorRT)
                       │
                       ▼
                 WebRTC Stream
                       │
                       ▼
                Client Avatar UI
```

Today we have the **right column** of that diagram (Whisper STT + GPT-4
chat + OpenAI TTS + Wav2Lip refinement + FFmpeg CPU + MP4-over-WebSocket
delivery). The improvements below close the gap in priority order. Each
item lists the **estimated effort** and the **infrastructure
prerequisites** so the team can pace the migration.

---

## Phase 0 — Already shipped (incremental fixes)

| Change | What it does | Where |
|---|---|---|
| 16 kHz audio resample before Wav2Lip | Eliminates the +25 ms lip-ahead drift caused by feeding the model 24 kHz audio | `services/video_engine.py :: _resample_wav_to_16k` |
| GPU auto-detect + larger batches | 10× faster inference on CUDA hosts (3-8 s vs 30-90 s/turn) | `services/video_engine.py :: _should_use_wav2lip_gpu` |
| Wav2Lip warm-up at app boot | Saves 5-12 s cold-import on the first candidate turn | `services/video_engine.py :: warmup_wav2lip_async` |
| Tightened `--pads` (0/15/0/0) | Better chin coverage across the 32-avatar pool | `_run_wav2lip` |
| 16 kHz audio file caching | First-time avatars don't re-resample on every turn | as above |

These ship with the existing infra. **No GPU is required** — they
auto-detect and degrade gracefully on CPU.

---

## Phase 1 — Latency-only wins (no new ML deps)

Each item is 1-3 days of work and doesn't require new model artifacts.

### 1.1 Replace WebSocket polling with Server-Sent Events for `video_ready`
- **What:** Push the MP4 URL the moment ffmpeg finishes muxing, instead
  of waiting for the next WS frame.
- **Saves:** ~150-400 ms per turn.
- **Risk:** Low — additive endpoint, frontend falls back to WS.

### 1.2 Streaming TTS → streaming render
- **What:** Start Wav2Lip the moment we have the first 1-second chunk
  of TTS audio, render in segments, concatenate at the end.
- **Saves:** First frame visible while the AI is still synthesising
  the second half of the sentence — cuts perceived latency by ~50 %.
- **Risk:** Medium — requires changing the audio buffering contract.

### 1.3 Pre-render a per-avatar "listening loop"
- **What:** At avatar publish time, generate a 6-second silent
  Wav2Lip render so the loop has computed mouth-close frames rather
  than the raw source.
- **Saves:** No flash when the talk video swaps back to idle.
- **Risk:** Low — one-time batch job.

---

## Phase 2 — Quality wins (small new deps)

### 2.1 Swap OpenAI TTS for **XTTS-v2** (open-source)
- **Why:** XTTS-v2 supports emotional styles per-utterance (the
  "Emotion Engine" box in the target diagram). Wav2Lip's sync also
  improves because XTTS-v2 produces 24 kHz audio with cleaner
  phoneme boundaries than OpenAI's compressed PCM.
- **Cost:** GPU strongly recommended (1.5 s/utterance on RTX 3090,
  8 s on CPU). Model is 2 GB on disk.
- **Effort:** ~1 week. Hot-swappable behind `services.openai_service.
  stream_tts_chunks(voice=…)` — no other code changes.

### 2.2 Add an emotion-detection layer on the candidate's last turn
- **What:** Run a sentiment / arousal classifier on the candidate's
  STT transcript, pass the dominant emotion as a style tag to XTTS-v2.
- **Effort:** ~2-3 days if XTTS-v2 is already in.

---

## Phase 3 — Architecture migration (multi-week)

These items require GPU infrastructure provisioning and model-zoo
operations that live OUTSIDE the codebase.

### 3.1 MuseTalk for realtime lip sync
- **Replaces:** Wav2Lip as the primary lip-sync engine.
- **Why:** MuseTalk runs in realtime (30 fps) on a single mid-range
  GPU and supports continuous audio streams, eliminating the
  "generate full MP4 then play it" round-trip entirely.
- **Cost:** A single RTX 4080 / A10G or better per concurrent
  session. ~3 GB model file. Python deps overlap heavily with
  Wav2Lip so the venv impact is minor.
- **Effort:** 2-3 weeks including the streaming-render bridge.

### 3.2 LivePortrait for head/eye motion
- **What:** Drives natural head turns + eye-line shifts from
  candidate's microphone level. Powers the "look attentive" feature
  during listening.
- **Stacks on:** MuseTalk output (LivePortrait operates on already-
  lip-synced frames).
- **Cost:** Another ~2 GB model; the same GPU can serve both.
- **Effort:** ~1 week once MuseTalk is in.

### 3.3 SadTalker for expression layer
- **What:** Adds micro-expressions (brow lift on surprise, slight
  smile on agreement) keyed off the GPT emotion token.
- **Cost:** Marginal — same GPU, same pipeline.
- **Effort:** ~3-5 days.

### 3.4 WebRTC delivery
- **Replaces:** MP4-over-WebSocket round-trip.
- **Why:** Sub-200 ms glass-to-glass latency, true full-duplex.
- **Requires:** A TURN server (Coturn or Cloudflare), DTLS certs,
  and a Daily.co-style signalling layer (or roll our own).
- **Effort:** 2-4 weeks if rolling own, 1 week if using a managed
  WebRTC provider.

### 3.5 FFmpeg with CUDA / TensorRT
- **What:** Hardware-accelerated H.264 encode + scale on the GPU.
- **Saves:** 50-80 % of the ffmpeg muxing cost per turn.
- **Cost:** ffmpeg recompiled with `--enable-cuda-nvcc --enable-libnpp`
  (or use an nvidia/ffmpeg container image).
- **Effort:** ~1 day for the build, half a day to wire env detection.

### 3.6 Socket.IO over plain WebSocket
- **Why the diagram lists it:** automatic reconnect, room/namespace
  semantics, and better mobile handling than raw WS.
- **Cost:** Frontend pulls in socket.io-client (~30 KB gz); backend
  swaps `fastapi.WebSocket` for `socketio.AsyncServer`.
- **Effort:** ~3-5 days. Mostly mechanical.

---

## Decision tree (what to ship next)

```
Is your team OK adding a CUDA-capable host?
├── No  →  Phase 0 (already shipped) + Phase 1.1 + Phase 1.3 are
│          the most you can do without GPU. Lip-sync quality is
│          already at Phase 0 ceiling.
└── Yes →  Skip Phase 1.2 (XTTS-v2 obsoletes most of the win),
           go to Phase 2.1 (XTTS-v2) first, then Phase 3.1
           (MuseTalk). Those two unlock the rest of the diagram.
```

---

## Where each item lives in code today

| Component | File | Status |
|---|---|---|
| WebSocket transport | `routes/websocket.py` | Phase 3.6 target |
| FastAPI orchestration | `routes/sessions.py`, `routes/websocket.py` | ✓ in place |
| Whisper STT | `services/openai_service.py :: transcribe_audio` | ✓ working |
| GPT-4 chat (→ GPT-5) | `services/openai_service.py :: stream_ai_response` | model name swap |
| Emotion engine | — | Phase 2.2 |
| TTS (→ XTTS-v2) | `services/openai_service.py :: stream_tts_chunks` | Phase 2.1 |
| Wav2Lip render | `services/video_engine.py :: get_or_generate` | ✓ Phase 0 improvements |
| Avatar listening | `static/html/people_hub_role_play.html` (CSS overlays) | ✓ in place |
| FFmpeg muxer | `services/video_engine.py :: _run_ffmpeg_dub` | Phase 3.5 |
| WebRTC delivery | — | Phase 3.4 |
| Frontend avatar player | `static/html/people_hub_role_play.html` (#video-stage) | ✓ in place |

This document is the single source of truth for the migration. Update
it whenever a phase ships so the team always knows where they are.
