# Enabling a real talking avatar (moving lips)

The base pipeline produces a voice + a "living" portrait (breathing zoom + head
sway). To make the **lips actually move in sync with the speech** — like HeyGen
/ Tavus — you need the neural lip-sync model **Wav2Lip**. The app already calls
it automatically once it's installed; you just have to put the model on disk.

## How the pipeline uses the models

For an uploaded photo, generation runs in this order (each step is skipped
gracefully if its model isn't installed):

1. **TTS** — script → voice (OpenAI TTS, then gTTS fallback)
2. **LivePortrait / SadTalker** *(optional)* — adds head/expression motion
3. **Wav2Lip** *(this guide)* — drives the mouth so the lips move with the voice
4. **CodeFormer** *(optional)* — sharpens the mouth/face after lip-sync
5. Watermark + thumbnail + save

Installing just **Wav2Lip** is enough to get moving lips. CodeFormer is
recommended too (it removes the slight blur Wav2Lip leaves around the mouth).

## 1. Install Wav2Lip (one command)

From the project root:

```bash
python models/install_wav2lip.py
```

This clones Wav2Lip into `models/wav2lip/`, downloads the two checkpoints, and
installs its Python deps (torch, opencv, librosa, …).

If a checkpoint download 404s, place the files manually:

- `models/wav2lip/checkpoints/wav2lip_gan.pth`
- `models/wav2lip/face_detection/detection/sfd/s3fd.pth`

(The installer prints working mirror links for each.)

## 2. Point the app at it

In `backend/.env` (these are already the defaults):

```
WAV2LIP_PATH=models/wav2lip
MODELS_PYTHON=python      # interpreter that has Wav2Lip's deps installed
```

If you installed Wav2Lip's deps into a separate virtualenv, set `MODELS_PYTHON`
to that venv's python, e.g. `MODELS_PYTHON=/path/to/wav2lip-venv/bin/python`.

## 3. Restart and regenerate

```bash
python run.py
```

Generate a new video in the Studio. In the backend logs you'll see
`Wav2Lip lip-sync applied`, and the avatar's lips will now move with the voice.
(Existing videos won't change — generate a fresh one.)

## Performance notes

- Wav2Lip runs on **CPU** (≈1–4 min for a short clip) or much faster on a CUDA
  GPU / Apple-Silicon MPS if PyTorch detects one.
- Keep resolution at **720p** for the fastest CPU renders.
- A clear, front-facing photo gives the best mouth tracking.

## Optional: CodeFormer (sharper mouth)

```bash
git clone https://github.com/sczhou/CodeFormer.git models/codeformer
cd models/codeformer && pip install -r requirements.txt && python basicsr/setup.py develop
# download weights as per the CodeFormer README
```

Set `CODEFORMER_PATH=models/codeformer` in `.env`. The pipeline runs it after
Wav2Lip when "Enable CodeFormer face enhancement" is checked in the Studio.

## Optional: LivePortrait / SadTalker (head + expression motion)

Install into `models/liveportrait/` with an entrypoint (`inference.py`) and a
driving template at `models/liveportrait/assets/driving.mp4`, then set
`LIVEPORTRAIT_PATH`. The pipeline animates the still first, then Wav2Lip syncs
the lips on top.

## Troubleshooting

- **"Wav2Lip present but missing entrypoint/checkpoint — skipping"** in logs:
  the `inference.py` or `wav2lip_gan.pth` isn't where it expects. Check the paths
  in step 1.
- **No face detected / Wav2Lip error**: use a clearer front-facing photo; the
  pipeline falls back to the living-portrait video so you still get output.
- **torch install issues on macOS**: `pip install torch torchvision` (CPU/MPS
  build installs automatically on recent pip).
