01 · Overview
Lifelike talking heads, without the data centre
Large video models make the most convincing digital humans, but they need clusters of server GPUs to run live. Small motion models run anywhere, but they look like puppets. Maya-Lite is built to have both: a 1.3B diffusion model that streams a talking head at up to 96 FPS on one consumer graphics card.
Two ideas carry it. Streaming-aware pre-training keeps lip-sync precise even though live audio arrives in short 1.32-second pieces. Oracle-guided bidirectional distillation cuts generation to a few steps while teaching the model to correct its own drift, so a stream can run indefinitely without the face slowly changing.
Maya-Lite comes in two variants built on the same backbone. Fast is tuned for speed and Pro for detail, so you can pick the trade-off your use case needs.
02 · Positioning
Why 1.3B
Existing talking-head models force a choice. Lightweight models such as SadTalker and Ditto drive an abstract motion rig, which is fast but loses the full texture of the face. Diffusion models such as Hallo3 and AniPortrait model every pixel, but run far below real time or can't stream at all.
We wanted every box ticked: streaming, real time, infinite length and a full pixel-level face, at a size that fits on one GPU.
| Model | Streaming | Real-time | Infinite length | Full-face pixel model | Size |
|---|---|---|---|---|---|
| SadTalker | ✓ | ✓ | ✓ | – | 0.2B |
| AniPortrait | – | – | ✓ | ✓ | 1.7B |
| EchoMimic-V3 | ✓ | – | – | ✓ | 1.3B |
| Ditto | ✓ | ✓ | ✓ | – | 0.2B |
| Hallo3 | ✓ | – | – | ✓ | 5B |
| Sonic | – | – | ✓ | ✓ | 1.5B |
| Maya-Lite | ✓ | ✓ | ✓ | ✓ | 1.3B |
Full-face pixel model: the face is generated in a pixel-level video latent rather than driven through abstract motion coefficients.
03 · Variants
Fast and Pro
Both variants share the same 1.3B diffusion transformer. They differ in the video autoencoder that turns pixels into the tokens the model works on, and that one choice sets the balance between speed and detail.
Speed first
Maya-Lite Fast
96 FPS
- 1× RTX 4090
- LTX-VAE · 32×32×8 compression
An aggressively compressed video latent (one token per 8,192 pixels, about 32× fewer than Pro) makes every step cheap. Built for kiosks, high-concurrency support and anywhere latency matters most.
Detail first
Maya-Lite Pro
10.81 FPS
- Real time on 2× RTX 5090
- Wan 2.1 VAE · 4×8×8 compression
A higher-fidelity latent keeps skin texture, teeth and fine detail. Pro posts the best lip-sync and video-coherence scores of any model tested. Built for brand presenters and premium experiences.
Same voice, same face, both variants
04 · Streaming
Built for streaming audio
In a live conversation the model never hears a full sentence in advance. Audio arrives in pieces of about 1.32 seconds, and speech encoders like wav2vec produce unstable, shifting features from clips that short. That shows up as mouths that lag or mumble.
- Earlier speech, or silence padding at the start of a stream
- Latest audio chunk, which drives the next frames
The other streaming problem is the cold start. The first chunk of a stream has no previous video to continue from, only the reference image. During training, one sample in ten starts from a single frame instead of a run of real motion, so the opening second of a stream looks as natural as the rest.
05 · Method
How Maya-Lite is trained
Inputs
Diffusion transformer × N · 1.3B
- Self-attentionbidirectional within chunk
- Reference cross-attention
- Audio cross-attentionmulti-layer speech features
- Feed-forward
Output
3D VAE decoderVideo chunkLTX-VAE for Fast, Wan 2.1 VAE for Pro
Stage 1 · 100,000 steps
Streaming-aware spatiotemporal pre-training
The model learns to talk from the 8-second audio cache, with the reference image concatenated directly onto its input for a pixel-aligned identity anchor, and with motion context that is sometimes a full run of frames and sometimes a single cold start frame.
Stage 2 · distillation
Oracle-guided bidirectional distillation
Distribution-matching distillation compresses sampling to a few steps and removes classifier-free guidance. Unlike standard distillation, the teacher is shown the real motion while the student works from its own output.
Learning from an oracle
Student · trainable
Teacher · frozen “oracle”
06 · Data
Training data
A 1.3B model has less room to absorb noise than a 14B one, so data quality matters more. We filtered more than 10,000 hours of raw talking-head footage down to 782 hours of clean, tightly aligned video.
- 330Kclips
- 782 hof aligned video
- 60Kspeakers
- 15languages
- 512²face-centred crops
The filtering pipeline
- DeduplicateSource IDs and MD5 hashes remove repeats across sources.
- SliceScene detection cuts footage into coherent 5–50 s clips.
- StandardizeEvery clip is re-timed to 25 FPS.
- CropFace detection centres a 512×512 crop on the speaker.
- Remove jump cutsOptical flow flags abrupt scene transitions.
- Drop faceless footageClips with too many frames missing a face are cut.
- Filter occlusionPose keypoints catch hands covering the face.
- Check lip-syncSyncNet scores discard poorly aligned audio.
- AnnotateLanguage, age, gender and ethnicity labels for every clip.
How it compares
| Dataset | Speakers | Clips | Hours | Resolution | Languages |
|---|---|---|---|---|---|
| MEAD | 60 | 281.4K | 39 | 384p | English |
| HDTF | 362 | 10K | 15.8 | 512p | – |
| Hallo3 | – | 101.5K | 70 | 720p | – |
| TalkVid | 7,729 | 281.4K | 1,244 | 1080p+ | 15 |
| Maya-Lite training set | 60K | 330K | 782 | 512p | 15 |
07 · Results
Benchmarks
We evaluate on HDTF and VFHQ against six talking-head models, in both streaming and offline modes. In streaming mode Maya-Lite Pro leads on lip-sync and video coherence on both benchmarks, and Maya-Lite Fast keeps most of that quality at 96 FPS.
| Model | FID ↓ | FVD ↓ | Sync-C ↑ | Sync-D ↓ | FPS ↑ |
|---|---|---|---|---|---|
| SadTalker* | 21.58 | 207.67 | 4.60 | 9.21 | 2.17 |
| AniPortrait△ | 19.83 | 242.29 | 1.89 | 11.91 | – |
| EchoMimic* | 9.00 | 155.71 | 3.56 | 10.22 | 0.81 |
| Ditto* | 12.35 | 199.13 | 3.57 | 10.49 | 45.04 |
| Hallo3* | 15.95 | 160.94 | 3.18 | 10.72 | 0.16 |
| Sonic△ | 13.53 | 113.31 | 5.17 | 8.69 | – |
| Maya-Lite Fast* | 11.37 | 126.52 | 4.21 | 9.49 | 96.00 |
| Maya-Lite Pro* | 9.97 | 111.38 | 5.73 | 8.77 | 10.81 |
| Maya-Lite Fast△ | 10.78 | 115.94 | 5.12 | 8.80 | – |
| Maya-Lite Pro△ | 8.31 | 103.14 | 6.04 | 8.46 | – |
* streaming · △ non-streaming (offline), where FPS doesn't apply · best result per column in bold.
- FID ↓
- Frame realism · Fréchet Inception Distance
- FVD ↓
- Video realism and coherence · Fréchet Video Distance
- Sync-C ↑
- Lip-sync confidence · SyncNet
- Sync-D ↓
- Lip-sync distance · SyncNet
- FPS ↑
- Frames per second · Single GPU
- Maya-Lite Fast96
- Ditto45.04
- Maya-Lite Pro10.81
- SadTalker2.17
- EchoMimic0.81
- Hallo30.16
25 FPS real-time threshold
Streaming costs very little quality. On HDTF, Maya-Lite Pro's lip-sync confidence moves only from 6.04 offline to 5.73 when streaming, still ahead of every other model in either mode. The audio cache and cold-start training are doing their job.
08 · Roadmap
What's next
Maya-Lite is optimized for the face and head. At 1.3B parameters it has less capacity for complex physical motion than our larger models, so large body movements and detailed hand gestures are less precise than the face itself. For full-body presenters, use Maya-2. Next, we're working on:
- Scale the architecture so Maya-Lite models the whole upper body, not just the face and head, without giving up real-time speed.
- Better large-amplitude movement and hand gestures, where a 1.3B model has less capacity than our larger models.
- Push the Pro variant to real time on a single consumer GPU.
Keep exploring
- ResearchMaya-2 (14B) Our full-body foundation model: 0.87 s start-up and 32 FPS at 14B parameters.
- ProductThrifty Studios Design, voice and deploy an artificial human in minutes.
- Business caseAI Humans ROI calculator Model the payback of putting AI humans to work with your own numbers.
- ResearchAll research Our approach, our team and the rest of our published work.
