01 · Overview
A 14B model that keeps up with a conversation
Large diffusion models make the most convincing digital humans, but they have always been too slow to talk to. Maya-2 closes that gap. It keeps the full 14 billion parameters and still starts streaming in 0.87 seconds at 32 frames per second.
Three ideas make that possible. Bidirectional streaming distillation keeps full attention inside every video chunk, so motion and detail stay coherent. Multi-step retrospective self-correction teaches the model to recover from its own small errors, so a stream can run indefinitely without drifting. And a full-stack acceleration suite (hybrid sequence parallelism, a parallel VAE and Hopper-native kernels) turns that quality into sub-second latency.
Because the student model mirrors its teacher so closely, training is light too: 1,000 steps of fine-tuning and 200 steps of distillation, roughly 23× less than the previous real-time approach.
02 · Architecture
Why bidirectional
Most streaming video models are autoregressive: each frame may only look backwards. That one-way dependency is the root of the identity drift, flicker and slow collapse you see when a digital human talks for more than a few minutes.
The hard problem in live generation isn't remembering more of the past. It's stopping small errors from compounding.
Maya-2 generates in chunks, and inside a chunk every frame can already see every other frame. So we keep the original model's bidirectional attention instead of forcing it into a causal shape. Each chunk plans its motion with both past and near-future context, which makes full-body gestures smoother and gives the stream a sturdier unit to build on. It also means the student matches its teacher's architecture exactly, which is why distillation converges so quickly.
03 · Stability
Infinite-length streaming
A support call, a lesson or a live stream doesn't end after ten seconds, so Maya-2 is evaluated on streams longer than five minutes, and tested out to 1,000 seconds of continuous generation.
Identity, lip-sync and background hold steady across the whole run. On the long benchmark Maya-2 posts the best lip-sync confidence (Sync-C 1.61) of any model tested, while also keeping the highest subject and background consistency among full-body generators.
- No cuts or resets between chunks
- Identity preserved from first frame to last
- Lip-sync that doesn't degrade over time
04 · Range
Any face, any style
One reference image is enough. Maya-2 animates photoreal people, illustrated characters and anime with the same lip-sync and the same natural upper-body motion, so your AI human can look exactly like your brand.
Stylised characters
Photoreal, full gesture
05 · Method
How Maya-2 is trained
Maya-2 is built on a 14B diffusion transformer. Speech features from wav2vec, identity features from CLIP and a few frames of recent motion condition every chunk it generates.
Inputs
Diffusion transformer × N · 14B
- Self-attentionbidirectional within chunk
- Reference cross-attention
- Audio cross-attention
- Feed-forward
Output
3D VAE decoder33-frame chunk28 new frames + 5 motion frames carried forward
Stage 1 · 1,000 steps
Latency-aware spatiotemporal adaptation
We fine-tune the model to work at the lower resolutions and shorter frame sequences that real time demands. Dynamic aspect-ratio bucketing keeps training data intact instead of padding or cropping it, so the model recovers fine detail and identity even at reduced resolution.
Stage 2 · 200 steps
Self-correcting bidirectional distillation
Distribution-matching distillation compresses sampling to 4 steps and removes classifier-free guidance entirely. The student is then made to generate several chunks in a row from its own outputs, so it learns to correct the drift it creates.
Learning from its own mistakes
06 · Systems
Real-time inference
Fast training alone doesn't make a 14B model real time. Maya-2 runs on a purpose-built inference stack for a single 8-GPU NVIDIA H800 node, and every stage of the loop was optimized in turn.
- Audio processing33 ms
- 4-step DiT denoising616 ms
- VAE frame decoding187 ms
- Motion-frame encoding14 ms
- Other overhead26 ms
~5×
Hybrid sequence parallelism
Ulysses and Ring Attention spread the DiT's attention workload across eight GPUs, cutting single-step inference about five-fold.
~5×
Parallel 3D VAE
Once the DiT is fast, decoding becomes the bottleneck. Slicing the spatial decode across GPUs removes it.
−20%
Hopper-native kernels
FlashAttention3 overlaps data movement with compute on NVIDIA Hopper, trimming attention latency versus FlashAttention2.
154 ms
Whole-graph compilation
torch.compile fuses the pipeline end to end, bringing one DiT step on 8 GPUs down from 193 ms.
Latency by GPU count (ms)
| GPUs | VAE encode | VAE decode | DiT step | DiT step (compiled) |
|---|---|---|---|---|
| 1 | 97 | 988 | 1,070 | 800 |
| 2 | 69 | 690 | 620 | 490 |
| 4 | 39 | 350 | 313 | 261 |
| 8 | 21 | 192 | 193 | 154 |
07 · Results
Benchmarks
We compare Maya-2 with six leading audio-driven avatar models on TalkBench, covering short clips and long streams. Maya-2 leads on aesthetics, image quality and lip-sync on short clips, and is the fastest model tested by a wide margin.
| Model | ASE ↑ | IQA ↑ | Sync-C ↑ | Sync-D ↓ | Subject-C ↑ | BG-C ↑ | Motion-S ↑ | Temporal-F ↑ | FPS ↑ |
|---|---|---|---|---|---|---|---|---|---|
| Ditto* | 3.10 | 4.37 | 1.04 | 12.58 | 99.80 | 99.23 | 99.75 | 99.86 | 21.80 |
| EchoMimic-V3 | 3.45 | 4.70 | 0.89 | 12.81 | 98.75 | 95.98 | 99.54 | 99.44 | 0.53 |
| StableAvatar | 3.05 | 3.01 | 0.78 | 12.12 | 98.34 | 96.46 | 99.44 | 99.06 | 0.64 |
| OmniAvatar | 3.06 | 3.01 | 1.32 | 11.85 | 98.64 | 96.95 | 99.60 | 99.66 | 0.16 |
| LiveAvatar* | 3.10 | 3.25 | 1.01 | 12.10 | 98.27 | 97.48 | 99.25 | 98.86 | 20.88 |
| InfiniteTalk△ | 3.09 | 3.04 | 1.25 | 11.89 | 98.71 | 97.04 | 99.54 | 99.34 | 5.10 |
| Maya-2* | 3.51 | 4.79 | 1.47 | 11.56 | 99.22 | 98.10 | 99.61 | 99.52 | 32.00 |
* real-time capable · △ LoRA version · best result per column in bold. Throughput (FPS) is a property of each model and is shown on the short split.
- ASE ↑
- Aesthetics score · Q-Align
- IQA ↑
- Image quality · Q-Align
- Sync-C ↑
- Lip-sync confidence · SyncNet
- Sync-D ↓
- Lip-sync distance · SyncNet
- Subject-C ↑
- Subject consistency · VBench
- BG-C ↑
- Background consistency · VBench
- Motion-S ↑
- Motion smoothness · VBench
- Temporal-F ↑
- Temporal flicker · VBench
- FPS ↑
- Frames per second · Throughput
A note on consistency scores. Ditto scores highest on subject and background consistency because it repaints only the face and keeps the body and background pixel-static. Maya-2 generates full-body motion, which naturally adds pixel variance, and still holds 99.22 subject consistency.
08 · Ablations
What made the difference
Letting the student roll out a random number of chunks during distillation, rather than a fixed number, gave the best quality and the most stable long streams at a moderate training cost.
| Strategy | ASE | IQA | Sync-C | Sync-D | Training cost |
|---|---|---|---|---|---|
| Fixed K = 1 | 3.42 / 3.08 | 4.69 / 2.99 | 1.39 / 1.12 | 11.79 / 12.47 | 2.33 h |
| Fixed K = 3 | 3.47 / 3.09 | 4.75 / 3.00 | 1.39 / 1.59 | 11.69 / 12.02 | 4.40 h |
| Fixed K = 5 | 3.50 / 3.10 | 4.79 / 3.01 | 1.35 / 1.44 | 11.87 / 12.33 | 6.40 h |
| Random K ∈ [1, 5] | 3.51 / 3.12 | 4.79 / 3.04 | 1.47 / 1.61 | 11.56 / 12.25 | 4.40 h |
Values shown as short / long split.
We also found that conditioning the teacher on the student's own predicted motion (with noise added, and left out of the loss) beats conditioning on ground truth, raising the aesthetics score from 3.46 to 3.51. Training on what the model will actually see at inference closes the gap between the two.
09 · Roadmap
What's next
Today Maya-2 needs a single 8× H800 node to hit its real-time numbers. The next step is efficiency rather than scale:
- Run Maya-2 on consumer-grade GPUs instead of an 8× H800 node, through pruning and quantization.
- Optimized attention mechanisms that keep the bidirectional quality at a fraction of the compute.
- Lower start-up latency again, so the first frame lands before a person notices the pause.
Keep exploring
- ProductThrifty Studios Design, voice and deploy an artificial human built on Maya-2.
- Business caseAI Humans ROI calculator Model the payback of putting AI humans to work with your own numbers.
- CustomersCustomer stories See where teams already run AI humans in support, training and healthcare.
- ResearchAll research Our approach, our team and the rest of our published work.
