Download README_lipsync.md from akamotaco/ppaso-tts-v1: direct link, hf CLI and curl.
- Browser
- Download file 10.9 kB
-
https://huggingface.co/akamotaco/ppaso-tts-v1/resolve/main/README_lipsync.md
- Command line
-
hf download hf://akamotaco/ppaso-tts-v1/README_lipsync.md
-
curl -L -o README_lipsync.md https://huggingface.co/akamotaco/ppaso-tts-v1/resolve/main/README_lipsync.md
Ppaso-TTS Lipsync β companion animation module
νκ΅μ΄: README_lipsync_ko.md Β· Main: README.md
Blendshape mapping detail: README_bs2d.md
Pure-numpy companion module ppaso_lipsync (in example/) turns
Ppaso-TTS output into per-frame 2D mouth + 1D eye-blink animation curves.
A single lipsync curve drives two render routes:
- atlas β 2D sprite swap (nearest viseme β one mouth sprite).
07_avatar_video.py - blendshape β 3D morph-target blend via
bs2d.08_vrm_video.py
Demo
Blendshape β 3D VRM, lipsync + body idle (samples/08_vrm_lipsync.mp4,
download):
Atlas β 2D sprite (samples/06_lipsync_intro.mp4,
download):
μ£μ§λλ°μ΄μ€μ© κ²½λν 립μ±ν¬ μ λλ©μ΄μ κΈ°λ₯. λμΌν 컀λΈκ° 2D μ€νλΌμ΄νΈμ 3D λΈλ λμμ μ λλ©μ΄μ μ λͺ¨λ ꡬλν©λλ€. (One curve drives both 2D sprite and 3D blendshape animation.)
How it works
TTS phones + durations
β
βΌ
viseme decompose β 9 anchor (A/I/U/E/O/EU/EO/M/sil) + j/w glide,
β palatalized / alveolar / sibilant / velar / glottal
βΌ
keyframes β 2D xy curve β shared lipsync curve, xβ[-1,1] yβ[-0.1,1]
β
ββββββββββββΊ atlas route : nearest viseme β sprite swap
β
ββββββββββββΊ blendshape route : bs2d β barycentric β morph blend
Eye blink (independent) β state_timeline {0:open,1:half,2:close}
+ curve_timeline [0,1]
The lipsync curve generation (decompose β keyframes β interpolate) is
shared. The two routes differ only in how they consume the (x, y) curve β
see the atlas/blendshape distinction below.
Route A β atlas (2D sprite)
Each frame the curve point picks the nearest viseme and the renderer shows that one mouth sprite (sprites cannot blend). Best for 2D character art with a sprite sheet.
cd example
python 07_avatar_video.py --text "μλ
νμΈμ." --out hi.mp4
- Sample 2D avatar:
example/avatar/(atlas.json + sprite sheet). - Curve β sprite mapping:
atlas.jsonviseme_to_sprite.
Route B β blendshape (3D VRM)
For 3D characters with morph-target (blendshape) visemes. Here the curve
goes through bs2d, which finds the triangle containing the point and
returns barycentric blend weights β so 3 blendshapes are mixed per frame,
giving continuous mouth motion (a sprite swap cannot do this).
lipsync (x,y) curve β bs2d(target="vrm") β {a,i,u,e,o,neutral} weight
β VRM blendShapeMaster morph β mesh deform β render
- vrm-spec β bs2d maps the curve onto a standard VRM 0.x avatar's 5 viseme + neutral expressions (Korean γ γ £γ γ γ β VRM a/i/u/e/o).
- Three render backends (
--backend):- gles (default) β EGL + OpenGL ES 3.0. Real GPU acceleration on
ARM SoCs (RK3576 Mali-G52 etc.). Needs
PyOpenGL. - egl β moderngl + EGL desktop GL 3.3. For x86 GPUs / Mesa llvmpipe;
ARM GPU drivers usually cap desktop GL at 3.1 so moderngl can't run
there β use
gleson-device. Needsmoderngl+glcontext. - numpy β pure-numpy software rasterizer, zero GPU/driver dependency (slow β fallback only).
- gles (default) β EGL + OpenGL ES 3.0. Real GPU acceleration on
ARM SoCs (RK3576 Mali-G52 etc.). Needs
- Bundled sample VRM:
example/avatar_vrm/AvatarSample_F.vrm(VRoid 'Vita', CC0).
cd example
python 08_vrm_video.py --text "μλ
νμΈμ." --out hi.mp4 # gles (default)
python 08_vrm_video.py --backend egl # x86 desktop GPU
python 08_vrm_video.py --backend numpy # no-GPU fallback
bs2d full-spec β vrm-spec, the projection / barycentric stages, and how to map your own character's blendshapes (9 or more) β see README_bs2d.md.
Body animation (optional)
09_vrm_live.py and 10_vrm_live_streaming.py add a skeletal body idle
on top of lipsync β GPU skinning via raster_gles_skinned. It needs a body
clip (--idle clip.npz); without one the body stays at rest pose
(lipsync and video still render fine).
Make a clip from any FBX skeletal animation with tools/fbx_to_clip.py
(needs bpy β pip install bpy):
python tools/fbx_to_clip.py idle.fbx avatar_vrm/idle.npz
python 09_vrm_live.py --idle avatar_vrm/idle.npz
The clip is not bundled β animation assets such as Mixamo are
royalty-free to embed in your rendered output, but their raw motion data is
generally not redistributable, so supply your own. BodyClipPlayer assumes
a loop-authored clip; fbx_to_clip.py drops the duplicate loop endpoint.
Quick start
from ppaso_tts import PpasoTTS
from ppaso_lipsync import generate_blink_timeline
tts = PpasoTTS('./', backend='onnx') # or 'rknn'
wav, lipsync = tts.synthesize_with_lipsync('μλ
νμΈμ.', fps=30)
# lipsync.xy_timeline β (n_frames, 2) β feeds atlas or blendshape route
# lipsync.keyframes β viseme keyframes (debug)
blink = generate_blink_timeline(duration=lipsync.duration, fps=30, seed=42)
# blink.state_timeline / blink.curve_timeline
Streaming (chunk-to-chunk continuity is automatic):
from ppaso_tts import PpasoTTS, StreamingTTS
stream = StreamingTTS(PpasoTTS('./', backend='rknn'))
for token in llm_token_stream():
for wav, lipsync in stream.feed_with_lipsync(token): # emits on .!?
play_audio(wav); animate(lipsync)
for wav, lipsync in stream.flush_with_lipsync():
...
Performance
Atlas route β RK3576 measured
End-to-end (TTS + lipsync + eye blink + sprite composite at 768Γ1152, no ffmpeg encode). RK3576, best-of-5 @ 30 fps.
| Case | Input | wav | Backend | E2E | RTF | Speed |
|---|---|---|---|---|---|---|
| short | 21 chars | 3.49 s | RK3576 NPU | 864 ms | 0.247 | 4.0Γ RT |
| ONNX CPU (4t) | 897 ms | 0.256 | 3.9Γ RT | |||
| medium | 41 chars | 7.26 s | RK3576 NPU | 1867 ms | 0.257 | 3.9Γ RT |
| long | 110 chars | 16.65 s | RK3576 NPU | 3668 ms | 0.220 | 4.5Γ RT |
Reproduce:
cd example && python bench_lipsync.py --backend rknn --model-dir ... Lipsync curve generation itself adds < 1 ms on top of TTS.
Blendshape route β RK3576 measured
End-to-end (TTS + lipsync + eye blink + morph + render at 512Γ768, no
ffmpeg encode β same E2E definition as the atlas route above).
RK3576 Mali-G52, gles backend, best-of-3 @ 30 fps. Two modes:
lipsync only (static skeleton β 08_vrm_video.py) and lipsync +
body (GPU skinning + body idle animation β 09_vrm_live.py).
β lipsync only β E2E by text length:
| Case | wav | frames | E2E | RTF |
|---|---|---|---|---|
| short | 3.51 s | 110 | 2222 ms | 0.63 |
| medium | 7.26 s | 223 | 4637 ms | 0.64 |
| long | 16.31 s | 497 | 9358 ms | 0.57 |
β‘ lipsync + body β E2E by text length:
| Case | wav | frames | E2E | RTF |
|---|---|---|---|---|
| short | 3.51 s | 110 | 2809 ms | 0.80 |
| medium | 7.26 s | 223 | 5580 ms | 0.77 |
| long | 16.31 s | 497 | 12314 ms | 0.76 |
β’ lipsync + body β medium (7.26 s, 223 frames) stage breakdown:
| Stage | medium |
|---|---|
| TTS + lipsync | 1100 ms |
| Eye blink | 0.5 ms |
| FK + pose (skeleton) | 1031 ms (4.6 ms/frame) |
| Morph (bs2d weight) | 19 ms |
| Render (GPU skinning) | 3430 ms (15.4 ms/frame) |
| E2E | 5580 ms Β· RTF 0.77 |
Reproduce:
cd example && python bench_vrm.py --backend gles --model-dir ..(lipsync only) β add--body <clip.npz>for lipsync + body (see Body animation above for making a clip).lipsync only sits at RTF β 0.6, lipsync + body at RTF β 0.77 β both comfortably faster than realtime. Body animation adds only the CPU FK + pose stage (
4.6 ms/frame): the skinning itself runs in the vertex shader (GPU), so render cost is essentially unchanged (15 ms/frame). Render still dominates (~60β70 % of E2E) and fits the 33 ms budget of 30 fps, so a streaming setup (TTS chunk β₯ render pipeline) stays realtime. The atlas route is lighter still (RTF ~0.25): 2D sprite paste vs full 3D mesh rasterization.Other backends (RK3576, lipsync only): numpy β pure-CPU rasterizer, ~11 s/frame (RTF ~345), a no-GPU fallback for stills only. egl (moderngl) β needs desktop GL 3.3, which ARM GPU drivers cap at 3.1 (Panfrost); x86 GPU / Mesa llvmpipe only, does not run on RK3576.
Coverage / Limitations
- Korean MFA G2P compatible β covers the phones the Ppaso G2P emits.
- ~11 unmapped phones out of 118 (loanword
f/v/z, rarecΚ°/dΚ/Κ/Κ, glideΙ₯, labializedΙΎΚ·). - Decomposition ratios (j-glide 30:70, w-glide 25:75) are a first pass.
- VRM route β VRM 0.x only; standard
blendShapeMasterviseme presets expected. VRM has no dedicated closed-lip viseme (see README_bs2d.md).
Resources
- Lipsync module:
example/ppaso_lipsync/β visemes / extractor / keyframes / interpolate / eyes / bs2d / pipeline. API docs:example/ppaso_lipsync/README.md - VRM render module:
example/ppaso_vrm/β loader / scene / anim / anim_body (skeleton) / raster (numpy + egl/moderngl + gles + gles_skinned). - bs2d mapping detail: README_bs2d.md
- Examples:
07_avatar_video.py(atlas) Β·08_vrm_video.py(blendshape) Β·09_vrm_live.py(blendshape + body) Β·10_vrm_live_streaming.py(streaming + body) Β·06_streaming_with_lipsync.pyΒ·bench_lipsync.py(atlas) Β·bench_vrm.py(blendshape) - Body-clip tool:
example/tools/fbx_to_clip.pyβ FBX β.npzbody clip - Sample assets:
example/avatar/(2D) Β·example/avatar_vrm/(3D VRM)
License
Apache 2.0 β same as Ppaso-TTS main.