ppaso-tts-v1 / README_lipsync.md
akamotaco's picture
feat(vrm): λ°”λ”” μ• λ‹ˆ 슀트리밍 + idle 클립 미동봉 (Mixamo λΌμ΄μ„ μŠ€)
cdf615f verified
|
Raw History Blame Contribute Delete
10.9 kB

Ppaso-TTS Lipsync β€” companion animation module

ν•œκ΅­μ–΄: README_lipsync_ko.md Β· Main: README.md

Blendshape mapping detail: README_bs2d.md

Pure-numpy companion module ppaso_lipsync (in example/) turns Ppaso-TTS output into per-frame 2D mouth + 1D eye-blink animation curves. A single lipsync curve drives two render routes:

  • atlas β€” 2D sprite swap (nearest viseme β†’ one mouth sprite). 07_avatar_video.py
  • blendshape β€” 3D morph-target blend via bs2d. 08_vrm_video.py

Demo

Blendshape β€” 3D VRM, lipsync + body idle (samples/08_vrm_lipsync.mp4, download):

Atlas β€” 2D sprite (samples/06_lipsync_intro.mp4, download):

μ—£μ§€λ””λ°”μ΄μŠ€μš© κ²½λŸ‰ν™” 립싱크 μ• λ‹ˆλ©”μ΄μ…˜ κΈ°λŠ₯. λ™μΌν•œ μ»€λΈŒκ°€ 2D μŠ€ν”„λΌμ΄νŠΈμ™€ 3D λΈ”λ Œλ“œμ‰μž… μ• λ‹ˆλ©”μ΄μ…˜μ„ λͺ¨λ‘ κ΅¬λ™ν•©λ‹ˆλ‹€. (One curve drives both 2D sprite and 3D blendshape animation.)

How it works

TTS phones + durations
        β”‚
        β–Ό
  viseme decompose   ← 9 anchor (A/I/U/E/O/EU/EO/M/sil) + j/w glide,
        β”‚              palatalized / alveolar / sibilant / velar / glottal
        β–Ό
  keyframes β†’ 2D xy curve     ← shared lipsync curve, x∈[-1,1] y∈[-0.1,1]
        β”‚
        β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Ί  atlas route      : nearest viseme β†’ sprite swap
        β”‚
        └──────────►  blendshape route : bs2d β†’ barycentric β†’ morph blend

  Eye blink (independent)  β†’ state_timeline {0:open,1:half,2:close}
                             + curve_timeline [0,1]

The lipsync curve generation (decompose β†’ keyframes β†’ interpolate) is shared. The two routes differ only in how they consume the (x, y) curve β€” see the atlas/blendshape distinction below.

Route A β€” atlas (2D sprite)

Each frame the curve point picks the nearest viseme and the renderer shows that one mouth sprite (sprites cannot blend). Best for 2D character art with a sprite sheet.

cd example
python 07_avatar_video.py --text "μ•ˆλ…•ν•˜μ„Έμš”." --out hi.mp4
  • Sample 2D avatar: example/avatar/ (atlas.json + sprite sheet).
  • Curve β†’ sprite mapping: atlas.json viseme_to_sprite.

Route B β€” blendshape (3D VRM)

For 3D characters with morph-target (blendshape) visemes. Here the curve goes through bs2d, which finds the triangle containing the point and returns barycentric blend weights β€” so 3 blendshapes are mixed per frame, giving continuous mouth motion (a sprite swap cannot do this).

lipsync (x,y) curve β†’ bs2d(target="vrm") β†’ {a,i,u,e,o,neutral} weight
                    β†’ VRM blendShapeMaster morph β†’ mesh deform β†’ render
  • vrm-spec β€” bs2d maps the curve onto a standard VRM 0.x avatar's 5 viseme + neutral expressions (Korean γ…γ…£γ…œγ…”γ…— β†’ VRM a/i/u/e/o).
  • Three render backends (--backend):
    • gles (default) β€” EGL + OpenGL ES 3.0. Real GPU acceleration on ARM SoCs (RK3576 Mali-G52 etc.). Needs PyOpenGL.
    • egl β€” moderngl + EGL desktop GL 3.3. For x86 GPUs / Mesa llvmpipe; ARM GPU drivers usually cap desktop GL at 3.1 so moderngl can't run there β€” use gles on-device. Needs moderngl + glcontext.
    • numpy β€” pure-numpy software rasterizer, zero GPU/driver dependency (slow β€” fallback only).
  • Bundled sample VRM: example/avatar_vrm/AvatarSample_F.vrm (VRoid 'Vita', CC0).
cd example
python 08_vrm_video.py --text "μ•ˆλ…•ν•˜μ„Έμš”." --out hi.mp4   # gles (default)
python 08_vrm_video.py --backend egl             # x86 desktop GPU
python 08_vrm_video.py --backend numpy           # no-GPU fallback

bs2d full-spec ↔ vrm-spec, the projection / barycentric stages, and how to map your own character's blendshapes (9 or more) β€” see README_bs2d.md.

Body animation (optional)

09_vrm_live.py and 10_vrm_live_streaming.py add a skeletal body idle on top of lipsync β€” GPU skinning via raster_gles_skinned. It needs a body clip (--idle clip.npz); without one the body stays at rest pose (lipsync and video still render fine).

Make a clip from any FBX skeletal animation with tools/fbx_to_clip.py (needs bpy β€” pip install bpy):

python tools/fbx_to_clip.py  idle.fbx  avatar_vrm/idle.npz
python 09_vrm_live.py --idle avatar_vrm/idle.npz

The clip is not bundled β€” animation assets such as Mixamo are royalty-free to embed in your rendered output, but their raw motion data is generally not redistributable, so supply your own. BodyClipPlayer assumes a loop-authored clip; fbx_to_clip.py drops the duplicate loop endpoint.

Quick start

from ppaso_tts import PpasoTTS
from ppaso_lipsync import generate_blink_timeline

tts = PpasoTTS('./', backend='onnx')          # or 'rknn'
wav, lipsync = tts.synthesize_with_lipsync('μ•ˆλ…•ν•˜μ„Έμš”.', fps=30)
# lipsync.xy_timeline   β†’ (n_frames, 2)  β€” feeds atlas or blendshape route
# lipsync.keyframes     β†’ viseme keyframes (debug)

blink = generate_blink_timeline(duration=lipsync.duration, fps=30, seed=42)
# blink.state_timeline / blink.curve_timeline

Streaming (chunk-to-chunk continuity is automatic):

from ppaso_tts import PpasoTTS, StreamingTTS

stream = StreamingTTS(PpasoTTS('./', backend='rknn'))
for token in llm_token_stream():
    for wav, lipsync in stream.feed_with_lipsync(token):   # emits on .!?
        play_audio(wav); animate(lipsync)
for wav, lipsync in stream.flush_with_lipsync():
    ...

Performance

Atlas route β€” RK3576 measured

End-to-end (TTS + lipsync + eye blink + sprite composite at 768Γ—1152, no ffmpeg encode). RK3576, best-of-5 @ 30 fps.

Case Input wav Backend E2E RTF Speed
short 21 chars 3.49 s RK3576 NPU 864 ms 0.247 4.0Γ— RT
ONNX CPU (4t) 897 ms 0.256 3.9Γ— RT
medium 41 chars 7.26 s RK3576 NPU 1867 ms 0.257 3.9Γ— RT
long 110 chars 16.65 s RK3576 NPU 3668 ms 0.220 4.5Γ— RT

Reproduce: cd example && python bench_lipsync.py --backend rknn --model-dir ... Lipsync curve generation itself adds < 1 ms on top of TTS.

Blendshape route β€” RK3576 measured

End-to-end (TTS + lipsync + eye blink + morph + render at 512Γ—768, no ffmpeg encode β€” same E2E definition as the atlas route above). RK3576 Mali-G52, gles backend, best-of-3 @ 30 fps. Two modes: lipsync only (static skeleton β€” 08_vrm_video.py) and lipsync + body (GPU skinning + body idle animation β€” 09_vrm_live.py).

β‘  lipsync only β€” E2E by text length:

Case wav frames E2E RTF
short 3.51 s 110 2222 ms 0.63
medium 7.26 s 223 4637 ms 0.64
long 16.31 s 497 9358 ms 0.57

β‘‘ lipsync + body β€” E2E by text length:

Case wav frames E2E RTF
short 3.51 s 110 2809 ms 0.80
medium 7.26 s 223 5580 ms 0.77
long 16.31 s 497 12314 ms 0.76

β‘’ lipsync + body β€” medium (7.26 s, 223 frames) stage breakdown:

Stage medium
TTS + lipsync 1100 ms
Eye blink 0.5 ms
FK + pose (skeleton) 1031 ms (4.6 ms/frame)
Morph (bs2d weight) 19 ms
Render (GPU skinning) 3430 ms (15.4 ms/frame)
E2E 5580 ms Β· RTF 0.77

Reproduce: cd example && python bench_vrm.py --backend gles --model-dir .. (lipsync only) β€” add --body <clip.npz> for lipsync + body (see Body animation above for making a clip).

lipsync only sits at RTF β‰ˆ 0.6, lipsync + body at RTF β‰ˆ 0.77 β€” both comfortably faster than realtime. Body animation adds only the CPU FK + pose stage (4.6 ms/frame): the skinning itself runs in the vertex shader (GPU), so render cost is essentially unchanged (15 ms/frame). Render still dominates (~60–70 % of E2E) and fits the 33 ms budget of 30 fps, so a streaming setup (TTS chunk βˆ₯ render pipeline) stays realtime. The atlas route is lighter still (RTF ~0.25): 2D sprite paste vs full 3D mesh rasterization.

Other backends (RK3576, lipsync only): numpy β€” pure-CPU rasterizer, ~11 s/frame (RTF ~345), a no-GPU fallback for stills only. egl (moderngl) β€” needs desktop GL 3.3, which ARM GPU drivers cap at 3.1 (Panfrost); x86 GPU / Mesa llvmpipe only, does not run on RK3576.

Coverage / Limitations

  • Korean MFA G2P compatible β€” covers the phones the Ppaso G2P emits.
  • ~11 unmapped phones out of 118 (loanword f/v/z, rare cΚ°/dΚ‘/ʝ/ʁ, glide Ι₯, labialized ΙΎΚ·).
  • Decomposition ratios (j-glide 30:70, w-glide 25:75) are a first pass.
  • VRM route β€” VRM 0.x only; standard blendShapeMaster viseme presets expected. VRM has no dedicated closed-lip viseme (see README_bs2d.md).

Resources

  • Lipsync module: example/ppaso_lipsync/ β€” visemes / extractor / keyframes / interpolate / eyes / bs2d / pipeline. API docs: example/ppaso_lipsync/README.md
  • VRM render module: example/ppaso_vrm/ β€” loader / scene / anim / anim_body (skeleton) / raster (numpy + egl/moderngl + gles + gles_skinned).
  • bs2d mapping detail: README_bs2d.md
  • Examples: 07_avatar_video.py (atlas) Β· 08_vrm_video.py (blendshape) Β· 09_vrm_live.py (blendshape + body) Β· 10_vrm_live_streaming.py (streaming + body) Β· 06_streaming_with_lipsync.py Β· bench_lipsync.py (atlas) Β· bench_vrm.py (blendshape)
  • Body-clip tool: example/tools/fbx_to_clip.py β€” FBX β†’ .npz body clip
  • Sample assets: example/avatar/ (2D) Β· example/avatar_vrm/ (3D VRM)

License

Apache 2.0 β€” same as Ppaso-TTS main.