MuseTalk API for Python
Python client for a MuseTalk-class lip sync API: real-time talking-head generation through a hosted endpoint
Run MuseTalk in the browser → View on GitHub
Install
pip install git+https://github.com/musetalk-dev/musetalk-api.git
export SYNEXA_API_KEY="sk-..."
Quickstart
import musetalk_api
output = musetalk_api.run({
"image_url": "https://example.com/input.png",
"audio_url": "https://example.com/input.png"
})
print(output)
Hosted models
- veed/fabric-1.0 — Fabric 1.0 turns a photo plus an audio track into a talking-head video. ($0.08 / run)
About MuseTalk
MuseTalk is Tencent Music Lyra Lab's real-time lip sync model: it masks the lower face, encodes it with a Stable Diffusion VAE and inpaints the mouth in one U-Net pass conditioned on Whisper audio features, reaching 30 fps and above on a V100. This client calls a hosted talking-head endpoint (veed/fabric-1.0) from Python with one dependency: send a portrait URL and an audio URL, get back a lip-synced video at $0.08 per run. MuseTalk's own weights can be self-hosted from the official repository.
FAQ
Is there a MuseTalk API?
Not from Tencent; MuseTalk is released as open weights. This client exposes the same lip sync capability through a hosted talking-head endpoint (veed/fabric-1.0) that you call over HTTPS.
How much does the MuseTalk API cost?
The hosted veed/fabric-1.0 model is $0.08 per run. Billing is per prediction; there is no hourly GPU charge.
Can I run MuseTalk without a GPU?
With this client, yes: generation happens on the hosted service and your code only makes HTTP requests. Self-hosting MuseTalk needs a CUDA GPU; its real-time figures assume a V100-class card.
Does this client work with the original MuseTalk repo or ComfyUI?
No. It does not load the TMElyralab/MuseTalk checkpoints and it is not a ComfyUI node. It is a network client for the hosted endpoint. If you need MuseTalk's real-time video-to-video path, run the official repository locally.
What input formats does it accept?
image_url (a portrait in .jpg/.png/.webp with one clear, forward-facing face) and audio_url (.mp3/.wav/.flac/.m4a/.ogg), plus an optional resolution. The output is a video URL. Video input is not accepted by the hosted endpoint.
Is this the official MuseTalk SDK?
No. This is an independent, community-maintained client and is not affiliated with Tencent Music Entertainment or VEED. The official project lives at https://github.com/TMElyralab/MuseTalk.