Step 1
Open the studio
Go to FairStack and open the talking head workspace.
Eight selectable avatar and lip-sync models. Drive a face with audio, keep identity stable, and pay per second — not a seat license.
Kling Avatar Standard · Prompt: Neon-lit portrait, subtle head turn, dialogue-ready framing
A talking-head model animates a face from an image plus audio (or lip-syncs an existing video). FairStack surfaces audio-driven and lip-sync models under one balance.
8 selectable talking head models · 113 on the platform
Every model below is non-hidden, non-fallback, generation-category — part of the 113 customers can actually run.
$0.0015/sec
MuseTalk 1.5 is a lip synchronization model that adds natural mouth movement to existing images or video at an ultra-affordable per-second rate. The model specializes in lip sync only, driving mouth movements from audio input without generating full body motion or head movement, keeping the processing focused and the cost extremely low. With per-second billing at $0.00111 per second, it is a low-cost lip sync model available on the platform. A full minute of lip sync costs approximately $0.067, making it practical for high-volume production, batch processing, and applications where hundreds or thousands of clips need lip synchronization. The model works with both static images and existing video. Compared to premium lip sync models like Sync Lipsync 2.0 Pro at $0.083 per second, MuseTalk 1.5 is at a much lower per-use price with proportionally simpler output. Against full talking head models like Kling Avatar at $0.25, it costs a fraction but only provides mouth movement rather than full facial and body animation. Best suited for budget lip sync at scale, adding speech to portrait photos, and high-volume video lip synchronization where ultra-low cost matters most. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.017/sec
Kling LipSync Audio-to-Video is Kuaishou's lip synchronization model that matches video lip movements to provided audio input. The model analyzes the audio's phonemes and drives realistic lip movements, jaw motion, and subtle facial expressions in the source video to match the provided speech, producing natural-looking synchronized output. Powered by Kling AI's video generation technology and delivered via fal.ai, the model preserves the identity and visual characteristics of the person in the video while modifying only the mouth and jaw area. The synchronization handles various speaking speeds and accents, with best results on clear, well-recorded audio. Compared to budget lip sync models like MuseTalk at $0.00111 per second, Kling LipSync delivers higher synchronization accuracy and more natural mouth shapes. Against full talking head generation models, it focuses specifically on accurate lip sync for existing video rather than generating new video content. Best suited for video dubbing, lip sync for translated content, and music video creation where matching lip movements to audio produces convincing synchronized video. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.048/sec
Kling Avatar Standard is Kuaishou's talking head generation model that creates video of a speaking avatar from an image and audio input. The model generates natural lip synchronization, head movement, and facial expressions driven by the provided audio, producing realistic presenter-style video from a single portrait photo. At $0.25 per generation, it provides professional-quality talking head content using Kling's established video generation technology. The model handles various portrait styles and produces natural-looking mouth movements, head tilts, and eye movements that make the avatar appear engaged and lifelike. Audio quality directly impacts lip sync accuracy. Compared to budget lip sync models like MuseTalk at $0.00111 per second, Kling Avatar delivers substantially higher visual quality with fuller head and face animation rather than just mouth movement. Against premium talking head services that charge per-minute subscription rates, FairStack's per-generation pricing provides better value for occasional use. Best suited for talking head videos, audio-driven avatar presentations, and content creation where natural-looking avatar video is needed from a portrait photo. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.096/sec
Kling Avatar Pro is Kuaishou's talking head model from the Kling ecosystem, distinguished by its multi-character support and flat-rate pricing at $0.25 per video. This flat-rate model makes it cost-effective for longer clips where per-second pricing from competitors would accumulate significantly, while multi-character capability enables more complex avatar scenes. The model achieves a lip sync score of 0.82, delivering reliable synchronization quality that sits in the mid-tier range. It inherits the visual quality standards of the broader Kling ecosystem, maintaining consistency for workflows that also use Kling image and video generation models. Compared to OmniHuman v1.5 which achieves the best emotional expressiveness at per-second pricing, Kling Avatar Pro offers multi-character scenes and flat-rate economics that serve different use cases. Against single-character models like Pixverse Lipsync and Sync Lipsync, Kling Avatar Pro is the only option supporting multiple characters in a single generation. Best suited for multi-character avatar scenes, longer talking head content where flat-rate pricing is advantageous, and workflows already within the Kling ecosystem requiring visual consistency across image and video generation. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.160/sec
Sync Lipsync 3.0 Image-to-Video animates a single still photo into a talking-head video driven by your audio track. Unlike the v2 family (which requires existing video), v3 needs only a photo and a voice clip — upload both URLs and get a lip-synced avatar back. Premium tier at $0.1333 per output second ($8/min), positioned above Sync Lipsync 2.0 Pro ($5/min).
$0.240/sec
WAN 2.2 Speech-to-Video is Alibaba's talking head generation model that creates realistic video from audio input and a reference portrait image. The model synthesizes natural facial expressions, lip movements, head motion, and subtle micro-expressions synchronized to the provided speech audio, producing convincing presenter-style video. The model generates natural-looking talking head content that goes beyond basic lip sync, adding head tilts, eye movements, brow expressions, and other facial dynamics that make the output more lifelike. Audio-to-video synchronization is tight, with lip movements closely tracking the phonemes in the speech input. Compared to basic lip sync models that only animate the mouth area, WAN 2.2 Speech-to-Video generates fuller facial animation for more engaging output. Against text-to-video talking head models, the audio-driven approach preserves the speaker's actual voice and delivery. Best suited for talking head video creation, podcast video generation, and virtual presenter content where realistic facial animation synchronized to speech audio creates convincing presenter videos. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.240/sec
EchoMimic V3 is a audio-driven model on FairStack. Pay $0.240/sec with a transparent 20% margin — no subscription.
$0.600/req
InfiniteTalk is an audio-driven talking head model hosted on RunPod's serverless infrastructure. It generates avatar videos from a portrait photo and audio input, offering both 480p and 720p resolution tiers at flat-rate pricing of $0.25 and $0.50 per video respectively. The RunPod hosting ensures reliable uptime and consistent availability. The model handles standard lip synchronization driven by audio input, producing talking head videos suitable for basic avatar content. Flat-rate pricing means cost is predictable regardless of video duration, which benefits longer clips where per-second models become expensive. Generation is slower than some competitors but delivers reliable results. Compared to Pixverse Lipsync and Sync Lipsync which charge per second, InfiniteTalk's flat-rate pricing becomes advantageous for clips longer than a few seconds. Against OmniHuman v1.5 which leads in emotional expressiveness, InfiniteTalk offers a simpler but more predictable option at a lower price. The model is best chosen when RunPod infrastructure reliability is a workflow requirement. Best suited for audio-driven character animation, longer talking head clips where flat-rate pricing provides savings, and workflows that prioritize RunPod's infrastructure reliability. Available on FairStack at infrastructure cost plus a 20% platform fee.
One balance. Pick a model. Pay the model cost plus 20%.
Step 1
Go to FairStack and open the talking head workspace.
Step 2
Choose from the selectable list on this page — capability flags decide who appears.
Step 3
Run the job. Credits never expire if you stop mid-project.
Step 4
The same balance covers image, video, voice, and music.
This hub lists every selectable talking-head model. /avatars focuses on audio-driven avatar workflows; /lip-sync focuses on lip-sync models.
Still have questions? We're here to help.