Step 1
Open the studio
Go to FairStack and open the ai video workspace.
Generate video from text, images, or existing clips with 49 selectable models. Pay per second with a transparent 20% margin — no plans, no tiers, credits never expire.
Sora 2 · Prompt: Slow dolly along a coastal cypress road at golden hour, soft marine haze, cinematic 35mm look
FairStack AI video is a single API and studio for generating clips from text, images, or video inputs. You pick a model, set duration and resolution, and pay the model cost plus 20%. FairStack never trains on your content.
49 selectable ai video models · 113 on the platform
Workflows
Every model below is non-hidden, non-fallback, generation-category — part of the 113 customers can actually run.
$0.123/sec
Seedance 2.0 Mini is ByteDance’s most affordable Seedance 2 model, built for high-volume video generation where cost matters more than maximum resolution. Like the rest of the Seedance 2 family it is multi-mode from a single endpoint: drive it with a text prompt, a single first frame, first+last frames for interpolation, or up to nine reference images plus video/audio references — FairStack auto-detects the mode from your input. It renders 480p or 720p clips from 4 to 15 seconds with native audio generation, and supports seven aspect ratios (including 21:9 and adaptive). At roughly $0.0625 per second it is the value pick of the Seedance 2 lineup, trading the 1080p/4K ceiling of the standard model for materially lower cost. Best for social-media shorts, rapid iteration, storyboard animation, and any workflow that generates a lot of video. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.198/sec
Seedance 2.0 Fast is a text to video model on FairStack. Pay $0.198/sec with a transparent 20% margin — no subscription.
$0.060/sec
Veo 3.1 Lite Image-to-Video is the most affordable entry into Google's Veo line, always generating video with audio. At $0.05/s for 720p with audio it brings Veo-family motion and sound quality to budget-sensitive workflows. 4/6/8s durations, 720p/1080p output, aspect auto-aligned to the input image.
$0.246/sec
Seedance 2.0 is a text to video model on FairStack. Pay $0.246/sec with a transparent 20% margin — no subscription.
$0.084/sec
Kling 2.5 Turbo Pro Image-to-Video is the value sweet spot of the entire i2v catalog — strong motion and fidelity at just $0.07/s. It derives output aspect directly from your input image (no reframing), supports 5/10s durations, an optional tail frame for controlled endings, and cfg_scale for prompt adherence.
$0.084/sec
WAN 2.6 T2V is Alibaba's latest text-to-video generation model, representing the most advanced entry in the WAN family. It generates 5-second video clips at 720p with notably strong artistic style and visual quality, making it particularly well-suited for stylized and creative content rather than strict photorealism. The model benefits from Alibaba's extensive research in video generation, delivering improved temporal coherence and motion quality over earlier WAN versions. It handles artistic prompts with distinctive flair, producing outputs with painterly qualities, rich color grading, and creative visual interpretations that set it apart from purely photorealistic competitors. Compared to Kling 3.0 which focuses on cinematic realism, WAN 2.6 T2V leans toward artistic expression and stylized output. Against Veo 3 which costs $0.20 per second, WAN 2.6 at $0.50 per generation offers a different aesthetic that may better serve creative and artistic content needs. The model is hosted on RunPod for reliable delivery. Best suited for artistic video content, music video concepts, stylized social media clips, and creative projects where visual artistry matters more than photorealism. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.162/sec
Kling 3.0 Pro from Kuaishou is the highest-rated video generation model globally by Elo score, delivering the best motion quality (0.90), excellent camera control (0.85), and support for clips up to 15 seconds. When video quality is the top priority and cost is secondary, this is the definitive choice. The model produces smooth, naturalistic movement that avoids the common artifacts of AI video — jitter, ghosting, and unnatural physics. Camera control follows complex instructions including pan, zoom, dolly, and orbital movements with high fidelity. Up to 15 seconds of continuous video is supported, the longest duration among top-tier models. Visual quality (0.88 score) is consistently high across subjects and scenes. Compared to Veo 3.1 Quality (the only model with native audio), Kling 3.0 Pro offers longer maximum duration and lower cost per second. Against Grok Imagine T2V (Elo #3), it delivers noticeably superior motion quality and camera control at a higher price. Direct pricing from Kling would cost approximately $0.40 per generation — the FairStack price of approximately $0.135 per second represents significant savings at infrastructure cost plus 20%. Best suited for premium video production where quality is paramount, long-form clips up to 15 seconds, cinematic content requiring smooth camera work, and professional video workflows. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.0051/sec
Bria Video Background Removal is Bria AI's specialized model for removing backgrounds from video content frame by frame while maintaining temporal consistency. The model separates foreground subjects from their backgrounds with clean edges, preventing the flickering and edge artifacts that occur when processing video frames independently. With per-second billing at $0.14 per second, costs scale proportionally with video duration. Bria AI distinguishes itself through commercially safe output, as the company trains its models exclusively on properly licensed data, eliminating copyright concerns for commercial use. The temporal consistency ensures smooth, flicker-free edges throughout the video. Compared to chroma-key green screen setups, AI background removal works with any footage regardless of the original background, eliminating the need for controlled filming environments. Against frame-by-frame image background removal tools, the video-native approach provides superior temporal consistency. Best suited for video background removal, green screen alternatives, and commercial content production where Bria's commercial safety and temporal consistency matter. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.018/sec
Grok Imagine T2V from xAI is ranked Elo #3 for video generation, delivering strong overall quality at an exceptional price point of $0.10 per generation. It represents the best value proposition in the premium video category — priced below Kling Pro or Veo while still producing high-quality output suitable for professional use. The model generates up to 10 seconds of video with good motion quality (0.84) and strong visual quality (0.85). The balanced capability profile means it handles diverse subjects and scenes without significant weaknesses in any particular area. As a newer entrant in the video generation space, it benefits from xAI's latest research while building a growing track record in production workflows. Compared to Kling 3.0 Pro (Elo #1), Grok T2V trades some motion smoothness and camera control precision for a lower cost per generation. Against Veo 3.1 Fast (Elo #2), it offers similar quality at a competitive price with up to 10-second duration. Versus Runway Gen-4 Turbo ($0.06), it costs more but delivers meaningfully higher quality. The Elo #3 ranking at $0.10 per generation makes it the clear choice when quality and value must both be strong. Ideal for general-purpose video generation, best-value premium video, 10-second clips, and any budget-conscious video workflow that still demands high quality. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.018/sec
Grok Imagine I2V is xAI's image-to-video model that animates still images with notably strong motion quality relative to its price point. At $0.10 per generation, it delivers animation quality that competes with models costing two to three times as much, making it one of the best value propositions in the image-to-video category. The model produces natural-looking animation with good temporal coherence and smooth motion transitions. Source image fidelity is well-preserved, with the AI adding movement that feels organic rather than artificial. The Grok architecture underpinning the model benefits from xAI's investment in multimodal understanding, which translates to better interpretation of what should move and how within a given image. Compared to premium options like Kling 2.6 I2V at $0.275 or Runway Gen-4 at $0.50, Grok Imagine I2V sacrifices some visual refinement but delivers surprisingly competitive motion quality. For social media content, prototyping, and any workflow where budget matters more than maximum polish, it represents excellent value. Best suited for value-oriented video animation, social media video content, and good-quality motion on a budget. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.024/sec
P-Video (Pruna) is a text to video model on FairStack. Pay $0.024/sec with a transparent 20% margin — no subscription.
$0.024/sec
P-Video I2V (Pruna) is a image to video model on FairStack. Pay $0.024/sec with a transparent 20% margin — no subscription.
$0.031/sec
Hailuo 2.3 Fast I2V is MiniMax's budget-friendly image-to-video model optimized for fast generation times. The model animates still images into video with quick turnaround at one of the lower price points in the I2V category, making it practical for high-volume production and rapid prototyping. At $0.08 per generation, it is among a low-cost I2V models available, approximately priced below mid-tier options. The Fast designation indicates shorter generation times compared to Pro-tier alternatives, though with proportionally lower visual quality. For quick animations, social media content, and draft work, the quality is adequate. Compared to Hailuo 02 Pro I2V at $0.225, the Fast variant saves 65% with faster generation, though with visibly lower quality in complex scenes. Against other budget I2V options like Runway Gen-3 I2V at $0.075, it offers competitive pricing with the benefit of Hailuo 2.3's newer architecture. Best suited for budget image animation, quick animation drafts, and fast I2V workflows where speed and cost matter more than maximum quality. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.042/req
Vidu Q3 Turbo T2V is Vidu's budget text-to-video model optimized for fast generation at significantly lower cost than the standard Q3 tier. The model supports up to 16 seconds of video duration with built-in audio generation, offering one of the longest maximum durations among budget video models. With per-second billing at $0.077 per second, it is substantially priced below most video generation alternatives. A 5-second clip costs approximately $0.39, while a full 16-second clip costs about $1.23. Multiple resolution options are available, with 720p providing the best quality-to-cost balance. Compared to standard Vidu Q3 at $0.154 per second, the Turbo variant halves the per-second cost while maintaining usable quality for drafts and social media. Against premium models like Kling 2.6 at $0.275 flat, it costs less per generation for longer content. Best suited for draft video previews, social media content, and high-volume video production where budget-friendly per-second pricing enables cost-effective generation at scale. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.042/req
Vidu Q3 Turbo I2V is Vidu's budget image-to-video model that animates still images into video at turbo speed with significantly lower cost than the standard Q3 tier. Like the T2V variant, it supports up to 16 seconds of duration with built-in audio generation. With per-second billing at $0.077 per second, it offers the same cost advantages as the T2V Turbo variant applied to image animation. The model preserves source image details while adding motion, producing usable animations at a predictable per-use price of premium I2V models. Best results are achieved at 720p resolution. Compared to standard Vidu Q3 I2V at $0.154 per second, the Turbo variant halves costs while maintaining adequate quality for drafts and social content. Against premium I2V models like Kling 2.6 I2V at $0.275, the dramatic cost difference enables high-volume animation production. Best suited for quick image animations at scale, social media content creation, and batch processing workflows where budget-friendly per-second pricing makes high-volume animation economically viable. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.060/req
PixVerse V5 Image-to-Video is the premium at a low price benchmark winner — clean, stylized motion with built-in style presets (anime, 3D, clay, comic, cyberpunk). Resolution tiers 360p–1080p, 5 or 8s durations (8s doubles cost; 1080p caps at 5s). Output aspect derives from your input image.
$0.060/sec
Grok Imagine Edit Video is xAI's AI-powered video editing model that applies targeted modifications to existing video content using natural language instructions. Users describe the desired changes, and the model transforms the video accordingly while preserving unmodified elements and maintaining temporal coherence across frames. With per-second billing at $0.07 per second, it is one of the most affordable per-second video editing options available. The Grok architecture benefits from xAI's investment in language understanding, translating editing instructions into precise visual modifications. Per-second billing keeps costs proportional to the source video duration. Compared to flat-rate V2V editing models like Runway Aleph at $0.20, Grok's per-second pricing is more economical for shorter clips. Against other Grok Imagine models, the Edit Video variant brings the same strong value proposition to the editing category. Best suited for AI-driven video editing, per-second editing workflows, and cost-effective video modifications where affordable per-second billing matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.067/sec
Kling AI Avatar v2 Standard is the cost-efficient avatar tier ($0.0562/s). Provide an image and an audio track and it generates a synchronized talking avatar — realistic humans, animals, cartoons, or stylized characters. Audio waveform drives facial animation and lip timing.
$0.084/req
Vidu Q3 T2V is Vidu's text-to-video model with flexible per-second billing, generating balanced-quality video from text prompts at proportional cost. The model delivers reliable output with good motion quality and scene coherence, providing a solid option for general-purpose video generation with the advantage of duration-flexible pricing. With per-second billing at $0.154 per second, users pay only for the exact duration they need. This pricing structure is advantageous for short clips where flat-rate models may overprice the output, and remains competitive for longer sequences. The model handles a variety of prompt types with consistent quality. Compared to flat-rate models like Kling 2.6 at $0.275, Vidu Q3's per-second pricing is more economical for shorter clips but can add up for longer sequences. Against other per-second models like Kling V3 Standard at $0.168 per second, Vidu Q3 offers a modest cost advantage. The model provides a good balance of quality and pricing flexibility. Best suited for flexible-duration video generation and general text-to-video workflows where per-second billing provides cost control. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.084/req
Vidu Q3 I2V is Vidu's image-to-video model with flexible per-second billing, animating still images into video at proportional cost. The model preserves source image details while adding natural-looking motion, providing a reliable option for general-purpose image animation with the advantage of duration-flexible pricing. With per-second billing at $0.154 per second, users pay only for the animation duration they need. This pricing structure benefits short clips where flat-rate models may overprice, and remains competitive for standard-length animations. The model handles various source image types including photographs and illustrations. Compared to flat-rate I2V models like Kling 2.6 I2V at $0.275, Vidu Q3's per-second pricing is more economical for short clips. Against other per-second options like Kling O3 Standard I2V at $0.168 per second, Vidu Q3 offers a modest cost advantage per second. Best suited for flexible-duration image animation and general I2V workflows where per-second billing provides cost control and pricing transparency. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.101/sec
Kling O3 Standard Reference-to-Video is Kuaishou's balanced reference-guided video generation model in the O3 generation. The model creates new video matching the visual style of provided reference clips at a moderate per-second price point, making reference-guided generation accessible for everyday workflows. With per-second billing at $0.168 per second, it offers 25% savings compared to the O3 Pro R2V tier while maintaining good style matching quality. The model interprets the visual characteristics of reference clips including color palette, lighting, motion style, and composition, producing new content that feels visually related to the source. Compared to O3 Pro R2V at $0.224 per second, the Standard tier trades some precision in style matching for meaningful cost savings. Against text-only video models, the reference-guided approach remains significantly more effective at achieving consistent visual style. Best suited for reference-guided video at a balanced price, everyday reference matching, and workflows where visual consistency with existing content matters but premium pricing is not justified. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.101/sec
Kling O3 Standard V2V Edit is Kuaishou's balanced video-to-video editing model in the O3 generation, offering good editing quality with per-second billing at a moderate price point. The Standard tier delivers reliable video editing results without the premium pricing of the Pro tier. With per-second billing at $0.168 per second, it offers 25% savings compared to the O3 Pro V2V Edit while maintaining good editing quality. The model handles common video editing tasks including style adjustment, element modification, and scene transformation with frame-to-frame consistency. Compared to O3 Pro V2V Edit at $0.224 per second, the Standard tier trades some editing refinement for meaningful cost savings. Against non-per-second editors like Luma Modify at $0.15, the per-second model provides more flexible pricing based on source video duration. Best suited for balanced video editing, general V2V transformations, and everyday editing workflows where good quality at a reasonable per-second price provides practical value. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.120/sec
Sora 2 Image-to-Video is OpenAI's widely-searched video model, animating a still image with synchronized audio. Supports 4/8/12/16/20-second durations, 720p output, and 16:9 or 9:16 aspect derived from your input. Billed at $0.10/s — strong quality and brand recognition at an accessible price.
$0.120/sec
Kling 3.0 Standard is the cost-optimized variant of Kuaishou's video generation model, sharing the same architecture and maximum 15-second duration as the Pro variant while delivering good quality at a lower price point. It bridges the gap between budget video models and the premium Kling Pro tier. The model produces solid video with good motion quality (0.82 score), maintaining natural-looking movement and reasonable camera following. Supporting the same 15-second maximum as Pro means users can generate longer clips without the premium cost. The Kling architecture provides consistent visual coherence throughout the clip duration. Compared to Kling 3.0 Pro, the Standard variant shows a noticeable quality gap in motion smoothness, fine detail, and camera precision. Against Runway Gen-4 Turbo ($0.06/5s), it offers longer duration and higher overall quality at a higher price. Versus Hailuo 2.3 Fast ($0.08/6s), it provides the longer 15-second option with better motion quality. It occupies the practical middle ground for users who need more than budget quality but less than premium pricing. Best suited for budget premium video work, longer clips where 15-second duration matters, good-enough quality workflows, and projects where Kling Pro's cost premium is not justified. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.120/sec
HeyGen Avatar 4 turns a photo with a clear face into a premium talking-avatar video. Provide text for the avatar to speak (with a voice) or supply your own audio for lip-sync. Supports stable/expressive talking styles, captions, backgrounds, 360p–1080p output, and aspect ratios 16:9/9:16/4:5/5:4/1:1. Billed $0.10/s.
$0.120/sec
Kling 3.0 Motion Control is a video to video model on FairStack. Pay $0.120/sec with a transparent 20% margin — no subscription.
$0.126/sec
WAN 2.6 I2V is the latest and most capable image-to-video model in Alibaba's WAN family at $0.35 per generation. It delivers the best motion quality and temporal coherence of any WAN model, representing the current peak of the WAN I2V architecture. The model produces video with noticeably smoother, more natural motion than earlier WAN versions. Temporal coherence is excellent, maintaining visual consistency of subjects and scenes throughout the clip without the flickering or morphing artifacts that affect lesser models. Source image fidelity is preserved with the best detail retention in the WAN family. Compared to WAN 2.2 I2V at $0.30, the 2.6 version delivers visible improvements in motion quality and coherence for $0.05 more. Against Runway Gen-4 at approximately $0.50 which competes in the premium I2V space, WAN 2.6 offers strong quality at a lower price. Kling v2.1 Pro I2V provides 1080p output but the comparison depends on whether resolution or motion quality matters more. Best suited for premium image animation, high-fidelity video from photographs, content requiring the best WAN motion quality, and workflows where the WAN 2.6 quality improvement justifies the premium over WAN 2.5. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.134/sec
Kling O3 Pro Reference-to-Video is Kuaishou's premium reference-guided video generation model in the O3 generation. The model creates new video content that matches the style, mood, and visual characteristics of provided reference clips, enabling consistent video series and brand-coherent content at premium quality. With per-second billing at $0.224 per second, the Pro tier delivers the highest-quality reference matching in the Kling O3 lineup. Users upload a reference video and the model generates new footage that maintains visual consistency with the source material, including color palette, lighting style, motion dynamics, and compositional approach. Compared to Kling O3 Standard R2V at $0.168 per second, the Pro tier delivers more precise style matching with better handling of subtle visual characteristics. Against text-only video models where style is described rather than demonstrated, reference-guided generation provides significantly tighter control over the visual output. Best suited for premium reference-guided video production and style-matched content where maintaining visual consistency with existing material is critical. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.134/sec
Kling O3 Pro V2V Edit is Kuaishou's premium video-to-video editing model in the O3 generation, delivering maximum-quality AI-powered edits to existing video content. The Pro tier produces the most refined editing output in the Kling V2V lineup, handling complex modifications while preserving temporal coherence and visual quality. With per-second billing at $0.224 per second, the Pro tier commands the highest price in the Kling V2V editing range. The premium reflects the compute required for high-fidelity video editing that maintains frame-to-frame consistency while applying significant visual modifications. The model handles style changes, element modification, and scene adjustments. Compared to Kling O3 Standard V2V Edit at $0.168 per second, the Pro tier delivers more refined edits with better handling of complex transformations. Against non-Kling V2V editors, it benefits from Kling's strong motion quality reputation applied to editing. Best suited for premium video editing and maximum-quality V2V transformations where the highest editing fidelity justifies the premium per-second pricing. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.138/sec
Kling AI Avatar v2 Pro delivers enhanced facial detail and smoother lip-sync precision over the Standard tier, at $0.115/s. Image + audio in, synchronized talking avatar out. The premium choice when avatar realism matters.
$0.168/sec
Kling V3 Standard First+Last Frame is Kuaishou's keyframe interpolation model from the latest V3 generation. Users provide two images representing the desired start and end states of a video, and the model generates smooth, natural motion to transition between them. This gives precise creative control over both the opening and closing frames of the generated content. With per-second billing at $0.168 per second, users pay proportionally for the transition duration they need. The V3 generation delivers improved interpolation quality over earlier Kling versions, with more natural motion paths and better handling of complex transformations between visually different keyframes. The model works well for morphing effects, scene transitions, and controlled animations. Compared to text-to-video models where camera movement and scene progression are left to the AI's interpretation, first-and-last-frame generation provides deterministic control over the video's narrative arc. Against Veo 3.1 Fast FLF at $0.10 per second, Kling V3 Standard FLF is moderately more expensive but benefits from Kling's strong motion generation reputation. Best suited for controlled video transitions, morph effects, and start-to-end animation sequences. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.174/sec
HappyHorse 1.1 is Alibaba's flagship video model, a strict upgrade over HappyHorse 1.0: lower price, built-in native audio with Foley, and multilingual lip-sync. The text-to-video mode renders 3-15 second clips at 720p or 1080p across nine aspect ratios (16:9 through 9:21). Served from Kie.ai at $0.1125/s (720p) or $0.145/s (1080p) with an automatic fal.ai fallback. Best all-rounder when synchronized audio and speech matter.
$0.174/sec
HappyHorse 1.1's image-to-video mode animates a single still image into a 3-15 second clip with synchronized native audio, Foley, and multilingual lip-sync. Aspect ratio follows the input image; output is 720p or 1080p. Served from Kie.ai at $0.1125/s (720p) or $0.145/s (1080p) with an automatic fal.ai fallback — materially priced below the 1.0 generation it replaces as the default.
$0.174/sec
HappyHorse 1.1's reference-to-video mode composes a new 3-15 second scene guided by 1-9 reference images — characters, products, or styles — with native audio, Foley, and multilingual lip-sync. Supports nine aspect ratios at 720p or 1080p. Served from Kie.ai at $0.1125/s (720p) or $0.145/s (1080p) with an automatic fal.ai fallback.
$0.180/req
PixVerse V5.6 T2V is PixVerse's latest text-to-video model, known for producing high-quality video with a distinctive creative and artistic sensibility. The model excels at stylized content, generating video with strong visual aesthetics, rich color palettes, and compositional choices that lean toward artistic expression rather than strict photorealism. At $0.45 per generation, it sits in the premium tier alongside models from Runway and Luma. The price reflects the model's strong visual quality and its particular strength in creative and artistic content. The V5.6 architecture represents significant improvements over earlier PixVerse versions in both quality and consistency. Compared to photorealism-focused models like Kling 2.6, PixVerse V5.6 is stronger at stylized, artistic, and creative content but may not match the natural look of dedicated photorealistic generators. Against Runway Gen-4 at $0.50, it offers similar pricing with a different creative strength profile. Best suited for creative video content, artistic video production, and visually appealing output where stylistic quality matters. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.180/req
PixVerse V5.6 I2V is PixVerse's latest image-to-video model, known for generating animation with strong creative and artistic visual quality. The model adds motion to source images while maintaining the artistic sensibility that characterizes PixVerse output, producing animated content with rich colors, strong compositions, and visually engaging movement. At $0.45 per generation, it sits in the premium tier of the I2V market. The price reflects the model's strong artistic output quality and the V5.6 generation's improvements in motion coherence and source image preservation. The model is particularly effective at animating illustration-style and artistic source images. Compared to photorealism-focused I2V models like Kling 2.6 I2V at $0.275, PixVerse V5.6 excels at creative and stylized animation but may not match the natural look of dedicated photorealistic models. Against the previous PixVerse V4.5 at $0.20, the V5.6 delivers notably better quality at a higher price. Best suited for creative image animation, artistic motion content, and projects where visual artistry matters more than photographic accuracy. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.180/req
PixVerse V5.6 Transition is PixVerse's specialized video transition model that creates smooth morphing effects between two or more images. The model generates creative visual transitions, seamlessly blending one image into another with fluid motion and artistic effects that go beyond simple crossfades. At $0.45 per generation, it matches the PixVerse V5.6 premium pricing tier. The model leverages the V5.6 architecture's strength in creative and artistic output to produce transitions that are visually compelling rather than mechanically interpolated. The morphing effects can handle significant visual differences between source images. Compared to first-and-last-frame models like Kling V3 FLF that create motion between two keyframes, PixVerse Transition focuses specifically on creative morphing effects rather than natural scene progression. This makes it better suited for stylistic transitions in music videos, presentations, and creative content. Best suited for image-to-image transitions, creative morphing effects, and visual effects production where artistic transitions enhance the content. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.180/sec
Veo 3.1 Fast Image-to-Video delivers the Veo 3.1 architecture optimized for speed and cost. It animates a still image with native audio at $0.15/s (with audio) or $0.10/s (no audio) for 720p/1080p, in 4/6/8s durations. The best value entry point into the Veo line for fast iteration on ad creative.
$0.180/req
WAN 2.2-5B Image-to-Video is the screen-safe cheap-tier winner — open-weight, ~$0.03 per generation, with reliable, clean motion that holds up for product and UI footage. Resolution 580p/720p, aspect ratio auto/16:9/9:16/1:1 (auto resizes and center-crops to the chosen aspect). 24fps, up to 5s.
$0.192/sec
OmniHuman 1.5 is ByteDance's best open audio-driven avatar model and the anchor of FairStack's lipsync/avatar lane. Give it a photo of a person plus an audio track and it synthesizes an expressive talking-head video where the character's emotions and movements correlate with the audio. 720p (to 60s) or 1080p (to 30s), optional prompt + mask + turbo mode. $0.16/s.
$0.202/sec
Kling O1 V2V Edit is Kuaishou's prior-generation video-to-video editing model, applying prompt-driven modifications to existing video while preserving motion and temporal coherence. It handles environment and object replacement, element insertion, and scene adjustment. With per-second billing, the O1 tier offers a cost-effective editing path for workflows that do not require the latest O3 refinements. The model maintains frame-to-frame consistency while applying visual modifications to the source clip. Compared to the O3 editors, O1 trades some refinement for a lower per-second price, making it a practical option for general video editing tasks. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.288/sec
HappyHorse 1.0 T2V is a text to video model on FairStack. Pay $0.288/sec with a transparent 20% margin — no subscription.
$0.288/sec
HappyHorse 1.0 I2V is a image to video model on FairStack. Pay $0.288/sec with a transparent 20% margin — no subscription.
$0.288/sec
HappyHorse 1.0 Reference-to-Video is a image to video model on FairStack. Pay $0.288/sec with a transparent 20% margin — no subscription.
$0.312/req
Seedance v1.5 Pro I2V is ByteDance's image-to-video model that animates still images into 5-second video clips at 720p resolution. Upload a reference image and the model generates natural motion while preserving the visual fidelity of the source material. Hosted on RunPod for reliable serverless infrastructure. The model excels at bringing static images to life with coherent motion that respects the original composition. Character animations, product reveals, and scene transitions maintain the color palette, lighting, and style of the input image. At 720p output, the results are suitable for social media and web content. Compared to Kling v2.1 Pro I2V which outputs at 1080p with stronger cinematic motion, Seedance v1.5 Pro I2V offers a more affordable entry point at $0.26 per generation. Against Runway Gen-4's $0.50 per generation, it provides reasonable quality at roughly half the cost. The tradeoff is shorter clip duration and slightly less sophisticated motion physics. Ideal for animating product images, creating social media video from existing photography, and character animation from reference art. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.336/sec
Happy Horse Video Edit is Alibaba's video-to-video editing model, applying prompt-driven edits to existing video. It serves as an alternative-provider editor in the FairStack v2v-edit lineup, offering a different model lineage from the Kling family for editing tasks such as style change and scene modification. With per-second billing, it provides a cost-effective editing option. The model preserves motion structure while applying the requested visual changes with frame-to-frame consistency. Best suited for general video editing where an Alibaba-lineage editor is preferred or as a fallback to the Kling editors. Available on FairStack at infrastructure cost plus a 20% platform fee.
$0.360/sec
Sora 2 Pro Image-to-Video is the premium Sora tier, adding 1080p (and true-1080p) output and enhanced detail over the base model. Durations 4–20s, $0.30/s at 720p and $0.50/s at 1080p. For when Sora quality needs to scale to full HD.
$0.480/sec
Veo 3.1 Image-to-Video is Google's flagship image animation model, the category leader for fidelity and native synchronized audio. It animates a still image into a richly detailed clip with generated sound, supporting 720p, 1080p, and 4K output and 4/6/8-second durations. Aspect ratio auto-aligns to your input image (16:9 or 9:16). Billed per second of output at the with-audio rate.
$1.53/req
Google Veo 3.1 Quality is the premium tier of Google's video generation model and the only video model that generates both 1080p video and synchronized audio in a single generation pass. No other model combines video and audio generation into one step, eliminating the need for separate audio production and synchronization. Top-tier visual quality (0.92 score) and motion quality (0.90) place it at the pinnacle of AI video generation capability. The 1080p output resolution is the highest available from any video model, delivering detail that is suitable for professional presentations and large-screen viewing. Native audio generation produces synchronized sound effects, ambient audio, and environmental sounds that match the visual content. Compared to Kling 3.0 Pro (Elo #1), Veo 3.1 Quality matches or exceeds visual quality while adding native audio and higher resolution. At approximately $1.25 per generation, it is the most expensive video model — roughly 10x the cost of mid-tier options. Against a workflow combining a separate video model plus audio generation, the one-step approach is simpler but may cost more. Maximum duration is 8 seconds per generation. Best suited for professional video production where both 1080p resolution and synchronized audio are required, premium marketing content, and workflows where the convenience of one-step video-plus-audio generation justifies the premium price. Available on FairStack at infrastructure cost plus a 20% platform fee.
One balance. Pick a model. Pay the model cost plus 20%.
Step 1
Go to FairStack and open the ai video workspace.
Step 2
Choose from the selectable list on this page — capability flags decide who appears.
Step 3
Run the job. Credits never expire if you stop mid-project.
Step 4
The same balance covers image, video, voice, and music.
FairStack lists 113 selectable models across all modalities. Video alone is 49 — text-to-video, image-to-video, and video-to-video — filtered by each model's real capability flags.
Video is charged per second at the model's rate times 1.20. We never print a single flat price for a clip, because duration and resolution change the cost. Example: Sora 2 image-to-video is $0.12/sec.
Not on FairStack. Sora 2 and Veo 3.1 Fast are image-to-video only on this platform. They appear on the image-to-video page and never on text-to-video.
Still have questions? We're here to help.