image-to-videoKling Video v3 Image to Video [Pro]
fal-ai/kling-video/v3/pro/image-to-videoKling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videofal-ai/kling-video/v3/pro/image-to-videoKling 3.0 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
image-to-videominimax/h3-max/image-to-videofal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
image-to-videominimax/h3-max/reference-to-videofal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
image-to-videominimax/h3-max-turbo/image-to-videofal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
image-to-videobytedance/seedance-2.5/reference-to-videoDreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.
image-to-videofal-ai/kling-video/v2.5-turbo/pro/image-to-videoKling 2.5 Turbo Pro: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.
image-to-videobytedance/seedance-2.5/image-to-videoDreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.
image-to-videofal-ai/kling-video/v3/standard/image-to-videoKling 3.0 Standard: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation, with custom element support.
image-to-videobytedance/seedance-2.0/image-to-videoByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.
image-to-videofal-ai/kling-video/v2.6/pro/image-to-videoKling 2.6 Pro: Top-tier image-to-video with cinematic visuals, fluid motion, and native audio generation.
image-to-videofal-ai/veo3.1/fast/image-to-videoGenerate videos from your image prompts using Veo 3.1 fast.
image-to-videominimax/h3/reference-to-videoMiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
image-to-videominimax/h3/image-to-videoMiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.
image-to-videobytedance/seedance-2.0/reference-to-videoByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.
image-to-videofal-ai/veo3.1/image-to-videoVeo 3.1 is the latest state-of-the art video generation model from Google DeepMind
image-to-videofal-ai/veo3.1/lite/image-to-videoVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
image-to-videoalibaba/wan-3.0/reference-to-videoWan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
image-to-videofal-ai/bytedance/seedance/v1.5/pro/image-to-videoGenerate videos with audio with Seedance 1.5 (supports start & end frame)
image-to-videogoogle/gemini-omni-flash/v1.1/image-to-videoGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint animates a still image into video with synchronized audio, extending a single frame into coherent motion that reflects the logic of the real world.
image-to-videoalibaba/wan-3.0/image-to-videoWan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
image-to-videofal-ai/bytedance/omnihuman/v1.5Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
image-to-videominimax/h3-max/lip-sync/image-to-videoH3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.
image-to-videofal-ai/kling-video/o3/pro/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
image-to-videoxai/grok-imagine-video/image-to-videoGenerate videos from images with audio using xAI's Grok Imagine Video model.
image-to-videofal-ai/kling-video/ai-avatar/v2/standardKling AI Avatar v2 Standard: Endpoint for creating avatar videos with realistic humans, animals, cartoons, or stylized characters
image-to-videobytedance/seedance-2.5/us/reference-to-videoUS-hosted ByteDance Seedance 2.5 generates video with native audio from up to 30 images, 10 videos, and 10 audio references. Supports reference-guided generation, video editing, and extension at 480p, 720p or 1080p.
image-to-videoxai/grok-imagine-video/v1.5/image-to-videoGenerate videos from images with audio using xAI's Grok Imagine 1.5 Video model.
image-to-videofal-ai/kling-video/o3/pro/reference-to-videoTransform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.