image-to-videoFLUX 3 Image to Video
blackforestlabs/flux-3/image-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videoblackforestlabs/flux-3/image-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.
video-to-videofal-ai/latentsyncLatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.
image-to-imagealibaba/qwen-image-3/editEdits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes
image-to-videofal-ai/bytedance/seedance/v1/pro/image-to-videoSeedance 1.0 Pro, a high quality video generation model developed by Bytedance.
video-to-videofal-ai/ffmpeg-api/composeCompose videos from multiple media sources using FFmpeg API.
text-to-imagefal-ai/flux-2/turboText-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities—all at turbo speed.
text-to-videofal-ai/kling-video/v3/standard/text-to-videoKling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.
image-to-imagefal-ai/flux-2-max/editFLUX.2 [max] delivers state-of-the-art image generation and advanced image editing with exceptional realism, precision, and consistency.
image-to-videoxai/grok-imagine-video/v1.5/lite/image-to-videoGenerate videos from images using xAI's Grok Imagine Video 1.5 Lite model.
image-to-imagefal-ai/qwen-image-edit-2511-multiple-anglesGenerates same scene from different angles (azimuth/elevation) with Qwen image Edit 2511 and the Lora Multiple Angles
image-to-3dtripo3d/h3.1/image-to-3dGenerate high-quality 3D models from a single image using Tripo H3.1.
image-to-imagetopaz/upscale/image/generativeProfessional generative image upscaling powered by Topaz Labs. Wonder 3.5 leads the range, with Redefine for prompt-guided detail and Recovery for extreme low-resolution sources. Best for rebuilding sharp detail in small or blurry images.
video-to-videofal-ai/kling-video/v2.6/standard/motion-controlTransfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
text-to-videominimax/h3/text-to-videoMiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.
image-to-videoveed/fabric-1.0VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
text-to-videofal-ai/veo3.1Veo 3.1 by Google, the most advanced AI video generation model in the world. With sound on!
text-to-videoalibaba/wan-3.0/text-to-videoWan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
text-to-videofal-ai/kling-video/lipsync/audio-to-videoKling LipSync is an audio-to-video model that generates realistic lip movements from audio input.
text-to-audiofal-ai/lyria2Lyria 2 is Google's latest music generation model, you can generate any type of music with this model.
video-to-videofal-ai/kling-video/v3/standard/motion-controlTransfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
text-to-speechfal-ai/minimax/voice-cloneClone a voice from a sample audio and generate speech from text prompts using the MiniMax model, which leverages advanced AI techniques to create high-quality text-to-speech.
image-to-imagefal-ai/recraft/vectorizeConverts a given raster image to SVG format using Recraft model.
image-to-videoalibaba/wan-3.0-prime/image-to-videoWan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly detailed animation.
text-to-audiocassetteai/music-generatorCassetteAI’s model generates a 30-second sample in under 2 seconds and a full 3-minute track in under 10 seconds. At 44.1 kHz stereo audio, expect a level of professional consistency with no breaks, no squeaks, and no random interruptions in your creations.
image-to-videofal-ai/bytedance/seedance/v1/pro/fast/image-to-videoImage to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost
text-to-imagefal-ai/gpt-image-1.5GPT Image 1.5 generates high-fidelity images with strong prompt adherence, preserving composition, lighting, and fine-grained detail.
text-to-imagexai/grok-imagine-image/v2.0/text-to-imageGenerate images from text using xAi's Grok Imagine 2.0 model.
image-to-videofal-ai/kling-video/v3/turbo/standard/image-to-videoKling 3.0 Turbo Standard animates a first and last frame reference image into 720P video with native audio, delivering quick, affordable image-driven motion for fast turnaround