video-to-videoKling O1 Edit Video [Standard]
fal-ai/kling-video/o1/standard/video-to-video/editEdit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
video-to-videofal-ai/kling-video/o1/standard/video-to-video/editEdit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
image-to-imagefal-ai/bria/background/replaceBria Background Replace allows for efficient swapping of backgrounds in images via text prompts or reference image, delivering realistic and polished results. Trained exclusively on licensed data for safe and risk-free commercial use
text-to-audiofal-ai/stable-audio-3/small/music/text-to-audioStable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.
image-to-imagefal-ai/ideogram/object-removalPrompt-free object removal from an image and mask, erasing objects with their shadows and reflections and reconstructing the scene cleanly.
image-to-imagefal-ai/phota/editPhota's model enables personalized photo editing, preserving identity while erasing distractions seamlessly.
image-to-videofal-ai/minimax/hailuo-2.3-fast/pro/image-to-videoMiniMax Hailuo-2.3-Fast Image To Video API (Pro, 1080p): Advanced fast image-to-video generation model with 1080p resolution
video-to-videoblackforestlabs/flux-3/extend-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.
image-to-imagefal-ai/florence-2-large/open-vocabulary-detectionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
text-to-imagefal-ai/ideogram/v3/generate-transparentGenerate images with transparent backgrounds using Ideogram Transparent model
image-to-imagefal-ai/gpt-image-1-mini/editGPT Image 1 mini combines OpenAI's advanced language capabilities, powered by GPT-5, with GPT Image 1 Mini for efficient image generation.
video-to-videofal-ai/wan-vace-14b/inpaintingVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
video-to-videofal-ai/sam-3-1/videoSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
image-to-imageideogram/v4/image-to-imageIdeogram V4.0q Image-to-Image transforms an input image with a text prompt, restyling and reworking the composition while preserving its core structure for prompt-faithful, high-fidelity edits.
video-to-videofal-ai/veo3.1/fast/extend-videoExtend Veo-Created Videos up to 30 seconds
text-to-speechfal-ai/zonos2Zonos2 is a text-to-speech model that clones a voice from a short sample and speaks naturally across many languages.
text-to-imagekrea/v2/medium/turbo/text-to-imageGenerate high-fidelity images extremely fast from text with Krea 2 Medium Turbo, supporting aspect ratio, creativity, seed controls, and optional style references.
image-to-imagemicrosoft/mai-image-2.5-pro/editApply precise, controllable edits to a reference image while preserving composition, typography, identity, and fine visual detail.
image-to-videofal-ai/pixverse/v5/image-to-videoGenerate high quality video clips from text and image prompts using PixVerse v5
text-to-videofal-ai/ltx-2.3/text-to-videoLTX-2.3 is a high-quality, fast AI video model available in Pro and Fast variants for text-to-video, image-to-video, and audio-to-video.
image-to-imagefal-ai/recraft/v3/image-to-imageRecraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.
text-to-videofal-ai/kling-video/v3/4k/text-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
text-to-video
image-to-imageluma/agent/uni-1/v1/editLuma Uni-1 Edit reworks a source image from a text instruction, preserving the original composition while applying style changes and following optional reference images to steer the result.
image-to-3dhitem3d/hi3d/v3.0/image-to-3dGenerate 3D models from a single image with Hi3D V3.0.
image-to-imagefal-ai/sam-3-1/image-rleSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
text-to-3dfal-ai/hunyuan-3d/v3.1/pro/text-to-3dGenerate 3D models from text prompts with Hunyuan 3D Pro
text-to-videolightricks/ltx-2.5/text-to-video/fastLTX-2.5 is Lightricks' open-source audio-video model. This endpoint generates synchronized video and audio from a text prompt in a single pass, in a speed-optimized mode built for rapid iteration and previews.
text-to-audiofal-ai/minimax-music/v2.5MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.