image-to-imageSeedream
bytedance/seedream/v5/lite/editImage editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagebytedance/seedream/v5/lite/editImage editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.
video-to-videofal-ai/kling-video/o1/video-to-video/editEdit an existing video using natural-language instructions, transforming subjects, settings, and style while retaining the original motion structure.
image-to-imagefal-ai/sam-3-1/imageSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
text-to-imagefal-ai/recraft/v4.1/text-to-imageRecraft V4.1 builds on the design-first foundation of V4 with sharper prompt control and cleaner composition. Tuned for brand systems and editorial work, it delivers production-ready raster images that hold up next to a designer's hand.
image-to-imagefal-ai/flux-pro/kontext/max/multiExperimental version of FLUX.1 Kontext [max] with multi image handling capabilities
image-to-imagefal-ai/fashn/tryon/v1.6FASHN v1.6 delivers precise virtual try-on capabilities, accurately rendering garment details like text and patterns at 864x1296 resolution from both on-model and flat-lay photo references.
text-to-imagefal-ai/gemini-25-flash-imageGoogle's famous original image generation and editing model, a.k.a Nano Banana
video-to-videotopaz/upscale/video/generativeProfessional generative video upscaling powered by Topaz Labs. Starlight models rebuild detail that is not in the source, with Fast variants at half the price. Best for low-quality, compressed or archive footage.
video-to-videofal-ai/workflow-utilities/trim-videoFFMPEG Utility for Trim Video
image-to-imagefal-ai/flux-2-pro/outpaintOutpainting generation with FLUX.2 [pro] from Black Forest Labs. Optimized for maximum quality, exceptional photorealism and artistic images.
text-to-speechfal-ai/qwen-3-tts/text-to-speech/1.7bBring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
image-to-video
image-to-videofal-ai/wan-25-preview/image-to-videoWan 2.5 image-to-video model.
image-to-imagefal-ai/flux-pulidAn endpoint for personalized image generation using Flux as per given description.
video-to-textopenrouter/router/videoRun any video-capable LLM with fal. Analyze, summarize, and understand video files using Gemini (Google) models. Supports mp4, mpeg, mov, webm, and YouTube links. Powered by OpenRouter.
image-to-imagefal-ai/evf-samEVF-SAM2 combines natural language understanding with advanced segmentation capabilities, allowing you to precisely mask image regions using intuitive positive and negative text prompts.
image-to-videofal-ai/wan/v2.2-a14b/image-to-videofal-ai/wan/v2.2-A14B/image-to-video
text-to-videoblackforestlabs/flux-3/text-to-videoFLUX 3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.
image-to-videofal-ai/veo3.1/lite/first-last-frame-to-videoVeo 3.1 Lite balances practical utility with professional capabilities, supporting Text-to-Video and Image-to-Video
text-to-audio
image-to-videofal-ai/kling-video/v3/4k/image-to-videoKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
text-to-videofal-ai/kling-video/v2.6/pro/text-to-videoKling 2.6 Pro: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation.
image-to-imagefal-ai/flux-2/turbo/editImage-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control—all at turbo speed.
text-to-imagefal-ai/recraft/v4/text-to-imageRecraft V4 was developed with designers to bring true visual taste to AI image generation. Built for brand systems and production-ready workflows, it goes beyond prompt accuracy delivering stronger composition, refined lighting, realistic materials, and a cohesive aesthetic. The result is imagery shaped by professional design judgment, ready for immediate real-world use without additional post-processing.
video-to-videoveed/lipsyncGenerate realistic lipsync from any audio using VEED's model.
image-to-3dtripo3d/h3.1/multiview-to-3dGenerate 3D models from multiple view images using Tripo H3.1.
video-to-videofal-ai/kling-video/v2.6/pro/motion-controlTransfer movements from a reference video to any character image. Pro mode delivers higher quality output, ideal for complex dance moves and gestures.
text-to-audiofal-ai/stable-audio-3/medium/text-to-audioStable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.