audio-to-audioACE Step Audio To Audio
fal-ai/ace-step/audio-to-audioGenerate music from a lyrics and example audio using ACE-Step
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
audio-to-audiofal-ai/ace-step/audio-to-audioGenerate music from a lyrics and example audio using ACE-Step
image-to-imagefal-ai/marigold-v2Estimate depth from a single image with Marigold V2, a diffusion-based depth model built on Qwen-Image-Edit, returning a colorized depth map.
text-to-textopenrouter/router/decisionsRun any decision model with fal, powered by OpenRouter.
video-to-videofal-ai/kling-video/o3/4k/video-to-video/referenceKling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
image-to-imagefal-ai/feynobgFeyNobg is a state of the art AI model for background removal from feyninc
image-to-imagebria/fibo-edit-1.5/editCommercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.
image-to-videofal-ai/ltx-video/image-to-videoGenerate videos from images using LTX Video
image-to-imagefal-ai/flux-2/klein/4b/base/editImage-to-image editing with FLUX.2 [klein] 4B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
text-to-videoalibaba/happy-horse/text-to-videoGenerate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.
image-to-3dtripo3d/triposplatTripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach
text-to-videofal-ai/wan/v2.2-a14b/text-to-video/turboWan-2.2 turbo text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
video-to-videofal-ai/sam-3/video-rleSAM 3 is a unified foundation model for promptable segmentation in images and videos. It can detect, segment, and track objects using text or visual prompts such as points, boxes, and masks.
text-to-imagefal-ai/qwen-image-2512/loraLoRA inference endpoint for Qwen Image 2512, an improved version of Qwen Image with better text rendering, finer natural textures, and more realistic human generation.
visionfal-ai/florence-2-large/more-detailed-captionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
image-to-imagefal-ai/ddcolorBring colors into old or new black and white photos with DDColor.
image-to-videofal-ai/pika/v2.2/image-to-videoTurn photos into mind-blowing, dynamic videos in up to 1080p. Experience better image clarity and crisper, sharper visuals.
video-to-audiofal-ai/kling-video/video-to-audioGenerate audio from input videos using Kling
llmopenrouter/router/openai/v1/embeddingsGenerate text embeddings using OpenAI-compatible API. Access embedding models like text-embedding-3-small, text-embedding-3-large (OpenAI), and other embedding models available through OpenRouter. Drop-in replacement for the OpenAI embeddings API. Powered by OpenRouter.
image-to-videoalibaba/happy-horse/v1.1/reference-to-videoHappy Horse 1.1 is Alibaba's #1-ranked video model. This reference-to-video endpoint turns up to 9 reference images into 1080p video with synchronized native audio and multilingual lip-sync for consistent characters.
visionfal-ai/moondream2/visual-queryMoondream2 is a highly efficient open-source vision language model that combines powerful image understanding capabilities with a remarkably small footprint.
text-to-videofal-ai/minimax/hailuo-2.3/standard/text-to-videoMiniMax Hailuo-2.3 Text To Video API (Standard, 768p): Advanced text-to-video generation model with 768p resolution
image-to-videofal-ai/wan/v2.7/reference-to-videoWan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
image-to-video
image-to-imagefal-ai/smart-resizeSmart image resize to arbitrary dimensions, powered by Nano Banana Pro with vision-LLM-guided prompting for composition-aware recomposition. Crop, cropping, resize ads.
image-to-imagefal-ai/florence-2-large/caption-to-phrase-groundingFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
audio-to-videofal-ai/elevenlabs/dubbingGenerate dubbed videos or audios using ElevenLabs Dubbing feature!
video-to-videotopaz/upscale/video/creativeProfessional creative video upscaling powered by Topaz Labs. Astra 2 reimagines fine detail and typically delivers 4K output. Best for cinematic shots that need maximum visual impact.
video-to-videomirelo-ai/sfx1.6/video-to-videoGenerate synced sounds for any video, and return it with its new sound track (like MMAudio). Now up to 60 seconds!