text-to-imageSDXL ControlNet Union
fal-ai/sdxl-controlnet-unionAn efficent SDXL multi-controlnet text-to-image model.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/sdxl-controlnet-unionAn efficent SDXL multi-controlnet text-to-image model.
image-to-videofal-ai/framepackFramepack is an efficient Image-to-video model that autoregressively generates videos.
image-to-imagefal-ai/image-apps-v2/product-holdingPlace products naturally in a person’s hands for realistic marketing visuals.
text-to-videofal-ai/kandinsky5-pro/text-to-videoKandinsky 5.0 Pro is a diffusion model for fast, high-quality text-to-video generation.
video-to-videofal-ai/flux-3-action/so101FLUX 3 Action turns what the robot sees into what it does next. Give it the scene camera image, the wrist camera image, the current SO-101 joint state and a plain-language instruction.
trainingfal-ai/flux-kontext-trainerLoRA trainer for FLUX.1 Kontext [dev]
text-to-imagefal-ai/sana/v1.5/1.6bSana v1.5 1.6B is a lightweight text-to-image model that delivers 4K image generation with impressive efficiency.
image-to-imagefal-ai/qwen-image-layered/loraQwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers. Use loras to get your custom outputs.
video-to-videotopaz/denoise/videoProfessional video denoising powered by Topaz Labs. Nyx models remove noise at source resolution, with Nyx Fast as a lighter, cheaper pass. Best for low-light and high-ISO footage.
image-to-imagebria/fibo-edit/colorizeImage colorization and color-grading model. Bring color to black-and-white photos or apply curated color treatments using simple style-based commands.
text-to-audiofal-ai/zonosClone voice of any person and speak anything in their voice using zonos' voice cloning.
audio-to-audiofal-ai/qwen-3-tts/clone-voice/0.6bClone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create speeches of yours!
text-to-imagefal-ai/janusDeepSeek Janus-Pro is a novel text-to-image model that unifies multimodal understanding and generation through an autoregressive framework
text-to-imagefal-ai/fooocusDefault parameters with automated optimizations and quality improvements.
image-to-imagefal-ai/bagel/editBagel is a 7B parameter multimodal model from Bytedance-Seed that can generate both images and text.
image-to-imagefal-ai/dreamomni2/editDreamOmni2 is a unified multimodal model for text and image guided image editing.
video-to-videobria/bria_video_eraser/erase/promptA high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency
image-to-videofal-ai/stable-videoGenerate short video clips from your images using SVD v1.1
fal-ai/flux/schnell/reduxFLUX.1 [schnell] Redux is a high-performance endpoint for the FLUX.1 [schnell] model that enables rapid transformation of existing images, delivering high-quality style transfers and image modifications with the core FLUX capabilities.
visionfal-ai/marlinMarlin is a 2B video VLM tuned for the two questions developers actually want to ask of their videos: what is happening, and when?
text-to-audiofal-ai/stable-audio-3/small/sfx/base/text-to-audioStable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
text-to-audiofal-ai/kokoro/italianA high-quality Italian text-to-speech model delivering smooth and expressive speech synthesis.
image-to-videofal-ai/vidu/q2/image-to-video/proUse the latest Vidu Q2 models which much more better quality and control on your videos.
video-to-videofal-ai/ltx-2.3-quality/extend-videoExtend high-quality video with audio from input video using LTX-2.3
video-to-videofal-ai/ltx-video-13b-distilled/multiconditioningGenerate videos from prompts, images, and videos using LTX Video-0.9.7 13B Distilled and custom LoRA
audio-to-audiofal-ai/ace-step/audio-outpaintExtend the beginning or end of provided audio with lyrics and/or style using ACE-Step
image-to-imagerundiffusion-fal/juggernaut-flux/pro/image-to-imageJuggernaut Pro Flux by RunDiffusion is the flagship Juggernaut model rivaling some of the most advanced image models available, often surpassing them in realism. It combines Juggernaut Base with RunDiffusion Photo and features enhancements like reduced background blurriness.
video-to-videofal-ai/pixverse/extendPixVerse Extend model is a video extending tool for your videos using with high-quality video extending techniques