text-to-imageFooocus Inpainting
fal-ai/fooocus/inpaintDefault parameters with automated optimizations and quality improvements.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
text-to-imagefal-ai/fooocus/inpaintDefault parameters with automated optimizations and quality improvements.
text-to-imagefal-ai/illusion-diffusionCreate illusions conditioned on image.
image-to-videofal-ai/pixverse/v5/effectsGenerate high quality video clips with different effects using PixVerse v5
image-to-video
video-to-videofal-ai/ltx-2.3-quality/outpaintOutpaint high-quality video using LTX-2.3
text-to-videominimax/h3/text-to-video/loraGenerate video with synchronized audio from a text prompt using MiniMax H3; load a trained LoRA at adjustable strength to lock in style, character, or motion.
image-to-imagefal-ai/flux-lora-depthGenerate high-quality images from depth maps using Flux.1 [dev] depth estimation model. The model produces accurate depth representations for scene understanding and 3D visualization.
video-to-videofal-ai/wan-22-vace-fun-a14b/depthVACE Fun for Wan 2.2 A14B from Alibaba-PAI
video-to-videofal-ai/kling-video/o1/standard/video-to-video/referenceKling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
image-to-imagefal-ai/ideogram/upscaleIdeogram Upscale enhances the resolution of the reference image by up to 2X and might enhance the reference image too. Optionally refine outputs with a prompt for guided improvements.
image-to-3dfal-ai/hunyuan_world/image-to-worldHunyuan World 1.0 turns a single image into a panorama or a 3D world. It creates realistic scenes from the image, allowing you to explore and view it from different angles.
image-to-imagefal-ai/moondream3-preview/segmentMoondream 3 is a vision language model that brings frontier-level visual reasoning with native object detection, pointing, and OCR capabilities to real-world applications requiring fast, inexpensive inference at scale.
text-to-imagefal-ai/hidream-o1-image/devUnified image generation with HiDream-O1-Image. Create, edit, and personalize high-resolution images up to 2K—single native model handles text-to-image, editing, and custom subjects without external components.
llmfal-ai/video-prompt-generatorGenerate video prompts using a variety of techniques including camera direction, style, pacing, special effects and more.
visionfal-ai/florence-2-large/captionFlorence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks
image-to-image
text-to-imagerecraft/v4/style/pro/text-to-vectorGenerates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.
text-to-videominimax/h3-max/styles/vhsGenerates 768p VHS-style video with audio from text prompts or an optional first-frame image. Supports 5–15 second clips and adjustable tape damage, from subtle analog noise to strong tracking distortion.
video-to-videobria/bria_video_eraser/erase/maskA high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency.
image-to-imagebria/genfill/v2The GenFill Route enables the generation of objects by prompt in a specific region of an image. You can define the area for object generation by using a mask that outlines the region where the object will be created. Our model is optimized to work seamlessly with blob-shaped masks.
trainingfal-ai/flux-2-klein-9b-base-trainerFine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.
image-to-imagefal-ai/vidu/q2/reference-to-imageVidu Reference-to-Image creates images by using a reference images and combining them with a prompt.
vision
image-to-imagefal-ai/fast-lightning-sdxl/image-to-imageRun SDXL at the speed of light
text-to-videominimax/h3-max/styles/16bit-pixelGenerates 768p video with audio in a 16-bit pixel-art style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
video-to-videofal-ai/wan-vace-14b/poseVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
text-to-imagefal-ai/recraft/v4.1/utility/pro/text-to-imageRecraft V4.1 Utility Pro pairs the high-resolution output of V4.1 Pro with a faster, cost-efficient runtime. Designed for studios shipping large-format work at scale, it makes premium-quality raster generation viable across full creative pipelines.
image-to-imageideogram/v4/tilingIdeogram V4.0q Tiling generates seamless, edge-matching textures and patterns that repeat infinitely in any direction, ideal for backgrounds, surfaces, and wallpapers.