image-to-imageWorkflow Utilities Extract Nth Frame
fal-ai/workflow-utilities/extract-nth-frameFFMPEG Untility for Extracting nth Frame
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-imagefal-ai/workflow-utilities/extract-nth-frameFFMPEG Untility for Extracting nth Frame
text-to-imagefal-ai/z-image/baseZ-Image is the foundation model of the Z- Image family, engineered for good quality, robust generative diversity, broad stylistic coverage, and precise prompt adherence.
video-to-videofal-ai/sam-3-1/video-rleSAM 3.1 builds comes with Object Multiplex, a shared-memory approach for joint multi-object tracking that delivers faster speeds with larger number of objects tracked.
image-to-video
text-to-imageluma/agent/uni-1/v1/text-to-imageLuma Uni-1 turns a text prompt into a single high-fidelity image, with control over aspect ratio and visual style, plus optional web-sourced and reference-image guidance for sharper grounding.
video-to-videofal-ai/kling-video/o1/video-to-video/referenceKling O1 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
image-to-imagetopaz/restore/imageProfessional image restoration powered by Topaz Labs. Recover 3 generatively rebuilds natural detail; Dust-Scratch V2 cleans film dust and scratches. Best for old, damaged or degraded photos.
image-to-textnvidia/nemotron-3-nano-omni/visionVision reasoning variant of NVIDIA's Nemotron 3 Nano Omni. 30B A3B hybrid Transformer-Mamba MoE - accepts an image plus a prompt and returns text.
video-to-audiomirelo-ai/sfx-v1.5/video-to-audioGenerate synced sounds for any video, and return the new sound track (like MMAudio)
image-to-imagegoogle/virtual-try-onGenerate realistic virtual try-on images from a person image and a clothing product image.
image-to-videoblackforestlabs/flux-3/image-to-video/draftFLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.
image-to-3dhitem3d/hi3d/v3.0/multi-view-to-3dGenerate 3D models from multiple view images using Hi3D V3.0.
image-to-imagefal-ai/flux-kontext-lora/inpaintFast inpainting endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image inpainting with reference images, while using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.
image-to-imagefal-ai/imageutils/marigold-depthCreate depth maps using Marigold depth estimation.
text-to-imagefal-ai/stable-diffusion-v35-largeStable Diffusion 3.5 Large is a Multimodal Diffusion Transformer (MMDiT) text-to-image model that features improved performance in image quality, typography, complex prompt understanding, and resource-efficiency.
image-to-3dmeshy/v7/multi-image-to-3deconstructs a high-fidelity textured 3D model from multiple angle views of one object, with game-ready topology and polygon control
text-to-audiomirelo-ai/sfx1.6/text-to-audioGenerate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.
text-to-videofal-ai/ltx-videoGenerate videos from prompts using LTX Video
text-to-imageluma/agent/uni-1/v1/maxLuma Uni-1 Max generates a single image at the model's highest fidelity, delivering richer detail and stronger prompt adherence than the base tier for hero-quality stills.
video-to-videopixelcut/video-background-removalPixelcut's Video Background Remover is an AI segmentation model that erases backgrounds frame by frame, with seamless temporal consistency.
audio-to-videofal-ai/flashtalkAudio-driven talking avatar generation powered by the SoulX-FlashTalk 14B model.
text-to-imagefal-ai/flux-1/devFLUX.1 [dev] is a 12 billion parameter flow transformer that generates high-quality images from text. It is suitable for personal and commercial use.
image-to-3dtripo3d/tripo/v2.5/multiview-to-3dState of the art Multiview to 3D Object generation. Generate 3D models from multiple images!
video-to-videobria/video/erase/maskHigh-fidelity mask-based video object removal with strong temporal consistency. Erase unwanted objects, people, or elements while preserving aesthetic quality. Trained on licensed data for risk-free commercial use.
video-to-video
video-to-videoxai/grok-imagine-video/extend-videoExtend videos with xAI's Grok Imagine video model
text-to-imagefal-ai/loraRun Any Stable Diffusion model with customizable LoRA weights.
video-to-videomirelo-ai/sfx-v1.5/video-to-videoGenerate synced sounds for any video, and return it with its new sound track (like MMAudio)