EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 40 · 28 per page
image-to-video
falREVIEW REQUIRED

CogVideoX-5B

fal-ai/cogvideox-5b/image-to-video

Generate videos from images and prompts using CogVideoX-5B

audio-to-audio
falREVIEW REQUIRED

Stable Audio 3 Medium Audio Outpainting

fal-ai/stable-audio-3/medium/audio-outpainting

Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.

musicextensioncontinuation
text-to-image
falREVIEW REQUIRED

DeepSeek Janus-Pro

fal-ai/janus

DeepSeek Janus-Pro is a novel text-to-image model that unifies multimodal understanding and generation through an autoregressive framework

stylized
image-to-image
falREVIEW REQUIRED

Playground v2.5

fal-ai/playground-v25/image-to-image

State-of-the-art open-source model in aesthetic quality

artisticstyle
image-to-image
falREVIEW REQUIRED

Kolors Image to Image

fal-ai/kolors/image-to-image

Photorealistic Image-to-Image

realismeditingdiffusion
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 SRPO [dev]

fal-ai/flux-1/srpo

FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

text-to-image
falREVIEW REQUIRED

Z Image Base Lora

fal-ai/z-image/base/lora

LoRA endpoint for Z-Image, the foundation model of the Z- Image family.

z-imagebaselora
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Krea [dev]

fal-ai/flux-1/krea/image-to-image

FLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

video-to-video
Black Forest LabsREVIEW REQUIRED

Flux 3 Action

fal-ai/flux-3-action/so101

FLUX 3 Action turns what the robot sees into what it does next. Give it the scene camera image, the wrist camera image, the current SO-101 joint state and a plain-language instruction.

roboticarm
video-to-video
falREVIEW REQUIRED

Wan VACE 14B

fal-ai/wan-vace-14b/reframe

VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

reframe
video-to-video
falREVIEW REQUIRED

Bernini-R Edit Video

fal-ai/bernini-r/edit-video

Edit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.

edittransformstylized
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 SRPO [dev]

fal-ai/flux-1/srpo/image-to-image

FLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.

image-to-image
falREVIEW REQUIRED

ControlLight

fal-ai/control-light

ControlLight is a LoRA fine-tune of FLUX.2 [klein] 9B that enhances low-light images while preserving scene structure and fine details, with a single alpha parameter that gives continuous control over enhancement strength from subtle to full brightening.

stylizedtransform
audio-to-audio
falREVIEW REQUIRED

Stable Audio 3 Medium Base Audio to Audio

fal-ai/stable-audio-3/medium/base/audio-to-audio

Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.

musicstyle-transferremix
image-to-video
falREVIEW REQUIRED

Vidu Image to Video

fal-ai/vidu/image-to-video

Vidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image

motionimage to video
training
Black Forest LabsREVIEW REQUIRED

Flux 2 Klein 9B Base Trainer

fal-ai/flux-2-klein-9b-base-trainer/edit

Fine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.

text-to-image
falREVIEW REQUIRED

Ovis Image

fal-ai/ovis-image

Ovis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.

ovis-imageartistic
video-to-video
briaREVIEW REQUIRED

Bria Video Eraser

bria/bria_video_eraser/erase/prompt

A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency

briaerase
audio-to-audio
falREVIEW REQUIRED

Tada

fal-ai/tada/3b/text-to-speech

A unified speech-language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment.

text-to-video
falREVIEW REQUIRED

LongCat Video

fal-ai/longcat-video/text-to-video/720p

Generate long videos in 720p/30fps from text using LongCat Video

image-to-image
briaREVIEW REQUIRED

Embed Product

bria/embed-product

Seamlessly embed products into any scene with pixel-perfect control, automatic perspective, and natural lighting. Trained on licensed data - risk-free for advertising and eCommerce production.

product-shotadvertising
text-to-video
falREVIEW REQUIRED

Kandinsky 6.0 Lite

fal-ai/kandinsky6-lite/text-to-video

Kandinsky 6.0 Lite is a lightweight, high-speed text-to-video model from Kandinsky Lab, built for fast, efficient generation with strong prompt adherence.

text-to-videokandinskylightweight
audio-to-audio
falREVIEW REQUIRED

Stable Audio 3 Medium Audio Inpainting

fal-ai/stable-audio-3/medium/audio-inpainting

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.

musiceditingrestoration
image-to-video
moonvalleyREVIEW REQUIRED

Marey Realism V1.5

moonvalley/marey/i2v

Generate a video starting from an image as the first frame with Marey, a generative video model trained exclusively on fully licensed data.

text-to-image
falREVIEW REQUIRED

Omnigen V2

fal-ai/omnigen-v2

OmniGen is a unified image generation model that can generate a wide range of images from multi-modal prompts. It can be used for various tasks such as Image Editing, Personalized Image Generation, Virtual Try-On, Multi Person Generation and more!

multimodaleditingtry-on
text-to-audio
falREVIEW REQUIRED

CSM-1B

fal-ai/csm-1b

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.

conversationaltext to speech
image-to-image
falREVIEW REQUIRED

Post Processing Blur

fal-ai/post-processing/blur

Apply Gaussian or Kuwahara blur effects with adjustable radius and sigma parameters

stylizedtransform
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit Plus Lora Gallery

fal-ai/qwen-image-edit-plus-lora-gallery/add-background

Add a realistic scene behind the object with white background

stylizedtransform