EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 17 · 28 per page
text-to-audio
falREVIEW REQUIRED

MMAudio V2 Text to Audio

fal-ai/mmaudio-v2/text-to-audio

MMAudio generates synchronized audio given text inputs. It can generate sounds described by a prompt.

audiofast
video-to-video
falREVIEW REQUIRED

Heygen Lipsync - Precision

fal-ai/heygen/v3/lipsync/precision

Replace or dub audio on an existing video with high-accuracy avatar-inference lip-sync.

lipsyncstylizedtransform
text-to-image
AlibabaREVIEW REQUIRED

Wan

fal-ai/wan/v2.7/text-to-image

Generate high-quality images from text prompts using the WAN 2.7 model with advanced prompt understanding and detailed output.

wantext-to-imageimage-generation
video-to-video
falREVIEW REQUIRED

Infinitalk

fal-ai/infinitalk

Infinitalk model generates a talking avatar video from an image and audio file. The avatar lip-syncs to the provided audio with natural facial expressions.

stylizedtransform
image-to-image
briaREVIEW REQUIRED

Bria Increase Resolution: Upscale Images up to 4x Without Losing Detail | fal

bria/increase-resolution

Upscale any image 2x or 4x, up to 8192×8192, with Bria Increase Resolution. Preserves the original content — no regeneration, no altered details. Commercial-safe

utilityediting
image-to-image
falREVIEW REQUIRED

Florence-2 Large

fal-ai/florence-2-large/ocr-with-region

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

ocrmultimodalvision
video-to-audio
soniloREVIEW REQUIRED

V1.1

sonilo/v1.1/video-to-music

Analyzes your video’s pacing, mood, and timing to generate a frame-synced, licensed, commercial-use-safe soundtrack in seconds.

stylizedtransformlipsync
llm
OpenAIREVIEW REQUIRED

OpenRouter Responses [OpenAI Compatible]

openrouter/router/openai/v1/responses

The OpenRouter Responses API with fal, powered by OpenRouter, provides unified access to a wide range of large language models - including GPT, Claude, Gemini, and many others through a single API interface.

image-to-image
falREVIEW REQUIRED

Bria Product Shot

fal-ai/bria/product-shot

Place any product in any scenery with just a prompt or reference image while maintaining high integrity of the product. Trained exclusively on licensed data for safe and risk-free commercial use and optimized for eCommerce.

product photography
image-to-3d
falREVIEW REQUIRED

Trellis

fal-ai/trellis/multi

Generate 3D models from multiple images using Trellis. A native 3D generative model enabling versatile and high-quality 3D asset creation.

stylized
image-to-image
falREVIEW REQUIRED

Image Editing Photo Restoration

fal-ai/image-editing/photo-restoration

Restore and enhance old or damaged photos by removing imperfections, adding color while preserving the original character and details of the image.

stylizedtransform
audio-to-audio
falREVIEW REQUIRED

FFmpeg API [Merge Audios]

fal-ai/ffmpeg-api/merge-audios

Merge audios into a single audio using FFmpeg API!

ffmpeg
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 9B LoRA

fal-ai/flux-2/klein/9b/edit/lora

Image-to-image editing with FLUX.2 [klein] 9B from Black Forest Labs and custom LoRA. Precise modifications using natural language descriptions and hex color control.

image-to-3d
falREVIEW REQUIRED

Hyper3D - Rodin V2.5 - Image to 3D - Fast

fal-ai/hyper3d/rodin/v2.5/fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

image-to-3d
audio-to-audio
falREVIEW REQUIRED

DeepFilterNet 3

fal-ai/deepfilternet3

Enhance speech audio by removing background noise and upsampling to 48KHz

speech-enhancement
image-to-image
falREVIEW REQUIRED

Segment Anything Model 2

fal-ai/sam2/auto-segment

SAM 2 is a model for segmenting images automatically. It can return individual masks or a single mask for the entire image.

segmentationmask
image-to-video
KlingREVIEW REQUIRED

Kling O1 Reference Image to Video [Pro]

fal-ai/kling-video/o1/reference-to-video

Transform images, elements, and text into consistent, high-quality video scenes, ensuring stable character identity, object details, and environments.

image-to-image
falREVIEW REQUIRED

Midas Depth Estimation

fal-ai/imageutils/depth

Create depth maps using Midas depth estimation.

depthutility
image-to-3d
falREVIEW REQUIRED

Hunyuan 3D Rapid Image to 3D

fal-ai/hunyuan-3d/v3.1/rapid/image-to-3d

Rapidly generate 3D models from images using Hunyuan 3D.

3dhunyuanimage-to-3d
text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Kontext [max]

fal-ai/flux-pro/kontext/max/text-to-image

FLUX.1 Kontext [max] text-to-image is a new premium model brings maximum performance across all aspects – greatly improved prompt adherence.

image-to-video
KlingREVIEW REQUIRED

Kling O1 First Frame Last Frame to Video [Standard]

fal-ai/kling-video/o1/standard/image-to-video

Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.

image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 [dev] with Controlnets and Loras

fal-ai/flux-general/inpainting

FLUX General Inpainting is a versatile endpoint that enables precise image editing and completion, supporting multiple AI extensions including LoRA, ControlNet, and IP-Adapter for enhanced control over inpainting results and sophisticated image modifications.

loracontrolnetip-adapter
image-to-video
ByteDanceREVIEW REQUIRED

OmniHuman

fal-ai/bytedance/omnihuman

OmniHuman generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.

image-to-videolipsync
image-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.2 [klein] 9B Base

fal-ai/flux-2/klein/9b/base/edit

Image-to-image editing with Flux 2 [klein] 9B Base from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.

image-to-image
briaREVIEW REQUIRED

Bria Virtual Try-On (FIBO-Edit-1.5)

bria/fibo-edit-1.5/virtual-try-on

Bria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.

virtual try-onfashion techapparele-commerce
text-to-image
IdeogramREVIEW REQUIRED

Ideogram V2

fal-ai/ideogram/v2

Generate high-quality images, posters, and logos with Ideogram V2. Features exceptional typography handling and realistic outputs optimized for commercial and creative use.

realismtypography
text-to-video
ByteDanceREVIEW REQUIRED

Seedance 1.0 Pro

fal-ai/bytedance/seedance/v1/pro/text-to-video

Seedance 1.0 Pro, a high quality video generation model developed by Bytedance.