EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 10 · 28 per page
image-to-video
MiniMaxREVIEW REQUIRED

H3 Max Camera Controls

minimax/h3-max/camera-controls

H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space

stylizedtransformediting
text-to-audio
falREVIEW REQUIRED

ACE Step

fal-ai/ace-step

Generate music with lyrics from text using ACE-Step

text-to-audiotext-to-music
audio-to-audio
falREVIEW REQUIRED

Audio Understanding

fal-ai/audio-understanding

A audio understanding model to analyze audio content and answer questions about what's happening in the audio based on user prompts.

utilityaudio
image-to-video
GoogleREVIEW REQUIRED

Veo 3.1

fal-ai/veo3.1/reference-to-video

Generate Videos from images using Google's Veo 3.1

text-to-image
Black Forest LabsREVIEW REQUIRED

FLUX.1 Kontext [pro]

fal-ai/flux-pro/kontext/text-to-image

The FLUX.1 Kontext [pro] text-to-image delivers state-of-the-art image generation results with unprecedented prompt following, photorealistic rendering, and flawless typography.

text-to-image
RecraftREVIEW REQUIRED

Recraft V4.1 Text to Vector

fal-ai/recraft/v4.1/text-to-vector

Recraft V4.1 Vector turns prompts into fully editable SVGs with structured layers and clean geometry. Built for logos, icons, and illustration systems, it produces artwork that goes straight from generation into Figma or Illustrator.

stylizedtransformtypography
video-to-video
Black Forest LabsREVIEW REQUIRED

Flux 3 FAST Edit Video

blackforestlabs/flux-3/edit-video

FLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.

stylizedtransformlipsync
audio-to-audio
ElevenLabsREVIEW REQUIRED

ElevenLabs Audio Isolation

fal-ai/elevenlabs/audio-isolation

Isolate audio tracks using ElevenLabs advanced audio isolation technology.

audio
text-to-image
falREVIEW REQUIRED

Krea 2 Text to Image Turbo LoRA

fal-ai/krea-2/turbo/lora

Generate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.

stylizedtransformtypographyrealism
text-to-audio
soniloREVIEW REQUIRED

Sonilo V1.1 Text to Music

sonilo/v1.1/text-to-music

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

stylizedtransformlipsync
text-to-audio
falREVIEW REQUIRED

Kokoro TTS

fal-ai/kokoro/american-english

Kokoro is a lightweight text-to-speech model that delivers comparable quality to larger models while being significantly faster and more cost-efficient.

speech
image-to-3d
tripo3dREVIEW REQUIRED

Tripo P2 Image to 3D

tripo3d/p2/image-to-3d

Tripo P2 generates 3D models from a single image, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.

3dimage-to-3d3d-generationlow-poly
video-to-video
KlingREVIEW REQUIRED

Kling O3 Reference Video to Video [Pro]

fal-ai/kling-video/o3/pro/video-to-video/reference

Kling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.

video-to-video
text-to-video
ByteDanceREVIEW REQUIRED

Bytedance Seedance V1.5 Pro Text To Video

fal-ai/bytedance/seedance/v1.5/pro/text-to-video

Generate videos with audio with Seedance 1.5

bytedanceseedanceaudio
image-to-video
GoogleREVIEW REQUIRED

Gemini Omni Flash

google/gemini-omni-flash/reference-to-video

Generates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.

stylizedtransformlipsync
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit

fal-ai/qwen-image-edit

Endpoint for Qwen's Image Editing model. Has superior text editing capabilities.

image-editingimage-to-imagehigh-quality-text
text-to-speech
MiniMaxREVIEW REQUIRED

MiniMax Speech 2.8 [Turbo]

fal-ai/minimax/speech-2.8-turbo

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

image-to-video
xAIREVIEW REQUIRED

Grok Imagine Video 1.5 Reference to Video

xai/grok-imagine-video/v1.5/reference-to-video

Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.

stylizedtransformlipsync
text-to-image
ByteDanceREVIEW REQUIRED

Seedream

bytedance/seedream/v5/lite/text-to-image

Text to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.

text-to-imagebytedanceseedream-5.0-lite
image-to-video
falREVIEW REQUIRED

sync-3 Avatar Image to Video

fal-ai/sync-lipsync/v3/image-to-video

sync-3 image to video turns a single still into a talking character, and works with any illustration or animated frame paired with a voice track

animationlip synctext-to-speech
text-to-image
GoogleREVIEW REQUIRED

Nano Banana Lite

google/nano-banana-lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

image-to-3d
falREVIEW REQUIRED

Hunyuan3d V3

fal-ai/hunyuan3d-v3/image-to-3d

Transform your photos into ultra-high-resolution 3D models in seconds. Film-quality geometry with PBR textures, ready for games, e-commerce, and 3D printing.

video-to-video
falREVIEW REQUIRED

Depth Anything Video

fal-ai/depth-anything-video

Generates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.

video to videomotionedit
text-to-speech
falREVIEW REQUIRED

Inworld TTS-1.5 Max

fal-ai/inworld-tts

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

text-to-speechinworldtts
video-to-video
veedREVIEW REQUIRED

VEED Lipsync

veed/lipsync/v2

Generate production-quality lipsync from any audio using VEED's most advanced model yet.

veedlipsyncvideo-to-videoavatar
image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit Plus

fal-ai/qwen-image-edit-plus

Endpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.

image-editingimage-to-imagehigh-quality-text
video-to-video
decartREVIEW REQUIRED

Lucy 2.5

decart/lucy-2-5/realtime

Real-time, prompt-driven video editing over WebRTC. Restyle, swap backgrounds, and add or replace objects live on a webcam or streamed feed at interactive latency.

realtimevideo-to-videowebrtc
image-to-video
KlingREVIEW REQUIRED

Kling O1 First Frame Last Frame to Video [Pro]

fal-ai/kling-video/o1/image-to-video

Generate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.