EXPLORE / MODEL DIRECTORY

Explore models.Know what runs.

Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.

3Kaista-ready endpointsPublic catalog connected
PUBLIC REFERENCE CATALOG

All model endpoints

Page 51 · 28 per page
image-to-image
falREVIEW REQUIRED

Image Preprocessors

fal-ai/image-preprocessors/scribble

Scribble preprocessor.

preprocessutilityeditingcontrolnet
image-to-image
falREVIEW REQUIRED

DocRes

fal-ai/docres

Enhance low-resolution, blur, shadowed documents with the superior quality of docres for sharper, clearer results.

image-enhancement
training
falREVIEW REQUIRED

ERNIE-Image Trainer

fal-ai/ernie-image-trainer

LoRA trainer for ERNIE-Image, Baidu's powerful 8B-parameter text-to-image model.

lorapersonalizationtrainer
text-to-audio
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/text-to-audio/lora

Text to Audio high-quality using LTX-2.3 with Lora

text-to-audio
video-to-video
briaREVIEW REQUIRED

Video

bria/video/erase/keypoints

High-fidelity keypoint-driven video object removal - minimal input, strong temporal consistency. Trained on licensed data for risk-free commercial video editing.

briavideoerasekeypoints
audio-to-video
falREVIEW REQUIRED

LTX-2.3 22B

fal-ai/ltx-2.3-22b/audio-to-video/lora

Generate video with audio from audio, text and images using LTX-2.3 and custom LoRA

vision
falREVIEW REQUIRED

Sa2VA 4B Video

fal-ai/sa2va/4b/video

Sa2VA is an MLLM capable of question answering, visual prompt understanding, and dense object segmentation at both image and video levels

multimodalvision
training
falREVIEW REQUIRED

LTX-2.3 22B Video to Video Trainer

fal-ai/ltx23-v2v-trainer

Train LTX-2.3 22B for video transformation or video-conditioned generation.

ltx2-videofine-tuningvideo-to-video
vision
falREVIEW REQUIRED

Florence 2 Large Region To Category

fal-ai/florence-2-large/region-to-category

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

multimodalvision
json
falREVIEW REQUIRED

Omnilottie

fal-ai/omnilottie/image-to-lottie

Convert your assets into lottie using Omnilottie.

lotties
audio-to-video
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/audio-to-video/lora

Generate high-quality video with audio from audio, text and images using LTX-2.3 and custom LoRA

audio-to-videolora
audio-to-audio
mirelo-aiREVIEW REQUIRED

Mirelo SFX1.6

mirelo-ai/sfx1.6/inpaint-audio

Erase and replace any moment in your audio with AI-driven precision.

audio-to-audiosfx
image-to-video
falREVIEW REQUIRED

Ltx 2.3 Quality

fal-ai/ltx-2.3-quality/image-to-video/lora

Generate high-quality video with audio from images using LTX-2.3 and custom LoRA

image-to-video
vision
OpenAIREVIEW REQUIRED

Isaac 0.1 [OpenAI Compatible Endpoint]

perceptron/isaac-01/openai/v1/chat/completions

OpenAI spec compatible endpoint of Isaac-01 which is a multimodal vision-language model from Perceptron for various vision language tasks.

multimodalvision
json
falREVIEW REQUIRED

Omnilottie

fal-ai/omnilottie/video-to-lottie

Convert your assets into lottie using Omnilottie.

lottie
video-to-video
briaREVIEW REQUIRED

Bria Video Eraser

bria/bria_video_eraser/erase/keypoints

A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency.

briaerase
video-to-video
briaREVIEW REQUIRED

Bria's VRMBG 3.0 Realtime

bria/video/background-removal/realtime

Remove video backgrounds in real time with Bria’s VRMBG 3.0 model. Built for live streaming, real-time video apps, content creation, and low-latency workflows that need fast, accurate background removal.

briavideobackground-removalrealtime
training
falREVIEW REQUIRED

Wan-2.1 LoRA Trainer

fal-ai/wan-trainer/t2v-14b

Train custom LoRAs for Wan-2.1 T2V 14B

loratraining
speech-to-text
falREVIEW REQUIRED

Speech-to-Text

fal-ai/speech-to-text/turbo/stream

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

streaming
text-to-speech
falREVIEW REQUIRED

Maya

fal-ai/maya/stream

Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.

text-to-speechtts
image-to-image
falREVIEW REQUIRED

Florence-2 Large

fal-ai/florence-2-large/region-to-segmentation

Florence-2 is an advanced vision foundation model that uses a prompt-based approach to handle a wide range of vision and vision-language tasks

multimodalvisionsegmentation
training
falREVIEW REQUIRED

LTX 2.3 Trainer (V2) - Video-to-Video IC-LoRA

fal-ai/ltx23-trainer-v2/ic-lora/v2v

Train an IC-LoRA that learns a video-to-video transformation from paired before/after clips, conditioned at inference on a reference (control) video.

text-to-video
falREVIEW REQUIRED

Heygen

fal-ai/heygen/v2/video-agent

Heygen Text to Video Generation Model

text-to-video
text-to-video
falREVIEW REQUIRED

Infinity Star

fal-ai/infinity-star/text-to-video

InfinityStar’s unified 8B spacetime autoregressive engine to turn any text prompt into crisp 720p videos - 10× faster than diffusion models.

text-to-video
audio-to-audio
falREVIEW REQUIRED

Workflow Utilities Impulse Response

fal-ai/workflow-utilities/impulse-response

FFMPEG Utility for Impulse Response

image-to-image
AlibabaREVIEW REQUIRED

Qwen Image Edit 2509 Lora Gallery

fal-ai/qwen-image-edit-2509-lora-gallery/face-to-full-portrait

Generate full portrait from a cropped face photo

stylizedtransform
image-to-image
falREVIEW REQUIRED

DocRes-dewarp

fal-ai/docres/dewarp

Enhance wraped, folded documents with the superior quality of docres for sharper, clearer results.

image-enhancement
image-to-image
falREVIEW REQUIRED

Image Preprocessors

fal-ai/image-preprocessors/pidi

PIDI (Pidinet) preprocessor.

detectionpreprocessutilitycontrolnet