ElevenLabs Speech to Text
fal-ai/elevenlabs/speech-to-textGenerate text from speech using ElevenLabs advanced speech-to-text model.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
fal-ai/elevenlabs/speech-to-textGenerate text from speech using ElevenLabs advanced speech-to-text model.
video-to-videominimax/h3-max/recastRecast the people in a video using reference photos with H3 Max, while preserving the source motion, camera, cuts, and audio.
video-to-videofal-ai/kling-video/v3/pro/motion-controlTransfer movements from a reference video to any character image. Cost-effective mode for motion transfer, perfect for portraits and simple animations.
image-to-videofal-ai/kling-video/o3/standard/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
image-to-imagepixelcut/background-removalPixelcut’s Background Remover enables fast, ultra high-quality removal of backgrounds from images. Perfect for e-commerce and image editing workflows. Powered by advanced AI for clean, perfect cutouts every time.
image-to-videobytedance/seedance-2.0/fast/image-to-videoByteDance's most advanced image-to-video model, fast tier. Lower latency and cost with synchronized audio, start and end frame control, and motion prompts.
image-to-imagefal-ai/flux-2/editImage-to-image editing with FLUX.2 [dev] from Black Forest Labs. Precise modifications using natural language descriptions and hex color control.
text-to-videofal-ai/veo3.1/fastFaster and more cost effective version of Google's Veo 3.1!
text-to-imagefal-ai/flux-2/flashText-to-image generation with FLUX.2 [dev] from Black Forest Labs. Enhanced realism, crisper text generation, and native editing capabilities— in a flash.
image-to-videogoogle/gemini-omni-flash/v1.1/reference-to-videoGemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video from combined multimodal references, images, videos and text together. Reasoning across all inputs to produce a single coherent result, with characters retaining their face, clothing, and voice throughout
image-to-imagefal-ai/flux-pro/kontext/maxFLUX.1 Kontext [max] is a model with greatly improved prompt adherence and typography generation meet premium consistency for editing without compromise on speed.
audio-to-audiofal-ai/demucsSOTA stemming model for voice, drums, bass, guitar and more.
video-to-videotopaz/upscale/video/precisionProfessional video upscaling powered by Topaz Labs. Precision models (Proteus, Artemis, Iris, Dione, Theia, Gaia, Rhea) enhance footage up to 4x while staying faithful to the source. Best for clean, natural upscales of real-world footage.
video-to-videofal-ai/ffmpeg-api/merge-videosUse ffmpeg capabilities to merge 2 or more videos.
image-to-3dfal-ai/trellis-2Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.
text-to-imagefal-ai/recraft/v3/text-to-imageRecraft V3 is a text-to-image model with the ability to generate long texts, vector art, images in brand style, and much more. As of today, it is SOTA in image generation, proven by Hugging Face's industry-leading Text-to-Image Benchmark by Artificial Analysis.
image-to-image
text-to-videobytedance/seedance-2.0/text-to-videoByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.
image-to-videofal-ai/kling-video/v3/turbo/pro/image-to-videoGenerate high quality 1080p videos from images using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.
image-to-imagefal-ai/ffmpeg-api/extract-frameffmpeg endpoint for first, middle and last frame extraction from videos
video-to-videofal-ai/bytedance-upscaler/upscale/videoUpscale videos with Bytedance's video upscaler.
image-to-imagefal-ai/qwen-image-edit-2511Endpoint for Qwen's Image Editing 2511 model.
video-to-videofal-ai/sync-lipsync/v2/proGenerate high-quality realistic lipsync animations from audio while preserving unique details like natural teeth and unique facial features using the state-of-the-art Sync Lipsync 2 Pro model.
image-to-videobytedance/seedance-2.0/fast/reference-to-videoByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.
image-to-imagebytedance/seedream/v5/flash/editSeedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.
image-to-imagefal-ai/ideogram/remove-backgroundRemove backgrounds from existing images with Ideogram's remove background feature. Isolate subjects cleanly for compositing and creative reuse.
image-to-imagefal-ai/imageutils/rembgRemove the background from an image.
image-to-videofal-ai/kling-video/v2.5-turbo/standard/image-to-videoKling 2.5 Turbo Standard: Top-tier image-to-video generation with unparalleled motion fluidity, cinematic visuals, and exceptional prompt precision.