image-to-videoKling O1 First Frame Last Frame to Video [Pro]
fal-ai/kling-video/o1/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videofal-ai/kling-video/o1/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.
video-to-videofal-ai/depth-anything-videoGenerates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.
image-to-imagefal-ai/qwen-image-edit-plusEndpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.
text-to-imagefal-ai/flux-2-flexText-to-image generation with FLUX.2 [flex] from Black Forest Labs. Features adjustable inference steps and guidance scale for fine-tuned control. Enhanced typography and text rendering capabilities.
text-to-videogoogle/gemini-omni-flashCreates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.
text-to-audiosonilo/v1.1/text-to-sound-effectsGenerates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
video-to-videofal-ai/kling-video/o3/pro/video-to-video/referenceKling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
image-to-imagefal-ai/flux-pro/kontext/multiExperimental version of FLUX.1 Kontext [pro] with multi image handling capabilities
speech-to-textfal-ai/elevenlabs/forced-alignmentAlign the transcript and your audio recording using Elevenlab's forced alignment feature!
image-to-videoxai/grok-imagine-video/reference-to-videoGenerate videos using multiple reference images with xAI's Grok Imagine video model
image-to-videofal-ai/minimax/hailuo-2.3-fast/standard/image-to-videoMiniMax Hailuo-2.3-Fast Image To Video API (Standard, 768p): Advanced fast image-to-video generation model with 768p resolution
video-to-videominimax/h3-max-turbo/extend-videoExtend an existing video with H3 Max Turbo: add 1 to 15 seconds of prompt-guided footage, with optional audio references and output resolutions from 480p to 2K.
image-to-imagefal-ai/flux-2-flex/editImage editing with FLUX.2 [flex] from Black Forest Labs. Supports multi-reference editing with customizable inference steps and enhanced text rendering.
text-to-speechfal-ai/minimax/speech-02-turboGenerate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
vision
video-to-videobria/video/background-removal/v3Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.
text-to-videofal-ai/pixverse/v6/text-to-videoPixverse's latest v6 Model.
image-to-imagefal-ai/kling-image/o3/image-to-imageKling Omni 3: Top-tier image-to-image with flawless consistency.
image-to-imagefal-ai/z-image/turbo/image-to-imageGenerate images from text and images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.
audio-to-audiofal-ai/elevenlabs/voice-changerChange the voices in your audios with voices in ElevenLabs!
image-to-3dtripo3d/tripo/v2.5/image-to-3dState of the art Image to 3D Object generation. Generate 3D model from a single image!
unknownopenrouter/router/audioRun any audio capable LLM with fal. Process audio files — transcription, analysis, understanding, understand— using Gemini (Google) models. Supports wav, mp3, aiff, aac, ogg, flac, m4a. Powered by OpenRouter.
image-to-3dfal-ai/sam-3/3d-objectsSAM 3D enables precise 3D reconstruction of objects from real images, while accurately reconstructing their geometry and texture.
image-to-imagefal-ai/flux-kontext-loraFast endpoint for the FLUX.1 Kontext [dev] model with LoRA support, enabling rapid and high-quality image editing using pre-trained LoRA adaptations for specific styles, brand identities, and product-specific outputs.
video-to-videofal-ai/birefnet/v2/videoVideo background removal version of bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
text-to-audiofal-ai/elevenlabs/text-to-dialogue/eleven-v3Generate realistic audio dialogues using Eleven-v3 from ElevenLabs.
video-to-videoveed/subtitlesVEED’s Subtitles API transforms raw footage into polished, publish-ready content with professional burned-in subtitles starting at a base rate of $0.10 per minute.
trainingfal-ai/krea-2-trainerTrain a custom LoRA on your own images to teach Krea 2 a new subject, character, or style. Provide a set of training images (and an optional trigger word), and the trainer outputs LoRA weights you can use for inference with the Krea 2 LoRA endpoint.