image-to-videoH3 Max Camera Controls
minimax/h3-max/camera-controlsH3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videominimax/h3-max/camera-controlsH3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space
text-to-audio
audio-to-audiofal-ai/audio-understandingA audio understanding model to analyze audio content and answer questions about what's happening in the audio based on user prompts.
image-to-videofal-ai/veo3.1/reference-to-videoGenerate Videos from images using Google's Veo 3.1
text-to-imagefal-ai/flux-pro/kontext/text-to-imageThe FLUX.1 Kontext [pro] text-to-image delivers state-of-the-art image generation results with unprecedented prompt following, photorealistic rendering, and flawless typography.
text-to-imagefal-ai/recraft/v4.1/text-to-vectorRecraft V4.1 Vector turns prompts into fully editable SVGs with structured layers and clean geometry. Built for logos, icons, and illustration systems, it produces artwork that goes straight from generation into Figma or Illustrator.
video-to-videoblackforestlabs/flux-3/edit-videoFLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.
fal-ai/elevenlabs/audio-isolationIsolate audio tracks using ElevenLabs advanced audio isolation technology.
text-to-imagefal-ai/krea-2/turbo/loraGenerate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.
text-to-audiosonilo/v1.1/text-to-musicGenerates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.
text-to-audiofal-ai/kokoro/american-englishKokoro is a lightweight text-to-speech model that delivers comparable quality to larger models while being significantly faster and more cost-efficient.
image-to-3dtripo3d/p2/image-to-3dTripo P2 generates 3D models from a single image, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.
video-to-videofal-ai/kling-video/o3/pro/video-to-video/referenceKling O3 Omni generates new shots guided by an input reference video, preserving cinematic language such as motion, and camera style to produce seamless scene continuity.
text-to-videofal-ai/bytedance/seedance/v1.5/pro/text-to-videoGenerate videos with audio with Seedance 1.5
image-to-videogoogle/gemini-omni-flash/reference-to-videoGenerates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.
image-to-imagefal-ai/qwen-image-editEndpoint for Qwen's Image Editing model. Has superior text editing capabilities.
text-to-speechfal-ai/minimax/speech-2.8-turboGenerate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
image-to-videoxai/grok-imagine-video/v1.5/reference-to-videoGenerate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.
text-to-imagebytedance/seedream/v5/lite/text-to-imageText to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.
image-to-videofal-ai/sync-lipsync/v3/image-to-videosync-3 image to video turns a single still into a talking character, and works with any illustration or animated frame paired with a voice track
text-to-imagegoogle/nano-banana-liteNano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
image-to-3dfal-ai/hunyuan3d-v3/image-to-3dTransform your photos into ultra-high-resolution 3D models in seconds. Film-quality geometry with PBR textures, ready for games, e-commerce, and 3D printing.
video-to-videofal-ai/depth-anything-videoGenerates depth maps from video using Video Depth Anything (CVPR 2025). Produces per-frame depth estimation with temporal consistency across frames. Supports 3 model sizes (Small, Base, Large), 5 colormaps including grayscale, side-by-side comparison with the original video, and raw depth export as .npz. Useful for 3D reconstruction, video effects, compositing, and scene understanding.
text-to-speechfal-ai/inworld-ttsText to Speech Endpoint for Inworld's TTS-1.5 Max.
video-to-videoveed/lipsync/v2Generate production-quality lipsync from any audio using VEED's most advanced model yet.
image-to-imagefal-ai/qwen-image-edit-plusEndpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509. Has superior text editing capabilities and multi-image support.
video-to-videodecart/lucy-2-5/realtimeReal-time, prompt-driven video editing over WebRTC. Restyle, swap backgrounds, and add or replace objects live on a webcam or streamed feed at interactive latency.
image-to-videofal-ai/kling-video/o1/image-to-videoGenerate a video by taking a start frame and an end frame, animating the transition between them while following text-driven style and scene guidance.