image-to-videoCogVideoX-5B
fal-ai/cogvideox-5b/image-to-videoGenerate videos from images and prompts using CogVideoX-5B
Browse the public fal model directory. Models become callable through Kaista Cloud only after supply, pricing, and compatibility review.
image-to-videofal-ai/cogvideox-5b/image-to-videoGenerate videos from images and prompts using CogVideoX-5B
audio-to-audiofal-ai/stable-audio-3/medium/audio-outpaintingStable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.
text-to-imagefal-ai/janusDeepSeek Janus-Pro is a novel text-to-image model that unifies multimodal understanding and generation through an autoregressive framework
image-to-imagefal-ai/playground-v25/image-to-imageState-of-the-art open-source model in aesthetic quality
image-to-imagefal-ai/kolors/image-to-imagePhotorealistic Image-to-Image
text-to-imagefal-ai/flux-1/srpoFLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
text-to-imagefal-ai/z-image/base/loraLoRA endpoint for Z-Image, the foundation model of the Z- Image family.
image-to-imagefal-ai/flux-1/krea/image-to-imageFLUX.1 Krea [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
video-to-videofal-ai/flux-3-action/so101FLUX 3 Action turns what the robot sees into what it does next. Give it the scene camera image, the wrist camera image, the current SO-101 joint state and a plain-language instruction.
video-to-videofal-ai/wan-vace-14b/reframeVACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.
video-to-videofal-ai/bernini-r/edit-videoEdit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping the rest of the scene intact.
image-to-imagefal-ai/flux-1/srpo/image-to-imageFLUX.1 SRPO [dev] is a 12 billion parameter flow transformer that generates high-quality images from text with incredible aesthetics. It is suitable for personal and commercial use.
image-to-imagefal-ai/control-lightControlLight is a LoRA fine-tune of FLUX.2 [klein] 9B that enhances low-light images while preserving scene structure and fine details, with a single alpha parameter that gives continuous control over enhancement strength from subtle to full brightening.
audio-to-audiofal-ai/stable-audio-3/medium/base/audio-to-audioStable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.
image-to-videofal-ai/vidu/image-to-videoVidu Image to Video generates high-quality videos with exceptional visual quality and motion diversity from a single image
trainingfal-ai/flux-2-klein-9b-base-trainer/editFine-tune FLUX.2 [klein] 9B from Black Forest Labs with custom datasets. Create specialized LoRA adaptations for specific editing tasks.
text-to-imagefal-ai/ovis-imageOvis-Image is a 7B text-to-image model specifically optimized for quick, high quality text rendering.
video-to-videobria/bria_video_eraser/erase/promptA high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency
audio-to-audiofal-ai/tada/3b/text-to-speechA unified speech-language model that synchronizes speech and text into a single, cohesive stream via 1:1 alignment.
text-to-videofal-ai/longcat-video/text-to-video/720pGenerate long videos in 720p/30fps from text using LongCat Video
image-to-imagebria/embed-productSeamlessly embed products into any scene with pixel-perfect control, automatic perspective, and natural lighting. Trained on licensed data - risk-free for advertising and eCommerce production.
text-to-videofal-ai/kandinsky6-lite/text-to-videoKandinsky 6.0 Lite is a lightweight, high-speed text-to-video model from Kandinsky Lab, built for fast, efficient generation with strong prompt adherence.
audio-to-audiofal-ai/stable-audio-3/medium/audio-inpaintingStable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.
image-to-videomoonvalley/marey/i2vGenerate a video starting from an image as the first frame with Marey, a generative video model trained exclusively on fully licensed data.
text-to-imagefal-ai/omnigen-v2OmniGen is a unified image generation model that can generate a wide range of images from multi-modal prompts. It can be used for various tasks such as Image Editing, Personalized Image Generation, Virtual Try-On, Multi Person Generation and more!
text-to-audiofal-ai/csm-1bCSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.
image-to-imagefal-ai/post-processing/blurApply Gaussian or Kuwahara blur effects with adjustable radius and sigma parameters
image-to-imagefal-ai/qwen-image-edit-plus-lora-gallery/add-backgroundAdd a realistic scene behind the object with white background