Skip to main content
FLUX 3 is Black Forest Labs’ multimodal foundation model, announced July 23, 2026 and currently in Early Access. It jointly learns from images, videos, and audio within a unified architecture built on Self-Flow, their approach for aligning multimodal generation and understanding in the same model. Instead of treating each modality in isolation, FLUX 3 learns a shared representation of the world: how objects hold together, how things move, and how events sound. Capabilities and limits may change during the Early Access rollout. For video, FLUX 3 creates highly diverse clips with native audio up to 20 seconds long in a single generation. All outputs come with synchronized audio generation, including ambient sound, speech, and effects. It supports text-to-video and image-to-video generation, with multi-shot output that chains individual clips into longer sequences and flexible aspect ratios.

What FLUX 3 Video is good at

  • Native audio: Every video comes with synchronized audio, including ambient sound, speech, and effects
  • Up to 20 seconds: Generates long clips in a single pass
  • Text and image input: Generates video from a text prompt, or continues from a starting frame image
  • Flexible formats: 720p or 1080p resolution, 24fps, aspect ratios from 9:16 to 21:9
  • Multi-shot output: Chains individual clips into longer sequences for extended storytelling

Example outputs

Text-to-video generation from a single prompt, with synchronized audio, motion, and physics: Image-to-video generation from a starting frame and a text prompt:

Use it

FLUX 3 Video workflows

Run the text-to-video and image-to-video workflows in ComfyUI, locally or on Comfy Cloud

Call it from code

Call it over HTTP through Comfy Router, with copy-paste Python, TypeScript and cURL snippets
Both paths run on the same Comfy account and are billed in the same credits. See Partner Nodes pricing for per-model rates and Concurrency limits for how many requests you can have in flight.