Black Forest Labs’ multimodal frontier model for visual creation.
Generate video with native audio from text or images — up to 20 seconds at 720p or 1080p.
FLUX 3 is Black Forest Labs’ new multimodal foundation model, the next step beyond the FLUX and FLUX.2 image families. Instead of treating still images as an isolated task, FLUX 3 jointly learns from images, video, and audio inside a unified architecture. The goal is not just prettier frames — it is a stronger sense of how scenes look, how motion unfolds, and how events sound. That shared backbone powers creative generation across stills and short-form video. On Flux.art you can use FLUX 3 Video now: text-to-video and image-to-video with optional native audio, draft previews, and clips up to 20 seconds.
FLUX 3 is designed as a single multimodal system trained across image, video, and audio signals — not a loose stack of separate models glued together after training.
By learning how scenes look, move, and sound together, FLUX 3 captures structure, motion, and visual causality rather than isolated pixel patterns.
The same model family targets high-quality image generation and short video with optional audio — helping keep look and feel consistent across a campaign.
From marketing visuals to social short-form and brand storytelling, FLUX 3 is aimed at production-oriented creative use cases, not research demos alone.
FLUX 3 marks a shift from “image-only” generators toward models that treat vision as multimodal and continuous in time. For teams shipping content every week, that change affects quality, consistency, and workflow design.
Headline capabilities creators should know when researching FLUX 3 for image and video workflows.
Generate short video clips up to about 20 seconds with optional native audio. Built for expressive faces, event-linked sound, and multilingual scenarios in ads and social short-form.
Image synthesis and editing with stronger complex-prompt following and multilingual text rendering than earlier FLUX versions, plus a wide stylistic range for marketing, product, and design assets.
One architecture trained across image, video, and audio — designed so stills and motion can share a more consistent understanding of objects, lighting, and scene structure.
Handles complex prompts and high-accuracy text in multiple languages — critical for posters, UI mockups, packaging, and multilingual campaigns.
Supports a wide range of output styles and aspect ratios so teams can explore photoreal, illustrated, and branded looks without switching model families for every format.
Aimed at creative tooling, media, design, and e-commerce storytelling — the same jobs teams already run with FLUX.2-class image models, expanded into short-form video.
Map FLUX 3 to concrete creative jobs so your team can plan campaigns and content roadmaps.
Concept spots, product teases, and social clips that need coherent motion plus audio in one pass. Multilingual expression support is especially useful for global campaigns.

With FLUX 3 Image, expect continuity with FLUX.2 strengths — prompt following, text rendering, and editable stills — while multimodal training improves consistency across a campaign’s still and motion kit.
Show materials, lighting, and short motion around a product without rebuilding the entire shoot. Useful for launch teasers, variant testing, and lifestyle scenes that match your brand look.
Generate FLUX 3 Video on Flux.art in a few steps.
Go to the AI Video Generator and select FLUX 3 for text-to-video or image-to-video.
Choose 5–20 seconds (or Auto), 720p/1080p, native audio, and optional Draft mode for a fast preview.
Describe the scene, or upload up to 10 reference images — first starts the clip, last ends it.
Run the job, preview the MP4 with synced audio, then download or iterate with Draft mode.
Common questions about FLUX 3 and how it compares to FLUX.2.
Generate cinematic clips with native audio from text or images — live now on Flux.art.