Replaced two other subscriptions with this. The integration story is what sold me: it fits into the tools I already use instead of asking me to move.
FLUX
Multimodal FLUX models for image, video, audio, and action prediction

About FLUX
Black Forest Labs develops FLUX models for visual intelligence across image, video, audio, and action prediction. Its models are designed to understand, reason about, and act in the world, with FLUX 3 presented as a multimodal model.
The models can be accessed through a browser playground, an API, or open-weight deployments on a user's own infrastructure. The site describes support for text, images, and keyframes as inputs, with image generation, video clips of up to 20 seconds, generated audio, and robotics-oriented visual and control predictions.
The API is positioned for production workloads, while open-weight access supports deployment, fine-tuning, and customization. Enterprise options, API pricing, documentation, and licensing are available through the site.
Similar tools
OpenMontage
— Open-source agentic video production from brief to editable renderOpen-source agentic video production from brief to editable render
Midjourney
— AI image and video models with a community-funded research labAI image and video models with a community-funded research lab

HyperFrames
— Open-source framework for rendering HTML into deterministic MP4 videoOpen-source framework for rendering HTML into deterministic MP4 video
Suno
— AI music creation from prompts, lyrics, and audioAI music creation from prompts, lyrics, and audio

VoiceStudio
— Local voice cloning, dubbing, transcription, and text-to-speechLocal voice cloning, dubbing, transcription, and text-to-speech
ElevenLabs
— AI voice, audio, creative, and conversational agent platformAI voice, audio, creative, and conversational agent platform