Skip to main content
Seed Audio 1.0 is ByteDance’s universal audio generation model, now available as a built-in node in ComfyUI. Unlike traditional text-to-speech systems that only read words aloud, Seed Audio understands the full spectrum of sound. It can generate speech, music, sound effects, ambient atmospheres, and multi-speaker dialogue from a single text prompt.

What Seed Audio 1.0 is good at

  • Multi-modal audio generation: Speech, music, sound effects, and ambient audio from one prompt. Describe what you hear, and Seed Audio produces it
  • Voice cloning: Clone voices from up to 3 reference clips (tagged @Audio1, @Audio2, @Audio3 in the prompt) for multi-character dialogue or consistent narration
  • Preset voices: Pick from built-in TTS 2.0 voices when you need reliable, high-quality narration without a reference clip
  • Image-driven voice: Derive a voice from a character image, matching the tone and style to the visual subject
  • Fine-grained control: Adjust speech rate, pitch, loudness, and sample rate independently for each generation
  • Multi-speaker dialogue: Name characters inline in the prompt and Seed Audio handles voice assignment, turn-taking, and emotional delivery
  • Up to 2 minutes per run: Generate extended audio clips suitable for narration, podcast clips, video dubbing, and sound design
With these capabilities, Seed Audio 1.0 fits a wide range of use cases including video voiceovers, multi-character audio dramas, product demos, e-learning narration, game audio prototyping, and accessibility content.

Use it in ComfyUI

Seed Audio 1.0 workflows

Run the text-to-audio, voice cloning, and image-driven audio workflows in ComfyUI, locally or on Comfy Cloud