MMAG (Multi‑Control Mixed Audio Generation) is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It addresses the lack of evaluation resources for emerging mixed-audio generation systems that synthesize complex acoustic scenes containing speech, music, and sound effects simultaneously.
① Main Set
text‑only control for holistic assessment — ~4,000 clips
② Voice Cloning Set
speaker identity control via short voice prompt — ~690 clips
③ Timestamp Set
temporal control via timestamp‑detailed captions — ~1,800 clips
We select cross‑domain audio samples from AudioCaps, VGGSound, and MECAT, and construct detailed captions covering speech transcriptions, speaker attributes (age, gender, dialect, emotion), sound events, musical information (genre, instruments, mood), and temporal ordering. All annotations are curated through a multi‑expert pipeline with manual verification.
Comparisons with Other Benchmarks
Comparison of MMAG with existing benchmarks. Domain columns indicate the audio categories covered. Condition columns indicate the types of annotations or control signals provided.
🔧 Construction
Annotation Pipeline
Overview of the MMAG annotation pipeline. Expert models extract fine-grained attributes from audio clips. An LLM aggregates these annotations into coherent overall and timestamped captions, while speech segments are further processed to construct voice prompts for voice cloning evaluation. Human inspection is performed to ensure annotation quality.
Caption Example Comparison
Illustrative example of two caption types for the same audio clip. Overall captions annotate speech, sound, and music, along with relative temporal order. Timestamped captions provide precise timestamps for foreground events and speech segments.
🎧 Main Set
📌 Sample #1
📝 Caption:
A young adult male speaker in a neutral manner discusses creative methods in American English in a quiet indoor studio setting: "Tests here of colors, different goals. Zach really wanted the costume to feel ancient and broken down, so we're playing with that with our paint work here." Intermittent electronic music plays throughout, incorporating bass, drums, and synthesizer with dark, deep, and funny moods.
GT
AuDirector
JavisDiT++
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
Ming‑0.5B
Ming‑16.8B
📌 Sample #2
📝 Caption:
A middle-aged American man is neutrally conducting a vehicle review outdoors in American English: "Hello everyone. Today, we're going to take a quick walk around and look at this 2011 Jeep Wrangler Sport." Ambient traffic noise and wind continue throughout the recording with occasional microphone wind noise. The environment suggests an outdoor urban setting like a dealership or street.
GT
AuDirector
JavisDiT++
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
Ming‑0.5B
Ming‑16.8B
📌 Sample #3
📝 Caption:
A young adult male with a US accent provides neutral-toned finger-snapping instructions in English: "Your middle finger. Step 3. Let your ring finger and pinky rest against your palm. Step 4. Now quickly, swipe your middle finger down so that it hits." Electronic pop music happy with synthesizer plays throughout. A finger snap occurs at the end. Recorded in an indoor studio environment.
GT
AuDirector
JavisDiT++
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
Ming‑0.5B
Ming‑16.8B
📌 Sample #4
📝 Caption:
An elderly male narrator with a Scottish accent confidently spoke in English: "By ones and by twos. A raider infiltrates." The fearful electronic music played by drums and synthesizers ran through the entire background sound, accompanied by continuous cricket chirping, establishing an outdoor nighttime setting.
GT
AuDirector
JavisDiT++
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
Ming‑0.5B
Ming‑16.8B
📌 Sample #5
📝 Caption:
Energetic rock music with guitar and drums in a fun mood accompanies the clip. Near the end, a whoosh sound effect occurs, captured in a studio environment.
GT
AuDirector
JavisDiT++
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
Ming‑0.5B
Ming‑16.8B
🎤 Voice Cloning Set
📌 Sample #1
📝 Caption:
A male speaker aged 30-40 with an American accent speaks neutrally in English: "And you can find these depending on which one you're looking for." Background sizzling persists throughout most of the recording with intermittent handling dishes. The environment is a quiet indoor space suggesting a kitchen.
🎙️ Reference Prompt Audio(ground‑truth speaker)
Generated samples:
GT
AuDirector
UniAVGen
Ming‑0.5B
Ming‑16.8B
📌 Sample #2
📝 Caption:
A 20-30-year-old male speaker with an American accent explains a guitar technique in English with a neutral tone: "Minor and D, and our first strumming pattern we're gonna do is just four strums straight down, like this." Acoustic guitar unbeat playing accompanies the explanation. The audio appears to be recorded in a quiet indoor space.
🎙️ Reference Prompt Audio(ground‑truth speaker)
Generated samples:
GT
AuDirector
UniAVGen
Ming‑0.5B
Ming‑16.8B
📌 Sample #3
📝 Caption:
A happy 30-40 year old female speaker with a US accent presents a children's toy in English: "Six noises and it's very easy to move around on all surfaces as you can see. This is for ages three and above made by Just Play and requires three double A." Mechanical whirring and ticking from the toy are audible in the background. The indoor setting remains generally quiet.
🎙️ Reference Prompt Audio(ground‑truth speaker)
Generated samples:
GT
AuDirector
UniAVGen
Ming‑0.5B
Ming‑16.8B
📌 Sample #4
📝 Caption:
In a kitchen setting with dishwashing sounds, a middle-aged male with an English accent speaks neutrally in English: "It's much easier to clean up a bowl when it's soft than when it's hard." Running water is heard throughout, while clattering dishes occur intermittently.
🎙️ Reference Prompt Audio(ground‑truth speaker)
Generated samples:
GT
AuDirector
UniAVGen
Ming‑0.5B
Ming‑16.8B
📌 Sample #5
📝 Caption:
A happy middle-aged male speaks about having taken hydrocodone pills during video recording in American English. He states: "It was the day that I did my video and I dropped a couple of uh hydrocodones so I'm on the." A tick-tock sound happens briefly during the speech, and whistling emerges afterward. The audio occurs in a quiet indoor environment with slight reverberation.
🎙️ Reference Prompt Audio(ground‑truth speaker)
Generated samples:
GT
AuDirector
UniAVGen
Ming‑0.5B
Ming‑16.8B
⏱️ Timestamp Set
📌 Sample #1
📝 Timestamped Caption:
From 0.90s to 1.15s, a dog barked. Around 4.90s to 5.14s, the dog barked again. A female speaker aged 30-40 in a neutral state stated in US English, from 5.50s to 6.90s, she said:" You can hang out in here and hide." From 6.9s to 7.3s, overlapping the end of speech, the dog barked again. From 8.60s to 9.20s, the dog barked once more. Throughout the recording, a slight echo persisted as background.
GT
AuDirector
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
📌 Sample #2
📝 Timestamped Caption:
From 0.18 to 0.97s, a young female speaker spoke fearfully in British English:"Oh dear, someone's.". From 2.11 to 10.00s, terrifying electronic music with a synthesizer playing. From 5.00 to 8.87 seconds, she continued:"Someone's been messing up the shit round here. Oh, here he is!" The environment suggests a controlled indoor setting with minimal background interference.
GT
AuDirector
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
📌 Sample #3
📝 Timestamped Caption:
A middle-aged American woman speaks neutrally in an indoor craft setting from 3.91 to 5.66s: " And then you end up with a nice fluffy rug." From 7.26 to 10.00 seconds, she continues "These work great as bathroom rugs, or if you want a larger rug,". Prior to this, 0.0 to 2.96 seconds feature rustling fabric throughout and a sewing machine starting from 1.28 to 2.96 seconds. Indoor ambience spans the entire recording.
GT
AuDirector
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
📌 Sample #4
📝 Timestamped Caption:
At the beginning, a deep ominous music with low-frequency reverberation starts and continues throughout. From 0.00s to 3.54s, a middle-aged male speaker with a US accent calmly states, "That's my inspiration to riding in the backcountry." From 3.54s to 8.28s, the deep timpani percussion occured. The entire recording suggests an open outdoor setting.
GT
AuDirector
LTX‑2
MOVA
Ovi
UniAVGen
Dasheng‑Base
Dasheng‑Fine
📌 Sample #5
📝 Timestamped Caption:
Throughout the recording: slow-tempo cello music with sad tones. From 2.415s to 3.58s, a young British man with a British accent spoke neutrally in English:"It's less obvious than going." From 5.50s to 8.10s, he said: "You know, you just kind of smooth the edges out there." Set in a quiet indoor environment with clear audio quality throughout.