Choose a Music Generation Approach
Decide between a finished track, stems, loops, and MIDI before you generate, because that choice determines everything you can still change afterwards.
Learning objectives
- Match the output format to the edits you will need to make later
- Judge whether stems are worth the extra step for a given project
- Decide the vocal question before generating rather than after
- Recognise the requests that no current music model handles well
ToolDix original visual
Frame
Name the outcome and constraints.
Build
Try one bounded workflow.
Review
Keep evidence, revise, and share.
The most consequential decision in AI music happens before you write a prompt, and most people make it by accident. Generating a finished stereo track is the default because it is the default button, and a finished stereo track is the format with the fewest remaining options. Once the mix is baked, you cannot lower the drums under a voice-over, you cannot extend the middle by eight bars, and you cannot remove the instrument that fights the dialogue.
Four output formats, four levels of remaining control
A finished stereo track is one file, ready immediately, and permanently mixed. It is the right choice when the music plays alone and at full length — a background bed for a long video, a mood piece, a placeholder for a pitch. It is the wrong choice for anything that has to sit under speech.
Stems are the same music delivered as separate instrument groups: drums, bass, harmony, melody, vocals. This is the format that makes music usable in production. You can duck one layer under dialogue, drop the drums for a quiet section, and rebalance a mix that was mastered too loud. If your music will share a timeline with a voice, stems are not a luxury.
Loops are short, seamlessly repeatable sections. They are the right format when duration is unknown or interactive — game audio, product demos of variable length, anything that has to stretch. Building from loops means arrangement becomes your job rather than the model's, which is more work and total control.
MIDI or symbolic output gives you notes rather than audio. You choose the instruments afterwards, edit any note, and change key or tempo without artifacts. It costs an extra production step and it is the only format where the model's output is fully editable.
The routing question is simple: what will you need to change after you hear it in context? Answer that honestly, because the answer is almost never "nothing," and the format decision is difficult to reverse.
Stems are worth more than they look
- Broad EQ notch across everything
- Sidechain the whole mix
- Or simply turn it down
- All three make the music smaller
- Duck the one conflicting layer
- Everything else stays at level
- Music stays big, voice stays clear
- Nobody notices the intervention
Teams skip stems because they add a step and because the finished track sounds good on its own. Then the voice-over arrives and the track fights it.
With a stereo file your options are a broad EQ notch, sidechain ducking of the whole mix, or lowering everything — all of which make the music sound smaller and none of which solve the actual conflict, which is usually one instrument occupying the same frequency range as the voice.
With stems you mute or duck that one layer and keep the rest at full level. The music stays big, the voice stays clear, and nobody notices the intervention. That is the whole reason professional libraries ship stems.
Two practical notes. Not every service provides stems, and separating a finished track after the fact with a source-separation tool works reasonably well but leaves artifacts on the isolated parts — usable for ducking, not for soloing. And where stems are available, download them at generation time; regenerating later gives you a different performance even from the same prompt.
Decide the vocal question first
Vocals change every downstream decision, so settle it before generating.
No vocals is the default for anything under speech, because two voices in the same range compete and the listener has to choose. Most commercial and explainer work needs instrumental.
Non-lexical vocals — hums, oohs, wordless choir — add human warmth without competing for attention semantically. Frequently the best answer when a track feels cold.
Full lyrics make the music the message. Right for anthems, titles, and content where the song is the point, and wrong under dialogue in every case.
Any voice resembling a specific real artist is a separate category with legal and platform consequences, and it is not a stylistic decision. Several jurisdictions have introduced protections for voice likeness, and major platforms remove this material. Do not build a deliverable on it.
Also decide the language and whether lyrics need to be intelligible, since intelligibility constrains the arrangement more than most people expect.
What models still handle badly
Worth knowing before you promise something:
- Exact duration. Models produce approximate lengths. Plan to edit rather than to specify.
- Precise hit points. Music that must land on a specific frame needs editing, not prompting.
- Long-form structure. Coherence over several minutes with genuine development remains difficult; assemble sections instead.
- A specific existing song. Beyond capability and squarely a rights problem.
- Complex time signatures and genre-accurate niche styles. Output tends toward the centre of the distribution.
Practice
Take three real projects and write the format decision for each with one sentence of justification. Then take the one that will share a timeline with speech and generate it twice: once as a stereo file, once as stems.
Build both against the voice-over. Time how long it takes to make each one work. The gap is the argument for stems, and it is usually large enough to end the discussion.
Common mistakes
Defaulting to a finished track. It is the format with the fewest remaining choices, chosen because it is the default button.
Deciding vocals after generating. It changes the arrangement, so it changes the generation.
Assuming you can separate stems later. You can, imperfectly, and the artifacts show on anything soloed.
Prompting for a duration. Generate longer than you need and edit. The edit was always going to happen.
Sources and license context
These references informed the lesson. ToolDix adds its own explanation, workflow, and practice rather than reproducing source material. Every link below leaves ToolDix and opens the publisher's own site in a new tab.
- Microsoft Muzic (opens github.com in a new tab)External · github.com (MIT)
- AudioCraft (opens github.com in a new tab)External · github.com (MIT license for code; model weights under separate terms)
Keep going
Read these next on ToolDix.
Original lessons that build on what you just read.