How-To · 7 min read

How to Add Music to Your Vocals With AI (Your Voice Stays Untouched)

Published 26 September 2026

Short answer

Use Jam on Autunes. Record your voice in the browser or upload a recording of at least 5 seconds, and Jam measures its tempo and key, then arranges a band underneath (drums, bass, guitar, keys and more, your choice) while your vocal stays exactly as you sang it. It costs 10 credits on any plan, including free; +20 credits extends the idea to a full ~3-minute song, and +5 lifts your voice out of a recording that already has music so the old backing can be replaced.

You have a tune in your head, or a voice note of a song you wrote, or a cover you sang over a karaoke track that you wish sounded like a real production. What you do not have is a band. Most AI music tools solve that by throwing your voice away: they generate a new singer who sings something 'inspired by' your recording. That is not what a singer wants.

Jam on Autunes does the opposite. Your recording is the fixed point. The AI listens to it, works out its tempo and key, and builds the instruments around it, so the finished track is your actual voice with a band that follows you.

What Jam does, step by step

Open Jam from the sidebar. You can record straight into the browser or upload a file (up to 50 MB). The take needs at least 5 seconds of audio; a verse and chorus is better, because the band has more to follow.

If you upload, Jam asks what is in the file. 'Just my voice' means an a cappella recording, and the band is built straight under it. 'My voice with music' means the recording already has backing, for example you sang over a karaoke track. In that case Autunes first lifts your voice out of the mix (5 extra credits), then arranges the new band under the isolated vocal. For uploads you also confirm that the recording is yours or that you have the rights to use it.

Then you choose what to make:

  • Add the band: your voice, untouched, with drums, bass, guitar and keys arranged under it in your tempo and key. You can pick the parts, from drums and bass to strings, brass, synth, percussion and backing vocals.
  • Length: 'My take's length' plays the band exactly under what you sang. 'Extend to ~3 min' (+20 credits) grows the idea into a full song.
  • Write a full song: a new, fully produced song in your take's tempo and key, where your lines become the hook. You read and approve the words before any credits are spent.
  • Lyrics: optional. Typing what you sang helps the band follow your phrasing; leave it empty and Jam transcribes your take.
  • Language: detected automatically, or pick it yourself for mixed-language takes.

What it costs

Jam is available on every plan, including free. A take with a band costs 10 credits, the same as a regular song. The add-ons are priced by the extra work they need on the GPU: +5 credits to separate your voice from an existing mix, +20 credits to extend to a full-length song. Free accounts get daily credits, so a basic Jam fits within a day's allowance.

Recording tips that make the band sound right

The band follows what it hears, so the recording decides how tight the result is. None of this needs a studio; a phone in a quiet room is enough.

  • Keep a steady tempo. Tapping your foot or counting '1, 2, 3, 4' in your head before you start keeps the pulse even, and an even pulse is what lets the drums lock in.
  • Record without music playing into the microphone. If you sing along to something, wear headphones so it does not leak into the take.
  • Hold the phone about a hand's length from your mouth and avoid rooms that echo. A clean, dry voice is easier to build under and mixes better.
  • Sing the whole idea in one take. Stopping and restarting changes the tempo, and the band cannot guess which tempo you meant.
  • Sing in the key you want the song in. Jam measures the key from your voice, and the band is arranged around it.

Jam or remix: which one do you need?

Both start from a recording, but they treat the voice in opposite ways. A remix or cover regenerates the whole song, singer included: the original is a reference, and a new vocal is sung from the lyrics. Jam never re-sings anything. Your recording stays exactly as it was and only the music is new.

That makes Jam the right tool whenever the performance itself matters: your own voice, a take with feeling you cannot repeat, or fast rap. Rap is the clearest case. A regenerated vocal can drift on dense, fast lines, while Jam keeps every bar because it keeps the recording. If you want the song to sound like a different singer, or to change the words entirely, that is a remix job instead.

Honest limits

Jam arranges around the take it gets. If the tempo wanders a lot within the take, the band follows the measured average and will sound loose in places. If the voice drifts off pitch, the band still plays in the measured key, which makes the drift more audible, not less. And Jam is for your own recordings or ones you have the rights to; it is not a way to rebuild someone else's released song.

Try it yourself

Make a full song from text — free to start, no skills needed.

Start creating free

Frequently asked questions

Is Jam free to use?

Yes, Jam works on every plan including free. A basic take with a band costs 10 credits; extending to a full song adds 20, and separating your voice from an existing mix adds 5.

Will the AI change my voice?

No. In 'Add the band' mode your vocal is kept exactly as recorded, and only the instruments are generated. Only the 'Write a full song' mode creates a new production built on your idea.

Can I upload a song where I sang over a karaoke track?

Yes. Choose 'My voice with music' and Autunes lifts your voice out of the recording first, then builds a new band under it, for 5 extra credits. Confirm you have the rights to the recording.

How long does my recording need to be?

At least 5 seconds, up to a 50 MB file. A verse and a chorus gives the band enough to follow; to turn a short idea into a full song, use 'Extend to ~3 min'.

Does Jam work for rap?

Yes, and it is the best option for fast rap on Autunes, because your recorded bars are kept as they are rather than re-sung by the model.

Related reads