This question gets a lot of confident marketing answers, so here is a straight one. Today's text-to-music models are built to generate audio from a written description, or to re-arrange audio you give them. Melody transfer — listening to a hum and composing a new song that follows that exact tune — is not a solved feature in the tools available right now. We tested it repeatedly on our own engine and on a hummed reference: the output follows rhythm and tempo well, but the melody you hummed does not survive.
So instead of selling a promise the technology does not keep, here is the workflow that does produce a finished song from something you sang.
What actually works: sing it, and the AI adds the band
Record or upload a real vocal take — you singing the song, not a hum — and the model builds a full arrangement underneath it: drums, bass, guitars, keys, whatever you ask for. Your voice stays exactly as you sang it. That is the mode the underlying engine is genuinely designed for, which is why it holds up.
Open autunes.com/jam, press Record, and sing while a live waveform draws what the microphone hears. Every recording becomes a take you can play back, keep or delete. You can also upload an existing recording (MP3, WAV, FLAC, M4A or OGG, up to 50MB). Pick a take, describe the sound you want, and press Add the band.
- Add the band: 10 credits
- Extend to roughly three minutes with an instrumental continuation: 20 credits more
- Your recorded vocal is the vocal in the finished track
- Available on every plan, including free
The other route: turn what you sang into a written song
If your recording is more of a sketch than a performance, Jam has a second mode. It listens to what you sang, writes complete lyrics around those lines — keeping your words as the chorus — and shows them to you before anything is generated. You edit the words until they are right, then the normal song pipeline produces a full track from them.
Be clear about what this does and does not carry over. Your words carry over. The melody does not: the tune is composed fresh to fit the style you asked for. For a lot of people that trade is fine, because the lyric was the idea they were protecting. Drafting the lyrics costs nothing; you only spend credits when you generate the song (10 credits on Zori5 Turbo, 15 on Zori5 XL).
How to get a usable take
Because your recording is the finished vocal in Add the band mode, recording quality matters more than it would for a guide track. A phone in a quiet room is enough; a fan, a TV or a live backing track in the background is not, because the model treats everything it hears as the vocal it must arrange around.
- Sing words rather than humming, so the arrangement has phrasing to follow
- Record somewhere quiet, close to the microphone, with nothing else playing
- Give at least 15 seconds; a full verse and chorus works better than a fragment
- Paste the lyrics you sang if you have them, so the band follows your phrasing
- For a rough idea rather than a performance, use the write-a-full-song mode instead