When we shipped voice cloning on the Zori models, the pitch was simple: sing for twenty seconds, no music, just you, and Autunes sings your song back in your voice. Same melody, same words, your timbre. We expected people to use it on the songs they had just generated. That is what it was built for, and that is where it shines.
What we did not fully expect was how fast people would test its limits. Within days of launch, a large share of the clone jobs coming through the pipeline were not Autunes songs at all. They were uploads: film songs, worship recordings, chart singles, a whole choir with reverb and a string section, with a twenty-second phone recording of someone singing in their kitchen as the reference. And the model went ahead and did it. It split the singer out of the mix, converted the lead vocal, and mixed the band back in.
It was a good demo of the model. It was a bad idea.
Two things became clear at once. First, the technology is genuinely capable: singing voice conversion on a full commercial mix, from a sample shorter than a TikTok, is not something most tools can do at all. Second, we had built a machine that let anyone put their own voice on a recording they do not own.
That second part is not a grey area. A commercial recording has two layers of rights sitting on it: the composition, owned by the writers and publishers, and the master, owned by the artist or label. Re-singing a master with a cloned voice creates a derivative of both without permission. It does not matter that the voice is yours. The recording is not, and the platform that processed it carries the risk alongside the person who uploaded it. We are a small team building a music product for India and the world, and being the company that made it easy to clone onto other people's masters is not a position we want to defend.
What changed on 17 September 2026
'Sing it in my voice' now only runs on songs generated on Autunes. Those songs are yours: you wrote or described them, the engine composed and sang them, and the license that comes with your plan covers them. On an uploaded track the button is greyed out with the reason spelled out, and the server refuses the job even if someone finds a way around the interface.
We used the same release to tighten the part of the feature that decides whether a clone sounds like you at all: the reference sample. Voice conversion is only as good as the twenty seconds it learns from, and a quiet, half-silent phone clip produces a clone that is barely intelligible. So the sample is now recorded inside Autunes, with a live level meter, and a take that is too short, too quiet or mostly silence is refused on the spot instead of costing you credits. What passes is trimmed and normalised before the model hears it, and the conversion runs at a higher step count than before. The words come through cleaner, the voice comes through closer.
- Works on: any song you generated on Autunes, in any language the platform sings.
- Does not work on: uploaded recordings, including your own covers of other people's songs.
- Reference: record 10 to 30 seconds a cappella in the app; saved voices can be reused on every new song.
- Cost: unchanged, and refunded automatically if the sample is rejected.
What this means for you
If you want to hear yourself on a song, make the song on Autunes first. Describe it or paste your lyrics, pick the language and style, generate, then open 'Sing it in my voice' on the result. The whole loop, from idea to a finished track in your own voice, takes a few minutes and you own every part of it.
If what you wanted was to sing over a favourite film song, that is exactly the use we have stepped away from, and we would rather tell you plainly than let the feature quietly fail. The model can do it. We have decided that it should not.
Where voice cloning goes from here
The restriction is on the input, not on the ambition. The next releases are about making the clone more yours: longer reference samples for people who want to capture their range, a cleaner path for duets where two saved voices share one song, and voice personas you can hand to collaborators on a shared project. All of it on music generated here, where the rights are simple and the result is something you can actually publish.