Sanskrit is unusually unforgiving for a text-to-music model. The difference between a short and a long vowel, or between an aspirated and an unaspirated consonant, is not a stylistic detail — it changes the word. A tool that treats Sanskrit as decorative Latin script will produce something that sounds vaguely devotional and is, to anyone who knows the text, wrong.
Sanskrit is a natively supported vocal language in Autunes, so it is sung by a model that was given Sanskrit rather than one guessing from English phonetics.
Write the shloka in Devanagari
Paste the shloka in Devanagari, not in Roman transliteration. Romanised Sanskrit forces the model to guess where the long vowels and aspirates go, and it will guess wrong often enough to be noticeable.
Keep the danda punctuation. The single danda and double danda mark where a line ends, and the model uses them to phrase — removing them tends to make the recitation run lines together.
- Devanagari script, not Roman transliteration.
- Keep the danda and double danda line breaks.
- Put each shloka as its own stanza, separated by a blank line.
Choose a style that leaves room for the words
Sanskrit recitation needs space. A dense, fast arrangement fights the text. Prompt for something slow and open — harmonium, tanpura, soft tabla, temple bells — and set a low tempo. The instrumentation should sit behind the voice rather than compete with it.
If you want plain chanting rather than a song, say so in the prompt and keep the instrumentation minimal.
- Slow tempo; devotional or classical framing.
- Harmonium, tanpura, tabla, bells — sparse rather than layered.
- Ask for clear diction explicitly in the prompt.
Check the recording before you publish it
Listen once with the text in front of you. The things to check are the aspirates, the long vowels and the sandhi joins — those are where an AI vocal is most likely to slip. If a specific word lands wrong, regenerating with the same prompt usually fixes it, because the model samples a different path each time.