An aarti is a harder test for an AI singer than an ordinary song. The words are fixed and familiar — listeners know exactly how each line should sound, so a mispronounced name is obvious in a way that a slightly odd pop lyric never is.
The good news is that the two things that matter most are both under your control: the script you paste in, and how explicitly you describe the arrangement.
Paste the text in Devanagari, stanza by stanza
Use Devanagari rather than Roman transliteration. Devanagari carries the vowel lengths and the aspirated consonants explicitly; Roman spelling makes the model guess, and divine names are exactly where a wrong guess is least forgivable.
Separate each stanza with a blank line. The model reads blank lines as structural breaks and uses them to phrase the melody, so an aarti pasted as one solid block tends to come out rushed and run-together.
- Devanagari script, with the danda punctuation kept.
- One stanza per block, blank line between blocks.
- Do not strip the refrain — repetition is part of the form.
Describe the arrangement, do not just say 'devotional'
A vague prompt gives you a generic result. Name the instruments — harmonium, tabla, manjira, temple bells — and set a slow tempo. Ask for a warm mix and clear diction in as many words.
Decide early whether you want a single voice or a congregational feel. An aarti sung by a group has a different character from a solo recitation, and the model will do either if you ask for it explicitly.
- Name instruments rather than relying on the word 'devotional'.
- Set a slow tempo — devotional material suffers when it is rushed.
- State whether you want solo or group vocals.
Keep it to one aarti per generation
Generate one aarti at a time rather than stringing several together into a single long text. Shorter, self-contained pieces come out cleaner and give you a take you can actually use; very long devotional texts are better handled as separate recordings that you sequence afterwards.
Listen with the text in front of you
Play the take once while reading along. You are checking for two things: whether every line is there, and whether the names are pronounced correctly. If one word lands wrong, regenerate — the model takes a different path each time and the same prompt will often fix it without any editing.