Suno’s logo. Image: Suno
Suno, the AI music generator, can now read your words aloud and score them at the same time. The company launched Speech on Thursday, a beta it calls “the first audio model that generates voice and music together as one cohesive track”, and opened it to every user.
How Speech works #
Speech is built into the Suno app. You type an idea, a poem or something you have written, then describe the kind of voice and the style of music you want behind it. The model produces the narration and the backing track in one take, rather than generating a voice and mixing it over a separate song.
Suno has posted a tutorial showing how it works:
In the launch post, Suno’s chief product officer Jack Brody says the team kept finding uses that “made us laugh, or occasionally surprised us with how moving they could be”. Its own examples include turning friends’ texts into “wildly overproduced dramatic readings”, giving voice notes “unnecessarily epic scores”, and making meditations, pep talks and bedtime stories for their children.
Suno says it tested Speech with a small group of users for a month before opening the beta. It didn’t say which plans include it, how many generations it allows or how the model was trained.
“Beta really does mean beta” #
The company is upfront that the feature is rough around the edges. Brody writes:
Occasionally, British accents can wander off to Australia and back. Dramatic s may be very dramatic. You will almost certainly discover uses for this that never occurred to us.
Jack Brody, chief product officer, Suno
Brody frames the launch as part of a push beyond songs into what Suno calls “creative entertainment”. “Music will always be at the heart of Suno and what we build,” he writes, but “our vision has always extended to other forms of human expression.”
A launch in the middle of a legal fight #
Speech arrives three weeks after Suno released v6, its first music models built with industry partners Warner Music Group, BMG and Believe. Not everyone signed up. Universal Music and Sony Music sued Suno again on September 18, accusing it of copying 60,202 of their recordings and saying v6 was partly trained on output from its earlier, unlicensed models. Their first lawsuit is still going.
In Germany, a Munich court ruled against Suno on July 31 in a case brought by the music rights society GEMA, ordering it to stop using the protected songs to train its models. And in September, musicians including Jason Isbell and David Lowery sued Suno in Massachusetts, arguing it trained its model to imitate artists by name and collected their “voiceprints” without consent.
That last claim is worth watching now that Suno is moving into the spoken word. Voice cloning has already landed AI companies in court: a Japanese voice actor sued TikTok over an AI copy of his voice, and an appeals court handed AI firms their first big copyright defeat last week.
Why it matters #
Speech takes Suno from songs into narration, putting it closer to voice AI companies than to music apps. It also widens the question hanging over all of its lawsuits: what the model learned from, and whose voices it can sound like.
Sources: Suno, “Introducing Speech (beta)”; Suno, “Introducing v6”; Music Business Worldwide; Variety; Rolling Stone.