
Suno added speech two weeks after UMG added 60,202 songs to its claim
Suno opened a public beta of Speech on Thursday, a feature that generates a spoken voice from a script or a prompt and lays AI background music under it in the same track. The Verge reads the move as diversification away from a music generator carrying lawsuits. The docket puts dates on that reading.
“Today, we're expanding what's possible in Suno with Speech: the first audio model that generates voice and music together as one cohesive track.”
— Jack Brody, Suno
Jack Brody, Suno's chief product officer
The product, in the parts that matter
- Two modes: Simple takes a description, Advanced takes your own script
- Advanced settings cover the voice's gender, speaking style and how much each generation varies
- The music bed is optional and switches off with a toggle
- A single generation runs to about eight minutes, roughly 1,200 words at normal speaking pace
- It ships on web and mobile at once
Eight minutes is the number to hold. That is a podcast segment, a long explainer or most of a conference talk, produced from one prompt with scoring attached. ElevenLabs has sold text to speech since 2023 and Adobe and DeepMind were in this years earlier, so the voice itself is familiar ground. Generating the voice and the music as one cohesive track is what Brody is claiming. Multimodal AI usually means text and images; here it means one model producing two audio layers that have to agree with each other, and the claim is about workflow. The market around it is not small either: we looked at a voice-typing startup valued at $2bn in July.
The other side of the launch
Warner settled with Suno in November 2025 in a deal reported at around $500m, licensing its catalogue for training and taking control over how AI likeness and outputs get handled. Universal and Sony did not follow. Reporting through April 2026 described their talks with Suno at a hard impasse over the walled garden question: whether tracks generated on Suno can travel beyond it. Then in mid-September Universal filed a second suit, adding 60,202 works to its copyright claim. Run the statutory arithmetic and that filing alone carries a theoretical ceiling of $9.03bn, at the $150,000 per work that US law allows for wilful infringement. Courts almost never award the ceiling. The number still sits 18 times above what Warner reportedly took to settle, and it is the figure behind the word diversification. Speech arrived about two weeks later.
What it walks into instead
Moving from recorded music to synthetic voice changes which rights holder sits on the other side of the table, and someone still sits there. We covered the Save Our Voices Now campaign in August, in which UK actors pressed for legal ownership of their own voice, and that fight is earlier in its life than the music one. A tool that lets you set a voice's gender and style, then run it for eight minutes, is the product that campaign was describing.
Brody was candid about the state of it, saying beta really does mean beta, that British accents occasionally wander off to Australia and back, and that dramatic pauses may be very dramatic. The question worth tracking is whether the second product ends up in the same courtroom as the first, and the campaign in London suggests someone has already asked it. Voice has been crossing into ordinary business workflows for a while, and we counted that shift in September.
Informational material, not investment advice. Settlement figures come from reporting, and the statutory ceiling is an upper bound rather than a forecast.

Comments (0)
No comments yet — be the first!
The market talks all day. We write when it says something
Short, and it tells you why it came
Related news

Bitcoin is at $86,397 and the rate bet moved before the data, not after it

SpaceX welded a valve with the tank full, on a business it is winding down

South Korea's retail cap is $70,000 a year, and it resets at every venue
Most readTop 7
Silicon Valley Workers Are Wearing Noise-Cancelling Masks to Dictate AI Prompts
335AI


