Paris-based AI startup Gradium, a spinout from the Kyutai research lab, has launched Voice Design, a feature that generates entirely new synthetic voices from text descriptions in seconds. Available via API and in Gradium Studio, the tool allows developers and creators to specify attributes like age, gender, accent, pitch, and job context without providing audio samples, eliminating traditional voice licensing and consent hurdles.

Voice Design accepts descriptions up to 500 characters across five languages and returns candidate audio embeddings within three to five seconds. Once a candidate embedding is promoted to a custom voice slot, it streams via Gradium’s standard text-to-speech endpoints at baseline latency. The service offers a free tier with five custom voice slots and paid plans accommodating up to 1,000 slots.

In blind pairwise testing conducted by Gradium across 7,627 human evaluations, the model achieved a 72.6% win rate against competing voice design platforms, outperforming ElevenLabs, Inworld, Fish Audio, and MiniMax. Regional accent benchmarks showed particularly high preference scores for Quebecois French and Rioplatense Spanish, addressing long-standing gaps in global voice catalogs.

Why it matters

  • Text-based voice synthesis bypasses audio cloning consent issues and expands options for regional accents in voice apps.

  • Startup Gradium poses direct competition to established text-to-speech leaders like ElevenLabs through dynamic prompt-driven voice generation.

  • Developers can rapidly prototype specialized voice personas across five languages without sourcing reference voice actors.

Source: marktechpost.com