Gradium Turns Voice Prompts Into Ready-to-Use Synthetic Speech

Gradium, a Paris-based voice AI company, has launched Voice Design, a feature that creates synthetic voices from written descriptions. Instead of choosing from a catalog, users can describe the voice they want and receive new candidates within seconds.
The feature is available in the Gradium API and Studio, and it is free on every plan, including the free tier. Gradium says Voice Design can turn a description of up to 500 characters into as many as five new voices, with no reference audio required.
This gives the company’s voice tools a new way to handle voice selection. Gradium already has 400 voices in its catalog, but Voice Design lets users create candidates from a written prompt rather than searching through existing options.
How Voice Design Works
Users can submit descriptions from 1 to 500 characters in English, French, Spanish, Portuguese or German. One request returns between one and five candidates, and those candidates are typically ready in three to five seconds.
Sampling is non-deterministic, so the same description can produce different results. That means users can receive a selection of candidates from one request instead of a single fixed output, giving them several generated voices to consider.
The feature is built into both Gradium’s API and Studio. A kept voice uses the same streaming Text-to-Speech endpoint as any catalog voice, with the same latency and output formats. In practical terms, Voice Design creates the voice, while the regular streaming endpoint handles its use after it has been kept.
Gradium places clear limits on candidates before they become kept voices. Audition text is capped at 100 characters, candidates are REST only, and the TTS WebSocket and Speech-to-Speech reject them. Unconverted candidates are deleted after 30 days.
Free Access Comes With Candidate Limits
Voice Design is available across all Gradium plans, but the number of candidates a plan can hold differs. The free tier holds five candidates, while paid plans hold 1,000 candidates.
That storage limit matters because the feature can return up to five voices from one request. A user on the free tier can keep a small set of candidates available, while paid plans provide room for a much larger collection before candidates are converted or removed.
The same access model applies across the API and Studio, so the feature is not limited to one interface. Users can create voices in either environment while working within the candidate restrictions set by Gradium.
Gradium Reports a 72.6% Test Win Rate
Gradium also shared results from a blind pairwise listening test focused on accent prompts. The test compared six voice design systems across five languages and included 7,627 comparisons.
Gradium reports a 72.6% win rate against the field and says it placed first in all five languages. The company also reports that its result was 13.6 points ahead of ElevenLabs.
The test was a vendor-run blind test, so the figures describe Gradium’s own evaluation of the systems. Still, the result gives a clear measure of what the company wants Voice Design to compete on: how well generated voices match written requests across multiple languages and accents.
Gradium also reports an 83.4% prompt adherence score on InstructTTSEval for English. Together, these results point to the company’s focus on following voice descriptions, not only producing speech that sounds natural.
What the Launch Changes
Voice Design moves synthetic voice creation from a catalog search toward a prompt-based process. A written description can produce up to five candidates in seconds, and the feature works across five languages without requiring reference audio.
The restrictions show that these candidates are not identical to catalog voices from the moment they appear. They can be auditioned only with up to 100 characters, work through REST, and cannot be used with the TTS WebSocket or Speech-to-Speech. Users also need to remember that unconverted candidates disappear after 30 days.
Even with those limits, the launch gives Gradium a direct route from description to synthetic voice. The feature became available on September 9, 2026, across the API and Studio, with access included on every plan and five candidates available on the free tier.




