VoxCPM voice cloning and TTS were tested here with six deliberately practical scripts: order data, narration, support copy, a short ad read, and two cloned-voice cases. The goal was not to crown a universal voice model. It was to see where this Wiro deployment stays intelligible, where extra diffusion steps help, and where text should be normalized before it reaches speech.
The six audio players below are the original Wiro-hosted outputs. No output was regenerated for this update. That matters for the timing and cost notes: the original run records are not attached to these media files, and the Wiro model documentation does not publish a price or a measured run time for this endpoint. Those fields are therefore marked as not recorded rather than filled with estimates.
Model page: VoxCPM on Wiro. Wiro exposes four relevant inputs: prompt, cfgValue, inferenceSteps, and the optional pair inputAudio plus referencePrompt. The reference file and its transcript must be supplied together, or both left empty. The hosted documentation describes cfgValue as guidance for prompt adherence: higher can adhere more closely, but may sound worse. Higher inferenceSteps can improve results at the cost of speed.
What the VoxCPM voice cloning and TTS test set checked
The first four samples test default synthesis, not cloning. They move from structured commercial text to more natural prose, then deliberately reduce the step count for a speed-versus-smoothness check. Tests five and six add a source speaker and an exact reference transcript. Together, they test whether a clean reference helps a new script retain a speaker-like character, and whether that advantage survives text packed with URLs, punctuation names, and code-like tokens.
- Default guidance:
cfgValue=2.0in five samples;2.3in the fast ad read. - Steps: 5, 10, or 20, depending on the test.
- Voice cloning: tests 5 and 6 use both a reference MP3 and its transcript.
- Run time and Wiro cost: not recorded for each existing output; the Wiro documentation does not state an endpoint price.
Upstream VoxCPM materials describe a tokenizer-free, diffusion-autoregressive TTS approach. The project’s official GitHub repository and its Hugging Face model card are useful technical context, but they describe the broader and newer VoxCPM line. They should not be treated as a benchmark or pricing claim for this specific Wiro-hosted endpoint.
Test 1: numbers, currency, and a tracking code
Parameters: cfgValue=2.0, inferenceSteps=10. Run time: not recorded. Cost: not recorded.
This is the operational test. The output handles a short confirmation with several categories that often break TTS: a five-digit number, a decimal amount, a range, and a hyphenated code. In the sample, the speech stays easy to follow and the pauses work for a transactional message. It is a good fit for concise order updates, IVR fragments, and status notifications. It is not proof that every identifier will be pronounced correctly, so production flows should still normalize account numbers, decimals, and abbreviations in the text layer.
Test 2: calm narration
Parameters: cfgValue=2.0, inferenceSteps=20. Run time: not recorded. Cost: not recorded.
This sample asks for phrasing rather than data accuracy. The longer sentences retain a measured pace and the pauses land at sentence boundaries. Compared with the 10-step business samples, the 20-step setting is the sensible quality-first choice for a short narration where waiting a little longer is acceptable. The sample alone cannot establish a numeric latency gain, because its original job timing was not retained.
Test 3: support message
Parameters: cfgValue=2.0, inferenceSteps=10. Run time: not recorded. Cost: not recorded.
Short clauses make this output work. Each instruction has room to land, and the tone remains consistent across the warning and the fallback instruction. This is the strongest pattern for support automation: write plain clauses, put the critical action in its own sentence, and avoid cramming a URL or an unspoken symbol into the same line. For a phone agent, pair this kind of output with the related realtime voice conversation guide.
Test 4: fast ad read
Parameters: cfgValue=2.3, inferenceSteps=5. Run time: not recorded. Cost: not recorded.
This is the speed stress case. Five steps are the lowest setting in this set, while guidance rises slightly to 2.3. The result remains understandable, but it sounds more synthetic than the higher-step narration. That is a useful trade: choose a low-step setup when an internal prototype or a disposable notification needs a quick turnaround. Do not choose it as the default for brand voice, an emotional read, or a long recording.
Test 5: clean-reference voice cloning
Parameters: reference audio plus exact reference transcript, cfgValue=2.0, inferenceSteps=10. Run time: not recorded. Cost: not recorded.
The clone carries more of the reference character than the default outputs. The clean input and matching transcript give the model a well-defined target. This is the setting to use when a consented speaker has supplied a tidy reference clip and the job needs continuity across new, relatively plain scripts. Keep the transcript exact; an inaccurate transcript changes the conditioning rather than merely serving as a label.
Test 6: token-heavy cloning
Parameters: token-heavy reference audio plus exact reference transcript, cfgValue=2.0, inferenceSteps=10. Run time: not recorded. Cost: not recorded.
This sample exposes the boundary. The cloned style survives better than an unconditioned read would, but code-like language remains awkward. URLs, underscores, punctuation names, and hashes are not natural prose. The fix belongs upstream: give callers a short URL, spell an identifier intentionally, or send the precise token by SMS or email. For recognition-side handling of recordings, see Realtime Speech to Text: 3 Smart Wiro Models in 2026 and Chatterbox Multilingual: Voice Cloning in 23 Languages.
Which VoxCPM setup to pick
| Need | Pick from this test | Why |
|---|---|---|
| Short operational messages | 10 steps, cfg 2.0 | Tests 1 and 3 keep compact instructions clear without the quality-first 20-step setting. |
| Short narration | 20 steps, cfg 2.0 | Test 2 gives the most natural pacing in this set. |
| Fast prototype output | 5 steps, cfg 2.3 | Test 4 trades polish for a lower-step workflow. |
| Consented speaker continuity | Reference audio plus exact transcript, 10 steps | Test 5 shows the clearest cloning case. |
| URLs and code strings | Normalize text first | Test 6 shows why cloning alone does not solve token pronunciation. |
Verdict
VoxCPM is most convincing here as a practical voice layer for short, well-written business speech and for consented cloning from a clean reference. Ten steps at cfg 2.0 is the useful middle ground in this set. Move to 20 steps when the delivery matters more than responsiveness. Drop to five only when the synthetic edge is acceptable. Above all, make the text speakable before sending it: a good reference clip cannot turn a raw URL into natural dialogue.