{"id":1248,"date":"2026-03-02T23:00:08","date_gmt":"2026-03-02T23:00:08","guid":{"rendered":"https:\/\/wiro.ai\/blog\/?p=1248"},"modified":"2026-09-27T19:39:16","modified_gmt":"2026-09-27T19:39:16","slug":"voxcpm-voice-cloning-and-tts-in-6-tests","status":"publish","type":"post","link":"https:\/\/wiro.ai\/blog\/voxcpm-voice-cloning-and-tts-in-6-tests\/","title":{"rendered":"VoxCPM: Voice Cloning and TTS in 6 Tests"},"content":{"rendered":"<p><strong>VoxCPM voice cloning and TTS<\/strong> were tested here with six deliberately practical scripts: order data, narration, support copy, a short ad read, and two cloned-voice cases. The goal was not to crown a universal voice model. It was to see where this Wiro deployment stays intelligible, where extra diffusion steps help, and where text should be normalized before it reaches speech.<\/p>\n<p>The six audio players below are the original Wiro-hosted outputs. No output was regenerated for this update. That matters for the timing and cost notes: the original run records are not attached to these media files, and the Wiro model documentation does not publish a price or a measured run time for this endpoint. Those fields are therefore marked as not recorded rather than filled with estimates.<\/p>\n<p>Model page: <a href=\"https:\/\/wiro.ai\/models\/openbmb\/voxcpm\">VoxCPM on Wiro<\/a>. Wiro exposes four relevant inputs: <code>prompt<\/code>, <code>cfgValue<\/code>, <code>inferenceSteps<\/code>, and the optional pair <code>inputAudio<\/code> plus <code>referencePrompt<\/code>. The reference file and its transcript must be supplied together, or both left empty. The hosted documentation describes <code>cfgValue<\/code> as guidance for prompt adherence: higher can adhere more closely, but may sound worse. Higher <code>inferenceSteps<\/code> can improve results at the cost of speed.<\/p>\n<h2>What the VoxCPM voice cloning and TTS test set checked<\/h2>\n<p>The first four samples test default synthesis, not cloning. They move from structured commercial text to more natural prose, then deliberately reduce the step count for a speed-versus-smoothness check. Tests five and six add a source speaker and an exact reference transcript. Together, they test whether a clean reference helps a new script retain a speaker-like character, and whether that advantage survives text packed with URLs, punctuation names, and code-like tokens.<\/p>\n<ul>\n<li><strong>Default guidance:<\/strong> <code>cfgValue=2.0<\/code> in five samples; <code>2.3<\/code> in the fast ad read.<\/li>\n<li><strong>Steps:<\/strong> 5, 10, or 20, depending on the test.<\/li>\n<li><strong>Voice cloning:<\/strong> tests 5 and 6 use both a reference MP3 and its transcript.<\/li>\n<li><strong>Run time and Wiro cost:<\/strong> not recorded for each existing output; the Wiro documentation does not state an endpoint price.<\/li>\n<\/ul>\n<p>Upstream VoxCPM materials describe a tokenizer-free, diffusion-autoregressive TTS approach. The project&#8217;s <a href=\"https:\/\/github.com\/OpenBMB\/VoxCPM\" target=\"_blank\" rel=\"noopener\">official GitHub repository<\/a> and its <a href=\"https:\/\/huggingface.co\/openbmb\/VoxCPM2\" target=\"_blank\" rel=\"noopener\">Hugging Face model card<\/a> are useful technical context, but they describe the broader and newer VoxCPM line. They should not be treated as a benchmark or pricing claim for this specific Wiro-hosted endpoint.<\/p>\n<h2>Test 1: numbers, currency, and a tracking code<\/h2>\n<p><strong>Parameters:<\/strong> <code>cfgValue=2.0<\/code>, <code>inferenceSteps=10<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-01-order-confirmation.mp3\"><\/audio><figcaption>Prompt: Your order 51723 is confirmed. Total: 1249.90 TL. Delivery window: 2 to 3 business days. Tracking: TR-508-AB.<\/figcaption><\/figure>\n<p>This is the operational test. The output handles a short confirmation with several categories that often break TTS: a five-digit number, a decimal amount, a range, and a hyphenated code. In the sample, the speech stays easy to follow and the pauses work for a transactional message. It is a good fit for concise order updates, IVR fragments, and status notifications. It is not proof that every identifier will be pronounced correctly, so production flows should still normalize account numbers, decimals, and abbreviations in the text layer.<\/p>\n<h2>Test 2: calm narration<\/h2>\n<p><strong>Parameters:<\/strong> <code>cfgValue=2.0<\/code>, <code>inferenceSteps=20<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-02-narration.mp3\"><\/audio><figcaption>Prompt: The street is quiet after midnight. A tram passes and the sound fades into the rain. The cafe sign flickers once, then holds steady.<\/figcaption><\/figure>\n<p>This sample asks for phrasing rather than data accuracy. The longer sentences retain a measured pace and the pauses land at sentence boundaries. Compared with the 10-step business samples, the 20-step setting is the sensible quality-first choice for a short narration where waiting a little longer is acceptable. The sample alone cannot establish a numeric latency gain, because its original job timing was not retained.<\/p>\n<h2>Test 3: support message<\/h2>\n<p><strong>Parameters:<\/strong> <code>cfgValue=2.0<\/code>, <code>inferenceSteps=10<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-03-support-message.mp3\"><\/audio><figcaption>Prompt: Hi. This is support. The reset link expires in 15 minutes. Do not share the code. If this was not requested, ignore this message.<\/figcaption><\/figure>\n<p>Short clauses make this output work. Each instruction has room to land, and the tone remains consistent across the warning and the fallback instruction. This is the strongest pattern for support automation: write plain clauses, put the critical action in its own sentence, and avoid cramming a URL or an unspoken symbol into the same line. For a phone agent, pair this kind of output with the related <a href=\"https:\/\/wiro.ai\/blog\/realtime-voice-conversation-wiro-2026\/\">realtime voice conversation guide<\/a>.<\/p>\n<h2>Test 4: fast ad read<\/h2>\n<p><strong>Parameters:<\/strong> <code>cfgValue=2.3<\/code>, <code>inferenceSteps=5<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-04-ad-read.mp3\"><\/audio><figcaption>Prompt: New drop. Same price. Faster shipping. Add to cart and check out in under 30 seconds.<\/figcaption><\/figure>\n<p>This is the speed stress case. Five steps are the lowest setting in this set, while guidance rises slightly to 2.3. The result remains understandable, but it sounds more synthetic than the higher-step narration. That is a useful trade: choose a low-step setup when an internal prototype or a disposable notification needs a quick turnaround. Do not choose it as the default for brand voice, an emotional read, or a long recording.<\/p>\n<h2>Test 5: clean-reference voice cloning<\/h2>\n<p><strong>Parameters:<\/strong> reference audio plus exact reference transcript, <code>cfgValue=2.0<\/code>, <code>inferenceSteps=10<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-01-en-clean.mp3\"><\/audio><figcaption>Reference transcript: For the shipping audit, order 48219 shipped on February 14 at 9:05 AM. Total weight 3.7 kilograms. Tracking code Z X dash 9 1 dash Delta.<\/figcaption><\/figure>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-05-voice-clone-clean.mp3\"><\/audio><figcaption>Clone prompt: Voice clone test. Ticket 77104 closed at 18:30. Refund amount 79.90 TL. Please reply with the last four digits of the card.<\/figcaption><\/figure>\n<p>The clone carries more of the reference character than the default outputs. The clean input and matching transcript give the model a well-defined target. This is the setting to use when a consented speaker has supplied a tidy reference clip and the job needs continuity across new, relatively plain scripts. Keep the transcript exact; an inaccurate transcript changes the conditioning rather than merely serving as a label.<\/p>\n<h2>Test 6: token-heavy cloning<\/h2>\n<p><strong>Parameters:<\/strong> token-heavy reference audio plus exact reference transcript, <code>cfgValue=2.0<\/code>, <code>inferenceSteps=10<\/code>. <strong>Run time:<\/strong> not recorded. <strong>Cost:<\/strong> not recorded.<\/p>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-02-en-tokens.mp3\"><\/audio><figcaption>Reference transcript: Email support plus wiro at acme dot dev. URL https colon slash slash api dot example dot com slash v1 slash run question mark mode equals fast ampersand retry equals 2. Error code E underscore C O N N underscore R E S E T. Commit seven f three a nine c one.<\/figcaption><\/figure>\n<figure>\n  <audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/voxcpm-06-voice-clone-tokens.mp3\"><\/audio><figcaption>Clone prompt: Second clone test. Open https colon slash slash status dot example dot com. If error code E underscore T I M E O U T appears, retry twice.<\/figcaption><\/figure>\n<p>This sample exposes the boundary. The cloned style survives better than an unconditioned read would, but code-like language remains awkward. URLs, underscores, punctuation names, and hashes are not natural prose. The fix belongs upstream: give callers a short URL, spell an identifier intentionally, or send the precise token by SMS or email. For recognition-side handling of recordings, see <a href=\"https:\/\/wiro.ai\/blog\/realtime-speech-to-text-wiro-2026\/\">Realtime Speech to Text: 3 Smart Wiro Models in 2026<\/a> and <a href=\"https:\/\/wiro.ai\/blog\/chatterbox-multilingual-voice-cloning-in-23-languages\/\">Chatterbox Multilingual: Voice Cloning in 23 Languages<\/a>.<\/p>\n<h2>Which VoxCPM setup to pick<\/h2>\n<table>\n<thead>\n<tr>\n<th>Need<\/th>\n<th>Pick from this test<\/th>\n<th>Why<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Short operational messages<\/td>\n<td>10 steps, cfg 2.0<\/td>\n<td>Tests 1 and 3 keep compact instructions clear without the quality-first 20-step setting.<\/td>\n<\/tr>\n<tr>\n<td>Short narration<\/td>\n<td>20 steps, cfg 2.0<\/td>\n<td>Test 2 gives the most natural pacing in this set.<\/td>\n<\/tr>\n<tr>\n<td>Fast prototype output<\/td>\n<td>5 steps, cfg 2.3<\/td>\n<td>Test 4 trades polish for a lower-step workflow.<\/td>\n<\/tr>\n<tr>\n<td>Consented speaker continuity<\/td>\n<td>Reference audio plus exact transcript, 10 steps<\/td>\n<td>Test 5 shows the clearest cloning case.<\/td>\n<\/tr>\n<tr>\n<td>URLs and code strings<\/td>\n<td>Normalize text first<\/td>\n<td>Test 6 shows why cloning alone does not solve token pronunciation.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h2>Verdict<\/h2>\n<p>VoxCPM is most convincing here as a practical voice layer for short, well-written business speech and for consented cloning from a clean reference. Ten steps at cfg 2.0 is the useful middle ground in this set. Move to 20 steps when the delivery matters more than responsiveness. Drop to five only when the synthetic edge is acceptable. Above all, make the text speakable before sending it: a good reference clip cannot turn a raw URL into natural dialogue.<\/p>\n<p><a href=\"https:\/\/wiro.ai\/models\/openbmb\/voxcpm\">Try VoxCPM on Wiro<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>VoxCPM voice cloning and TTS were tested here with six deliberately practical scripts: order data, narration, support copy, a short ad read,&hellip;<\/p>\n","protected":false},"author":4,"featured_media":1247,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[52],"tags":[94,105,62,68,104],"class_list":["post-1248","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-model-reviews","tag-audio","tag-openbmb","tag-text-to-speech","tag-voice-clone","tag-voxcpm"],"_links":{"self":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1248","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/comments?post=1248"}],"version-history":[{"count":2,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1248\/revisions"}],"predecessor-version":[{"id":4236,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1248\/revisions\/4236"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media\/1247"}],"wp:attachment":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media?parent=1248"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/categories?post=1248"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/tags?post=1248"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}