{"id":1235,"date":"2026-03-01T21:39:14","date_gmt":"2026-03-01T21:39:14","guid":{"rendered":"https:\/\/wiro.ai\/blog\/?p=1235"},"modified":"2026-09-27T20:55:57","modified_gmt":"2026-09-27T20:55:57","slug":"qwen3-asr-1-7b-speech-to-text-in-6-audio-tests","status":"publish","type":"post","link":"https:\/\/wiro.ai\/blog\/qwen3-asr-1-7b-speech-to-text-in-6-audio-tests\/","title":{"rendered":"Qwen3-ASR-1.7B: Speech-to-Text in 6 Audio Tests"},"content":{"rendered":"<p><strong>Qwen3-ASR-1.7B speech-to-text<\/strong> was tested with six short clips designed to expose the details that usually matter after an API response lands: digits, punctuation, spelled-out codes, URLs, noise, speed, and language choice. This was not a leaderboard test. It was a practical check of whether a transcript can move straight into an order workflow, support tool, or notes field without someone correcting the risky parts.<\/p>\n<p>Run the model here: <a href=\"https:\/\/wiro.ai\/models\/qwen\/qwen3-asr-1-7b\">Qwen3-ASR-1.7B on Wiro<\/a>. The source clips came from <a href=\"https:\/\/wiro.ai\/models\/resemble-ai\/chatterbox-multilingual\">Chatterbox Multilingual on Wiro<\/a>, which kept the scripts and synthetic voice setup consistent across the set.<\/p>\n<h2>What this Qwen3-ASR-1.7B speech-to-text test checks<\/h2>\n<p>A transcript can look good while still breaking a real task. An order number changed by one digit, a damaged postal code, or a missing underscore in an error code can force a manual follow-up. The six clips therefore separate ordinary spoken sentences from content where every character matters.<\/p>\n<nav><strong>In this article<\/strong><\/p>\n<ul>\n<li><a href=\"#setup\">Test setup and parameters<\/a><\/li>\n<li><a href=\"#results\">What each output shows<\/a><\/li>\n<li><a href=\"#runtime\">Runtime and cost<\/a><\/li>\n<li><a href=\"#when-to-use\">When to choose each model<\/a><\/li>\n<\/ul>\n<\/nav>\n<h2 id=\"setup\">Test setup and parameters<\/h2>\n<p>The input set used four scripts: clean English shipping dictation; English with an email address, URL, underscore-delimited error code, and commit-like token; Turkish e-commerce copy; and Spanish support copy. Test 5 reuses the first English script with added white noise. Test 6 reuses the token-heavy English clip at 1.35x speed. Using the same scripts for the altered versions makes the failure pattern easier to see.<\/p>\n<p>Each clip was transcribed with Qwen3-ASR-1.7B and an explicit language selection matching the clip where applicable: English, Turkish, or Spanish. The Wiro model page lists <code>batchSize=32<\/code> and <code>newTokens=256<\/code>, and those are the settings used here. Batch size is a memory guard rather than a quality dial; smaller values can help if an environment runs out of memory. The token limit caps transcript generation, so longer recordings may need a higher value. The documented input types include WAV, M4A, MP3, OGG, OPUS, and WebM.<\/p>\n<p>Chatterbox Multilingual generated the source speech. Its documented settings were left at their defaults for this test: language set per clip, <code>exaggeration=0.5<\/code>, <code>cfg_weight=0.5<\/code>, <code>temperature=0.8<\/code>, and <code>topP=1.0<\/code>. That matters because a TTS source with extreme prosody could turn this into a test of speech synthesis rather than recognition.<\/p>\n<h2 id=\"results\">Results: what each output actually shows<\/h2>\n<table>\n<thead>\n<tr>\n<th>Test<\/th>\n<th>Condition<\/th>\n<th>Result<\/th>\n<th>Practical reading<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>1<\/td>\n<td>Clean English<\/td>\n<td>All key numbers and the tracking code landed correctly.<\/td>\n<td>Good fit for ordinary English dictation.<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>English tokens<\/td>\n<td>The domain began as \u201cyiro\u201d and the protocol lost \u201chttp\u201d.<\/td>\n<td>Review technical strings.<\/td>\n<\/tr>\n<tr>\n<td>3<\/td>\n<td>Turkish<\/td>\n<td>Amount, delivery detail, and code changed substantially.<\/td>\n<td>Do not automate this case from one pass.<\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>Spanish<\/td>\n<td>Order number, time, postal code, and trailing text were wrong.<\/td>\n<td>Needs stronger validation for this sample.<\/td>\n<\/tr>\n<tr>\n<td>5<\/td>\n<td>English plus noise<\/td>\n<td>Matched the clean shipping transcript.<\/td>\n<td>This specific noise mix did not hurt it.<\/td>\n<\/tr>\n<tr>\n<td>6<\/td>\n<td>English tokens at 1.35x<\/td>\n<td>Repeated the technical-string weaknesses.<\/td>\n<td>Speed did not solve token recognition.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<h3>Test 1: English clean dictation<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-01-en-clean.mp3\"><\/audio><figcaption>Prompt script: For the shipping audit, order 48219 shipped on February 14 at 9:05 AM. Total weight 3.7 kilograms. Tracking code Z X dash 9 1 dash Delta.<\/figcaption><\/figure>\n<p>The output preserved order 48219, the date, time, decimal weight, and <code>ZX-91-DELTA<\/code>. It also added conventional punctuation and normalized \u201ca.m.\u201d. This is the strongest result in the set because the material contains both prose and values that must stay exact.<\/p>\n<pre>For the shipping audit, order 48219 shipped on February 14 at 9:05 a.m. Total weight 3.7 kilograms. Tracking code ZX-91-DELTA.<\/pre>\n<h3>Test 2: English with URL and token-like strings<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-02-en-tokens.mp3\"><\/audio><figcaption>Prompt script: Email support plus wiro at acme dot dev. URL https colon slash slash api dot example dot com slash v1 slash run question mark mode equals fast ampersand retry equals 2. Error code E underscore C O N N underscore R E S E T. Commit seven f three a nine c one.<\/figcaption><\/figure>\n<p>The transcript gets much of the spoken structure right, including <code>api.example.com<\/code>, the path, and the error-code letters. But \u201cwiro\u201d becomes \u201cyiro\u201d, the URL starts with \u201cs\u201d instead of \u201chttps\u201d, and the spoken digit becomes \u201ctwo\u201d. That is acceptable for a searchable support note, not for copying a URL or credential-like string into a system.<\/p>\n<pre>Email support plus yiro at acme.dev. URL s colon slash slash api.example.com slash v1 slash run question mark mode equals fast ampersand retry equals two. Error code e underscore c o n n underscore r e s e t. Commit seven f three a nine c one.<\/pre>\n<h3>Test 3: Turkish<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-03-tr.mp3\"><\/audio><figcaption>Prompt script: Sepet tutari 1.249,90 TL. Kargo kodu T R dash 508 dash A B. Teslimat 3 gun icinde. Iade suresi 14 gundur.<\/figcaption><\/figure>\n<p>This is a clear failure for transactional use. The output keeps the opening \u201cSepet tutar\u0131\u201d shape, but the amount changes, the shipping code collapses, and the delivery and return statements disappear. Language selection alone did not protect the values in this clip. Use human review or another model benchmarked on the target Turkish audio before sending these fields downstream.<\/p>\n<pre>Sepet tutar\u0131 180000 komisiki 16 TL kargo kodu TR016<\/pre>\n<h3>Test 4: Spanish<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-04-es.mp3\"><\/audio><figcaption>Prompt script: El pedido numero 1740 llego el martes a las 18:30. El codigo postal es 28013. Gracias por llamar.<\/figcaption><\/figure>\n<p>The Spanish output retains the broad sentence structure and closing phrase, but it changes the order number, time, and postal code. It also adds a stray final phrase. That combination makes it unsuitable for automatic CRM updates from this clip. It can still help an agent scan the topic of a call, provided the system treats IDs and times as untrusted.<\/p>\n<pre>El pedido n\u00famero B 740 lleg\u00f3 el martes a las de 8:40. El c\u00f3digo postal es 2800 S. Gracias por llamar. Hay fiends en tu.<\/pre>\n<h3>Test 5: English with added white noise<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-05-en-noisy.mp3\"><\/audio><figcaption>Prompt script: Same as test 1, mixed with white noise.<\/figcaption><\/figure>\n<p>For this particular mix, Qwen3-ASR-1.7B returned the same useful transcript as Test 1. That is encouraging, but it does not prove broad noise robustness. White noise is only one interference type; overlapping speakers, room echo, clipping, and a distant microphone deserve separate tests.<\/p>\n<pre>For the shipping audit, order 48219 shipped on February 14 at 9:05 a.m. Total weight 3.7 kilograms. Tracking code ZX-91-DELTA.<\/pre>\n<h3>Test 6: English token clip at 1.35x speed<\/h3>\n<figure><audio controls preload=\"metadata\" src=\"https:\/\/wiro.ai\/blog\/wp-content\/uploads\/2026\/02\/qwen3-asr-test-06-en-fast.mp3\"><\/audio><figcaption>Prompt script: Same as test 2, audio sped up 1.35x.<\/figcaption><\/figure>\n<p>The fast clip stays readable, but technical strings remain the weak point. \u201cwiro\u201d shifts to \u201cwireo\u201d, the protocol is still incomplete, and everything is flattened into a run-on sentence. The fact that the failure resembles Test 2 suggests that the token format, more than the 1.35x speed-up, drives the error.<\/p>\n<pre>Email support plus wireo at acme.dev. URL s colon slash slash api.example.com slash v1 slash run question mark mode equals fast ampersand retry equals two error code e underscore c o n n underscore r e s e t commit seven f three a nine c one<\/pre>\n<h2 id=\"runtime\">Runtime and cost on Wiro<\/h2>\n<p>The Qwen3-ASR-1.7B documentation includes a task-detail example with <code>elapsedseconds=6.0000<\/code>. That is an API example, not a measured runtime for these six clips, so it should not be read as a promise. The documentation does not provide a published per-output price for this model, and no new run was made for this revision. A precise run cost should come from the task record after an actual request, not from an assumed audio duration.<\/p>\n<h2 id=\"when-to-use\">When to pick Qwen3-ASR-1.7B or Chatterbox Multilingual<\/h2>\n<p>Pick Qwen3-ASR-1.7B when the job is speech-to-text, especially clean English notes, spoken operational sentences, and searchable call summaries. Keep validation around URLs, IDs, hashes, account numbers, and other strings where one character changes the meaning. The model page exposes explicit language selection and supports many audio file formats, which makes it a sensible starting point for controlled ingestion.<\/p>\n<p>Pick Chatterbox Multilingual when the job is to create the speech itself. It is a text-to-speech and voice-cloning model, not a transcript engine. Here it provided controlled multilingual source clips; that is different from proving how Qwen3-ASR-1.7B handles a customer\u2019s phone microphone. For a production pipeline, test both sides: voice quality and language coverage at generation time, then recognition accuracy on the recordings users actually send.<\/p>\n<h2>Further reading<\/h2>\n<p>The model family is documented at <a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-ASR-1.7B\" target=\"_blank\" rel=\"noopener\">Qwen3-ASR-1.7B on Hugging Face<\/a>. The TTS source model is available in the <a href=\"https:\/\/github.com\/resemble-ai\/chatterbox\" target=\"_blank\" rel=\"noopener\">Chatterbox open-source repository<\/a>. Both links were checked successfully before publication.<\/p>\n<h2>Related Wiro tests<\/h2>\n<ul>\n<li><a href=\"https:\/\/wiro.ai\/blog\/realtime-speech-to-text-wiro-2026\/\">Realtime Speech to Text: 3 Smart Wiro Models in 2026<\/a><\/li>\n<li><a href=\"https:\/\/wiro.ai\/blog\/cohere-transcribe-5-speech-to-text-tests-en-es-fr\/\">Cohere Transcribe: 5 Speech-to-Text Tests<\/a><\/li>\n<li><a href=\"https:\/\/wiro.ai\/blog\/chatterbox-multilingual-voice-cloning-in-23-languages\/\">Chatterbox Multilingual: Voice Cloning in 23 Languages<\/a><\/li>\n<\/ul>\n<h2>Verdict<\/h2>\n<p>Qwen3-ASR-1.7B handled the clean English logistics clip well and held that result on the white-noise variant. It was less dependable once the audio carried URLs, technical tokens, or the Turkish and Spanish transactional examples. Use it for text-first English transcription with a review layer for exact strings. Do not treat these six clips as evidence that it can safely write multilingual order data without checks.<\/p>\n<p><a href=\"https:\/\/wiro.ai\/models\/qwen\/qwen3-asr-1-7b\">Try Qwen3-ASR-1.7B on Wiro<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Qwen3-ASR-1.7B speech-to-text was tested with six short clips designed to expose the details that usually matter after an API response lands: digits,&hellip;<\/p>\n","protected":false},"author":4,"featured_media":1234,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[52],"tags":[101,94,100,63],"class_list":["post-1235","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-model-reviews","tag-asr","tag-audio","tag-qwen","tag-speech-to-text"],"_links":{"self":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1235","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/comments?post=1235"}],"version-history":[{"count":2,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1235\/revisions"}],"predecessor-version":[{"id":4273,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/posts\/1235\/revisions\/4273"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media\/1234"}],"wp:attachment":[{"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/media?parent=1235"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/categories?post=1235"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wiro.ai\/blog\/wp-json\/wp\/v2\/tags?post=1235"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}