Skip to content
Model Reviews

Qwen3-ASR-1.7B: Speech-to-Text in 6 Audio Tests

Qwen3-ASR-1.7B speech-to-text was tested with six short clips designed to expose the details that usually matter after an API response lands: digits, punctuation, spelled-out codes, URLs, noise, speed, and language choice. This was not a leaderboard test. It was a practical check of whether a transcript can move straight into an order workflow, support tool, or notes field without someone correcting the risky parts.

Run the model here: Qwen3-ASR-1.7B on Wiro. The source clips came from Chatterbox Multilingual on Wiro, which kept the scripts and synthetic voice setup consistent across the set.

What this Qwen3-ASR-1.7B speech-to-text test checks

A transcript can look good while still breaking a real task. An order number changed by one digit, a damaged postal code, or a missing underscore in an error code can force a manual follow-up. The six clips therefore separate ordinary spoken sentences from content where every character matters.

Test setup and parameters

The input set used four scripts: clean English shipping dictation; English with an email address, URL, underscore-delimited error code, and commit-like token; Turkish e-commerce copy; and Spanish support copy. Test 5 reuses the first English script with added white noise. Test 6 reuses the token-heavy English clip at 1.35x speed. Using the same scripts for the altered versions makes the failure pattern easier to see.

Each clip was transcribed with Qwen3-ASR-1.7B and an explicit language selection matching the clip where applicable: English, Turkish, or Spanish. The Wiro model page lists batchSize=32 and newTokens=256, and those are the settings used here. Batch size is a memory guard rather than a quality dial; smaller values can help if an environment runs out of memory. The token limit caps transcript generation, so longer recordings may need a higher value. The documented input types include WAV, M4A, MP3, OGG, OPUS, and WebM.

Chatterbox Multilingual generated the source speech. Its documented settings were left at their defaults for this test: language set per clip, exaggeration=0.5, cfg_weight=0.5, temperature=0.8, and topP=1.0. That matters because a TTS source with extreme prosody could turn this into a test of speech synthesis rather than recognition.

Results: what each output actually shows

Test Condition Result Practical reading
1 Clean English All key numbers and the tracking code landed correctly. Good fit for ordinary English dictation.
2 English tokens The domain began as “yiro” and the protocol lost “http”. Review technical strings.
3 Turkish Amount, delivery detail, and code changed substantially. Do not automate this case from one pass.
4 Spanish Order number, time, postal code, and trailing text were wrong. Needs stronger validation for this sample.
5 English plus noise Matched the clean shipping transcript. This specific noise mix did not hurt it.
6 English tokens at 1.35x Repeated the technical-string weaknesses. Speed did not solve token recognition.

Test 1: English clean dictation

Prompt script: For the shipping audit, order 48219 shipped on February 14 at 9:05 AM. Total weight 3.7 kilograms. Tracking code Z X dash 9 1 dash Delta.

The output preserved order 48219, the date, time, decimal weight, and ZX-91-DELTA. It also added conventional punctuation and normalized “a.m.”. This is the strongest result in the set because the material contains both prose and values that must stay exact.

For the shipping audit, order 48219 shipped on February 14 at 9:05 a.m. Total weight 3.7 kilograms. Tracking code ZX-91-DELTA.

Test 2: English with URL and token-like strings

Prompt script: Email support plus wiro at acme dot dev. URL https colon slash slash api dot example dot com slash v1 slash run question mark mode equals fast ampersand retry equals 2. Error code E underscore C O N N underscore R E S E T. Commit seven f three a nine c one.

The transcript gets much of the spoken structure right, including api.example.com, the path, and the error-code letters. But “wiro” becomes “yiro”, the URL starts with “s” instead of “https”, and the spoken digit becomes “two”. That is acceptable for a searchable support note, not for copying a URL or credential-like string into a system.

Email support plus yiro at acme.dev. URL s colon slash slash api.example.com slash v1 slash run question mark mode equals fast ampersand retry equals two. Error code e underscore c o n n underscore r e s e t. Commit seven f three a nine c one.

Test 3: Turkish

Prompt script: Sepet tutari 1.249,90 TL. Kargo kodu T R dash 508 dash A B. Teslimat 3 gun icinde. Iade suresi 14 gundur.

This is a clear failure for transactional use. The output keeps the opening “Sepet tutarı” shape, but the amount changes, the shipping code collapses, and the delivery and return statements disappear. Language selection alone did not protect the values in this clip. Use human review or another model benchmarked on the target Turkish audio before sending these fields downstream.

Sepet tutarı 180000 komisiki 16 TL kargo kodu TR016

Test 4: Spanish

Prompt script: El pedido numero 1740 llego el martes a las 18:30. El codigo postal es 28013. Gracias por llamar.

The Spanish output retains the broad sentence structure and closing phrase, but it changes the order number, time, and postal code. It also adds a stray final phrase. That combination makes it unsuitable for automatic CRM updates from this clip. It can still help an agent scan the topic of a call, provided the system treats IDs and times as untrusted.

El pedido número B 740 llegó el martes a las de 8:40. El código postal es 2800 S. Gracias por llamar. Hay fiends en tu.

Test 5: English with added white noise

Prompt script: Same as test 1, mixed with white noise.

For this particular mix, Qwen3-ASR-1.7B returned the same useful transcript as Test 1. That is encouraging, but it does not prove broad noise robustness. White noise is only one interference type; overlapping speakers, room echo, clipping, and a distant microphone deserve separate tests.

For the shipping audit, order 48219 shipped on February 14 at 9:05 a.m. Total weight 3.7 kilograms. Tracking code ZX-91-DELTA.

Test 6: English token clip at 1.35x speed

Prompt script: Same as test 2, audio sped up 1.35x.

The fast clip stays readable, but technical strings remain the weak point. “wiro” shifts to “wireo”, the protocol is still incomplete, and everything is flattened into a run-on sentence. The fact that the failure resembles Test 2 suggests that the token format, more than the 1.35x speed-up, drives the error.

Email support plus wireo at acme.dev. URL s colon slash slash api.example.com slash v1 slash run question mark mode equals fast ampersand retry equals two error code e underscore c o n n underscore r e s e t commit seven f three a nine c one

Runtime and cost on Wiro

The Qwen3-ASR-1.7B documentation includes a task-detail example with elapsedseconds=6.0000. That is an API example, not a measured runtime for these six clips, so it should not be read as a promise. The documentation does not provide a published per-output price for this model, and no new run was made for this revision. A precise run cost should come from the task record after an actual request, not from an assumed audio duration.

When to pick Qwen3-ASR-1.7B or Chatterbox Multilingual

Pick Qwen3-ASR-1.7B when the job is speech-to-text, especially clean English notes, spoken operational sentences, and searchable call summaries. Keep validation around URLs, IDs, hashes, account numbers, and other strings where one character changes the meaning. The model page exposes explicit language selection and supports many audio file formats, which makes it a sensible starting point for controlled ingestion.

Pick Chatterbox Multilingual when the job is to create the speech itself. It is a text-to-speech and voice-cloning model, not a transcript engine. Here it provided controlled multilingual source clips; that is different from proving how Qwen3-ASR-1.7B handles a customer’s phone microphone. For a production pipeline, test both sides: voice quality and language coverage at generation time, then recognition accuracy on the recordings users actually send.

Further reading

The model family is documented at Qwen3-ASR-1.7B on Hugging Face. The TTS source model is available in the Chatterbox open-source repository. Both links were checked successfully before publication.

Verdict

Qwen3-ASR-1.7B handled the clean English logistics clip well and held that result on the white-noise variant. It was less dependable once the audio carried URLs, technical tokens, or the Turkish and Spanish transactional examples. Use it for text-first English transcription with a review layer for exact strings. Do not treat these six clips as evidence that it can safely write multilingual order data without checks.

Try Qwen3-ASR-1.7B on Wiro.