Skip to content
Model Roundups

Speech-to-Text APIs in 2026: One Audio Clip, Two Modern Transcribers

Speech-to-text APIs: one clip, two different failure patterns

This speech-to-text APIs comparison puts the same short English MP3 through Qwen3-ASR-1.7B on Wiro and ElevenLabs Speech-to-Text on Wiro. The goal was not to declare a winner from one tiny clip. It was to see what happens when a transcript contains the details that often cause trouble in product copy, captions, support notes, and model demos: decimal numbers, comma-separated numbers, hyphenated terms, a brand name, and four unfamiliar proper nouns.

What this test set out to check

The source sentence is short on purpose: “Hi. This is a 2026 speech to text benchmark on Wiro. It includes numbers like 3.5, 720p, and 1,024. Proper nouns: Kling, Seedance, PixVerse, Hailuo. End.” It asks each transcriber to preserve more than ordinary English words. A useful result needs to hear 2026 rather than another number, keep 3.5 intact, distinguish 1,024 from 1024 where formatting matters, and avoid turning product names into plausible-sounding English words.

This is a diagnostic clip, not a word-error-rate benchmark. It has one speaker, no overlapping speech, and no independently measured noise level. That means it cannot establish overall accuracy, multilingual quality, diarization quality, timestamp quality, or performance on a two-hour recording. It can show how each output handled this exact cluster of risky tokens.

Audio sample and parameters

The same MP3 was used for both recorded outputs. Qwen3-ASR-1.7B was sent with language set to English, max inference batch size set to 32, and max new tokens set to 256. The language choice removes automatic language detection as a variable. The token ceiling is far above this clip’s needs, so it should not truncate the transcript. Qwen’s Wiro documentation lists accepted audio formats including WAV, M4A, MP3, OGG, OPUS, and WebM, plus a language selector with English and many other choices.

The ElevenLabs Wiro model exposes an audio URL input. No language, vocabulary, punctuation, speaker, or token-limit control is documented for this model page, so this comparison keeps the setup to the shared file rather than implying unavailable tuning. That difference matters in production: Qwen offers visible controls for a batch workflow, while the ElevenLabs listing keeps the request surface small.

For background on the underlying models, see the Qwen3-ASR-1.7B model card on Hugging Face and ElevenLabs Speech-to-Text documentation. Both links returned HTTP 200 when checked.

Qwen3-ASR-1.7B: what the output shows

Qwen3-ASR-1.7B transcript output
Hosted model output: Qwen3-ASR-1.7B transcript from the test clip.

Qwen returned the decimal 3.5 and 720p correctly. It also retained the final number as 1024, although the spoken reference presents it as 1,024. That is a formatting difference, not a change in value, but it still matters if a transcript will feed a price list, inventory record, or technical documentation.

The larger issue is the name-heavy tail. “Wiro” became “Weiro.” Kling became “hilling,” Seedance became “students,” and PixVerse became “pigs verse.” Hailuo survived in recognizable form. The opening phrase also drifted from “2026 speech to text” to “20th round six-page-to-text.” The transcript remains readable as a rough note, but it should not be published or pushed straight into a database without review.

The earlier recorded Wiro task took 45 seconds end to end. The available Wiro documentation does not publish a fixed per-output cost for this listing, and no completed new run was available to measure a current charge. Cost is therefore left unreported rather than guessed. Qwen’s public model card describes support for 30 languages and 22 Chinese dialects, language identification, offline and streaming inference, and long-audio transcription. Those wider claims are useful context, but this clip only tests a short English batch request.

ElevenLabs Speech-to-Text: what the output shows

ElevenLabs Speech-to-Text transcript output
Hosted model output: ElevenLabs Speech-to-Text transcript from the same test clip.

The ElevenLabs output preserved 3.5, 720p, and 1024. Its punctuation also separated the proper-name portion more clearly than the Qwen result. But the core identity errors remain. “Wiro” became “Weiro,” Kling became “hyelin,” Seedance became “sedents,” PixVerse became “pixvers,” and Hailuo became “hyluo.” The first phrase again transformed 2026 into a different string.

The recorded Wiro task completed in 4 seconds, much faster than the 45-second Qwen task on this particular clip. That is an observation from one prior run, not a latency guarantee. The Wiro model page documents the audio URL input but does not expose a per-output price, so no cost is claimed here. ElevenLabs’ official guide distinguishes batch transcription from realtime speech-to-text flows; this Wiro test belongs to the batch-style use case.

Comparison at a glance

Model Recorded run time Parameters used What held up Main caution
Qwen3-ASR-1.7B 45 seconds English; batch size 32; max new tokens 256 3.5 and 720p Several names and the opening year/phrase drifted
ElevenLabs Speech-to-Text 4 seconds Shared audio URL 3.5, 720p, and clearer separation of the list Every proper noun changed and the opening phrase drifted

Neither listing supplied a documented fixed per-output cost, so the comparison deliberately omits pricing. The two hosted transcript files above preserve the actual outputs used for the reading, rather than polishing them into a cleaner result.

Which speech-to-text API should you pick?

Pick ElevenLabs Speech-to-Text when low turnaround matters for short, simple audio and a minimal audio-URL request is attractive. The four-second recorded result makes it the practical first choice for quick internal notes, short clips, or a workflow where a person already checks the transcript before it ships.

Pick Qwen3-ASR-1.7B when the visible language selector and batch/token controls matter, or when the project needs a path toward broader multilingual and long-audio work. Its public documentation makes those capabilities explicit. The 45-second result here is slower, but this one sample does not reveal throughput under load or long-form behavior.

For either model, add a glossary or post-processing rule before publishing text that contains product names, customer names, tracking codes, SKUs, legal terms, or exact numerals. This clip shows why: the ordinary sentence is understandable, while the details a business may care about most are the first to drift.

More speech-to-text tests on Wiro

Run the same audio on Qwen3-ASR-1.7B and ElevenLabs Speech-to-Text to test the names, languages, and audio conditions that matter in your own workflow.