Try Seedance 2.5 Uncensored Video from ByteDance →
Models
Agents
Workflows
Studio
PricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
Generative Media AgentCreate and edit media by chattingWorkflow AgentBuild visual workflows with Agent
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

Qwen / Qwen3-TTS-12Hz-1.7B

Qwen3-TTS-12Hz-1.7B

byqwen

A fast inference text-to-speech model optimized for real-time audio generation with multi-language support.

Text to SpeechFast Inference
Model ID
Qwen3-TTS-12Hz-1.7B
Provider
qwen
Updated
1770730291
Qwen3-TTS-12Hz-1.7B
0
Comments
Average rating : 0 (0 users)
Providerqwen
ModelQwen3-TTS-12Hz-1.7B
Text to SpeechFast Inference
wiro playground—qwen/Qwen3-TTS-12Hz-1.7B
Reset to defaults

The prompt to generate audio from.

Emotion of speaker.

Sample outputs
qwen-qwen3-tts-12hz-1-7b-sample-1.mp3
qwen-qwen3-tts-12hz-1-7b-sample-3.mp3
Updated 1770730291

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text. Key features:

  • Powerful Speech Representation: Powered by the self-developed Qwen3-TTS-Tokenizer-12Hz, it achieves efficient acoustic compression and high-dimensional semantic modeling of speech signals. It fully preserves paralinguistic information and acoustic environmental features, enabling high-speed, high-fidelity speech reconstruction through a lightweight non-DiT architecture.

  • Universal End-to-End Architecture: Utilizing a discrete multi-codebook LM architecture, it realizes full-information end-to-end speech modeling. This completely bypasses the information bottlenecks and cascading errors inherent in traditional LM+DiT schemes, significantly enhancing the model’s versatility, generation efficiency, and performance ceiling.

  • Extreme Low-Latency Streaming Generation: Based on the innovative Dual-Track hybrid streaming generation architecture, a single model supports both streaming and non-streaming generation. It can output the first audio packet immediately after a single character is input, with end-to-end synthesis latency as low as 97ms, meeting the rigorous demands of real-time interactive scenarios.

  • Intelligent Text Understanding and Voice Control: Supports speech generation driven by natural language instructions, allowing for flexible control over multi-dimensional acoustic attributes such as timbre, emotion, and prosody. By deeply integrating text semantic understanding, the model adaptively adjusts tone, rhythm, and emotional expression, achieving lifelike “what you imagine is what you hear” output.

Speaker

Voice Description

Native language

Vivian

Bright, slightly edgy young female voice.

Chinese

Serena

Warm, gentle young female voice.

Chinese

Uncle_Fu

Seasoned male voice with a low, mellow timbre.

Chinese

Dylan

Youthful Beijing male voice with a clear, natural timbre.

Chinese (Beijing Dialect)

Eric

Lively Chengdu male voice with a slightly husky brightness.

Chinese (Sichuan Dialect)

Ryan

Dynamic male voice with strong rhythmic drive.

English

Aiden

Sunny American male voice with a clear midrange.

English

Ono_Anna

Playful Japanese female voice with a light, nimble timbre.

Japanese

Sohee

Warm Korean female voice with rich emotion.

Korean

Example prompts

Great starting points for Qwen3-TTS-12Hz-1.7B.

The old lighthouse blinked to life after years of silence. A fisherman on the shore noticed it and felt a strange sense of hope. No one remembered turning it on, yet its light guided lost boats safely home. By morning, the lighthouse was dark again, as if nothing had happened.Audio & Speech
Lena found a handwritten letter inside a book she bought from a second-hand shop. The letter was dated twenty years ago and described her exact life. Shaken, she searched for the author’s name but found nothing. That night, she noticed a new blank page waiting at the end of the book.Audio & Speech

API quick start

Run Qwen3-TTS-12Hz-1.7B with a single API call.

POST https://api.wiro.ai/v1/Run/Qwen/Qwen3-TTS-12Hz-1.7B
{
  "prompt": "A small robot learned to paint by watchin…",
  "instruction": "Excited",
  "language": "Auto",
  "speaker": "Aiden"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion