Try MiniMax H3 (Text-to-Video) (Image-to-Video) from MiniMax →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Audio & SpeechActive

nineninesix / kani-tts-2-en

kani-tts-2-en

bynineninesix

Generates natural-sounding speech from text with support for multi-speaker voice cloning and fast inference capabilities.

Text to SpeechVoice CloneFast Inference
Model ID
kani-tts-2-en
Provider
nineninesix
Updated
1776068638
kani-tts-2-en
0
Comments
Average rating : 0 (0 users)
Providernineninesix
Modelkani-tts-2-en
Text to SpeechVoice CloneFast Inference
wiro playground—nineninesix/kani-tts-2-en
Reset to defaults

Text to synthesize into speech.

Language/accent tag for multi-lingual models (e.g. en_us, en_nyork).

1 / 1
Maximum 1 audio allowed
nineninesix-kani-tts-2-input.mp3
URL

Reference audio file for voice cloning (10-20s recommended, any format); takes priority over speaker_embedding.

Sample outputs
nineninesix-kani-tts-2-en-sample-1.mp3
Updated 1776068638
## Overview Kani-TTS-2 is a text-to-speech model designed for fast inference and multi-speaker voice cloning. It generates natural-sounding speech from input text and supports reference audio for voice cloning. ## What you can build - Multi-character dialogue audio for games and animations - Voice-over content with personalized speaker voices - Automated podcast generation with custom voices - Interactive voice assistants with distinct speaker identities ## Inputs - **Dialogue**: Text to generate audio from (with speaker tags like [S1], [S2]) - **Speaker Reference Audios**: Optional reference audio files for voice cloning (up to 5 speakers) - **Speaker Text Prompts**: Optional text prompts corresponding to reference audios - **Text Normalize**: Option to normalize input text for better pronunciation ## Outputs - Generated audio files in supported formats (.wav, .mp3, .m4a, .webm) - Multi-speaker audio with distinct voice characteristics ## Recommended settings - Use reference audio files for consistent voice cloning - Enable text normalization for improved pronunciation - Test with short dialogue samples before full production ## Limitations - Voice cloning requires high-quality reference audio for best results - Performance may vary with longer dialogue inputs - Limited support for non-English languages ## Safety & compliance - Ensure all input text and audio references comply with copyright laws - Use generated audio only for intended purposes - Follow platform guidelines for audio content distribution

API quick start

Run kani-tts-2-en with a single API call.

POST https://api.wiro.ai/v1/Run/nineninesix/kani-tts-2-en
{
  "prompt": "Grounded in evolutionary biology and psyc…",
  "languageTag": "en_us",
  "inputAudio": "https://your-cdn.com/input.mp3",
  "inputEmbeddingSpeaker": "https://your-cdn.com/input.png"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion