Audio & SpeechActive
EMA Lightning Turkish TTS
EMA Lightning turns Turkish text into speech and can stream mono WAV audio up to 48 kHz. Adjust speaking speed and set a seed for repeatable takes.
Text to SpeechFast Inference
Model ID
ema-lightning
Provider
canberkkkkkk
Updated
1791379589
wiro playground—canberkkkkkk/ ema-lightning
Updated 1791379589
Overview
EMA Lightning is a Turkish text-to-speech (TTS) model by Canberkkkkkk. It normalizes written Turkish into spoken form, predicts letter timings, then generates compact audio latents with a small DiT model. A lightweight vocoder then decodes those latents into mono speech audio, up to 48 kHz. This setup is useful when you need clear Turkish speech from raw text, without voice recording.
What you can build
- Turkish IVR and call-center prompts, with 8 kHz or 16 kHz output
- Voiceovers for Turkish product videos and social clips
- Spoken notifications for Turkish apps (shipping updates, reminders, OTP reads)
- Screen-reader style narration for Turkish articles and documents
- Streaming speech in live apps, where audio can start before all text finishes
Inputs
- The Turkish text you want spoken, as a plain Unicode string. Numbers, dates, times, and currency are expanded and read aloud automatically.
- An output audio sample rate selection in Hz. Choose from 48,000, 24,000, 16,000, or 8,000 depending on quality needs.
- A speaking speed multiplier from 0.25 to 4.0. Use this to slow down or speed up delivery.
- An optional numeric seed (0 to 4,294,967,295) to make results repeatable for the same text and settings.
Outputs
- A mono WAV audio result at the chosen sample rate.
- Speech waveform content that stays within a normalized audio range.
- Basic generation metadata, including the final sample rate and the seed used.
Recommended settings
- Use 48,000 Hz when you want the best audio quality.
- Use 16,000 Hz or 8,000 Hz for phone lines and IVR systems.
- Keep speed at 1.0 for a natural pace.
- Set a fixed seed when you need the same take across reruns.
Limitations
- The model provides one fixed voice. It does not support voice cloning.
- It is designed for Turkish. Foreign words follow Turkish spelling and reading rules.
- It does not expose explicit emotion control.
- Text normalization can misread unusual abbreviations, rare formats, or domain-specific notation.
- Messy input text can reduce quality, like long ALL-CAPS blocks, inconsistent punctuation, or mixed-language fragments.
Safety & compliance
- Disclose synthetic speech to listeners when appropriate, especially in calls and public content.
- Do not use the voice for impersonation, fraud, or scams.
- For banking, health, legal, or emergency messaging, have a person review both text and audio.
API quick start
Run ema-lightning with a single API call.
POST https://api.wiro.ai/v1/Run/canberkkkkkk/ema-lightning
{
"prompt": "Merhaba, size nasıl yardımcı olabilirim?",
"sampleRate": 48000,
"speed": 1.0,
"seed": 0
}curl
curl -X POST "https://api.wiro.ai/v1/Run/canberkkkkkk/ema-lightning" \
-H "Content-Type: application/json" \
-H "x-api-key: YOUR_WIRO_API_KEY" \
--data-binary @- <<'JSON'
{
"prompt": "Merhaba, size nasıl yardımcı olabilirim?",
"sampleRate": 48000,
"speed": 1.0,
"seed": 0
}
JSON