Try Google Gemini Omni Flash Video Generator from Google →
Models
Agents
WorkflowsStudioPricingBlogDocs
ExploreDiscover models by categoryBrowse All ModelsBrowse the complete catalogSee FavoritesSign in to view saved models
OverviewThe platform at a glanceLearnSkills, knowledge, guardrailsAnatomyWhat makes agents reasonBuild Your AgentPick skills, set tier, deploy
Pre-built AgentsBrowse the catalog
Agent Usecases
Ad Campaign ManagerApp Event ManagerApp Review RepliesBarber BookingCustomer Win-BackEcommerce ListingsRestaurant Reviews
Sign InStart Building

Task History

Click to see output list

No tasks yet

Go to Models
Explore models/
Music GenerationActive

stabilityai / stable-audio-3-small-sfx

stable-audio-3-small-sfx

bystabilityai

Stable Audio 3 Small SFX by Stability AI generates stereo sound effects from text prompts, up to 120 seconds, for games, video, and product UX.

Text to MusicFast Inference
Model ID
stable-audio-3-small-sfx
Provider
stabilityai
Updated
1781090277
stable-audio-3-small-sfx
1
Comments
Average rating : 0 (0 users)
Providerstabilityai
Modelstable-audio-3-small-sfx
Text to MusicFast Inference
wiro playground—stabilityai/stable-audio-3-small-sfx
Reset to defaults

Prompt.

Music duration.

Sample outputs
stabilityai-stable-audio-3-small-sfx-sample-1.mp3
stabilityai-stable-audio-3-small-sfx-sample-2.mp3
Updated 1781090277
## Overview Stable Audio 3 Small SFX is a text-to-audio model by Stability AI. It uses latent diffusion to denoise audio in a compressed latent space. A semantic-acoustic autoencoder then decodes those latents into a stereo waveform. It’s built for sound effects, so you can design foley, hits, and ambiences from text. ## What you can build - Foley packs for games, film, and animation (footsteps, cloth, impacts) - UI and product sounds (clicks, toggles, alerts, success tones) - Ambient beds (room tone, wind, rain, traffic, crowd wash) - Cinematic transitions (risers, whooshes, drops, booms) - Rapid prototypes for storyboards, animatics, and gameplay mockups ## Inputs - A written description of the sound you want, as plain English text. - A target clip length in seconds. This model supports 1 to 120 seconds. - An optional “prompt expansion” switch. It can rewrite short prompts into longer ones. - A sampling-steps control. More steps can add detail, but changes taper off. - A guidance strength control. Higher values can push closer to the prompt. ## Outputs The model returns one generated audio clip as a downloadable file. The response also includes the final prompt used and a numeric seed for that output. Under the hood, Stable Audio 3 generates 44.1 kHz stereo audio. Some deployments deliver the result as a compressed file, such as MP3. ## Recommended settings - Use 3 to 15 seconds for most one-shot SFX like hits, clicks, and whooshes. - Use 15 to 60 seconds for ambiences and looping beds. - Keep sampling steps low for iteration. The distilled checkpoints look strong at 8 steps. - Turn prompt expansion on when your prompt is short or missing key details. - Start guidance strength near 1. Raise it only when prompt adherence is weak. ## Limitations - It’s not designed for speech or voice generation. - Prompts in languages other than English can perform worse. - Maximum clip length is 120 seconds per generation. - Highly detailed prompts can still produce artifacts like warble or smeared transients. - Very dynamic mixes may need post work (trim, fade, normalize, and EQ). - Vague prompts can produce generic textures instead of clear, named events. - Low-quality or noisy source audio can hurt results when you use editing workflows. ## Safety & compliance - Follow the Stability AI Community License terms for this model. - The pipeline uses a T5Gemma text component, which has its own use restrictions. - Don’t generate content that violates laws, platform rules, or rights you don’t own. - Treat outputs as generated media. Review them before publishing or shipping. - Some deployments include an audio safety checker. It may block certain prompts.

Example prompts

Great starting points for stable-audio-3-small-sfx.

Ambient sci-fi industrial techno, driving 128 BPM. Deep, thumping subterranean kick drums and hypnotic, metallic industrial percussion loops, surrounded by a vast, shifting wall of ethereal ambient drone synthesizers and distant, echoing space-station soundscapes. Dark, immersive, clinical digital precision, expansive soundstage.Music Generation
Aggressive rage-rap instrumental, upbeat 145 BPM. Distorted, buzzing Pierre Bourne-style neon synthesizer leads playing an energetic, looping hook, driving 808 basslines that punch through the mix, and clean, fast-hitting trap drum patterns. Euphoric, stadium-status energy, modern hip-hop festival vibe, loud digital master.Music Generation
Smooth pop and contemporary R&B blend, mid-tempo 95 BPM. Warm, sultry Fender Rhodes electric piano chords, a groovy, melodic walking bass guitar line, and crisp, polished acoustic drum patterns with a subtle swing. Smooth atmospheric pads floating in the background, intimate late-night aesthetic, radio-ready commercial pop mix.Music Generation

API quick start

Run stable-audio-3-small-sfx with a single API call.

POST https://api.wiro.ai/v1/Run/stabilityai/stable-audio-3-small-sfx
{
  "prompt": "Ambient sci-fi industrial techno, driving…",
  "scale": 1.0,
  "duration": 60,
  "promptExpansion": "False"
}
View full API docs

Discover, test, and run AI models, build workflows and agents with one unified API.

All systems operational
WiroAboutBlogCareersContact
ProductModelsAgentsPricingPartnerChangelogStatusFAQ
Getting StartedIntroductionAuthenticationProjectsCode ExamplesWiro MCP ServerSelf-Hosted MCPn8n IntegrationLLMs.txt
API ReferenceModelsRun a ModelModel ParametersTasksLLM & Chat StreamingWebSocketRealtime VoiceFiles
© 2026 Wiro AI. All rights reserved.
PrivacyTermsData Deletion