WiseOne AI logo
WiseOneAI

Private AI

Explore Audio Models

Browse and discover the best AI audio models for text to speech, speech to text, and music.

Provider logo

ElevenLabs Turbo V2.5

Text to Speech

High quality with lowest latency, ideal for real-time applications. Supports 32 languages while maintaining natural voice quality.

≈ $0.090 per audio

Try ElevenLabs Turbo V2.5Details
Provider logo

GPT-4o Mini TTS

Text to Speech

Ultra-low cost text-to-speech model with voice instructions support

≈ $0.012 per audio

Try GPT-4o Mini TTSDetails

xAI TTS

Text to Speech

Affordable Runware-hosted xAI text-to-speech with five voices, inline expressive controls, and multilingual auto-detection.

≈ $0.017 per audio

Try xAI TTSDetails
Provider logo

GPT-4o Mini TTS (2025-12-15)

Text to Speech

Latest snapshot of GPT-4o Mini TTS with voice instructions support

≈ $0.012 per audio

Try GPT-4o Mini TTS (2025-12-15)Details
Provider logo

Gemini 2.5 Flash Preview TTS

Text to Speech

Google Gemini native TTS. Single and multi-speaker support via prompt.

≈ $0.045 per audio

Try Gemini 2.5 Flash Preview TTSDetails
Provider logo

ByteDance Seed Speech TTS 2.0

Text to Speech

ByteDance Seed Speech TTS 2.0 for natural multilingual speech with voice instructions and delivery controls.

≈ $0.045 per audio

Try ByteDance Seed Speech TTS 2.0Details

Mureka v9 Generate BGM

Music

Generate background music with Mureka v9. Priced per generated item.

≈ $0.045 per audio

Try Mureka v9 Generate BGMDetails
Provider logo

Kokoro 82M

Text to Speech

High-quality multilingual text-to-speech model

≈ $0.002 per audio

Try Kokoro 82MDetails
Provider logo

Whisper Large V3

Speech to Text

OpenAI's state-of-the-art speech recognition model

≈ <$0.001 per minute

Try Whisper Large V3Details
Provider logo

ElevenLabs v3

Text to Speech

High-quality text-to-speech with enhanced controls and natural voices.

≈ $0.150 per audio

Try ElevenLabs v3Details
Provider logo

Gemini 3.1 Flash TTS Preview

Text to Speech

Google Gemini 3.1 Flash text-to-speech with inline audio tag and multi-speaker prompt support.

≈ $0.090 per audio

Try Gemini 3.1 Flash TTS PreviewDetails
Provider logo

Google Lyria 3 Pro Music

Music

Google Lyria 3 Pro generates premium music clips from a text prompt, with optional image guidance, negative prompts, and seed-based repeatability.

≈ $0.080 per audio

Try Google Lyria 3 Pro MusicDetails

Mureka O2 Generate Song

Music

Generate songs with Mureka O2. Priced per generated song.

≈ $0.150 per audio

Try Mureka O2 Generate SongDetails

ACE-Step v1.5 Turbo

Music

ACE-Step v1.5 Turbo is the faster, lower-cost Runware variant for full-song generation, with broad genre coverage, improved stylistic consistency, and text-guided music creation for creator workflows.

≈ $0.006 per audio

Try ACE-Step v1.5 TurboDetails
Provider logo

Inworld TTS 1.5 Max

Text to Speech

Inworld flagship TTS model with the best balance of quality and speed, plus enhanced alignment data.

≈ $0.018 per audio

Try Inworld TTS 1.5 MaxDetails
Provider logo

MiniMax Music 02

Music

MiniMax Music 02 is a compact MoE music generator (230B params, 10B active) tuned for speedy, cost-effective song creation. Provide a creative prompt plus formatted lyrics to render polished, full-length tracks with configurable bitrate and sample rate.

≈ $0.050 per audio

Try MiniMax Music 02Details
Provider logo

Whisper-1

Speech to Text

OpenAI's original Whisper model with full format support including SRT and VTT subtitles

≈ $0.009 per minute

Try Whisper-1Details
Provider logo

MiniMax Speech 2.8 HD

Text to Speech

Studio-quality HD text-to-speech with expressive delivery, emotion control, pronunciation customization, and fine-grained audio controls for production-ready speech.

≈ $0.150 per audio

Try MiniMax Speech 2.8 HDDetails