01 /TOOLS / whisper
Whisper
shipOpenAI's open-weights speech recognition — the default transcription model.
journal entry · observed by @whysanesanders · verified 2026-08-15
The reason paid transcription became a niche: large-v3 accuracy is production-grade across dozens of languages and it runs on a laptop. Every transcription pipeline starts here.
Known limitations
- Hallucinates text on silence and noisy segments.
- No built-in speaker diarization (pair with pyannote).
- large-v3 needs a decent GPU for real-time work.
Facts
- pricing
- free
- price note
- free (open weights); hosted API $0.006/min
- free tier
- yes
- open source
- yes
- api
- yes
- self-host
- yes
- category
- audio
Models used
Receipts
Same sector — audio
Deepgram
Speech-to-text API built for developers — fast, cheap, real-time.
audio
freemium · $200 free credits; then from ~$0.0043/min (Nova models)
APIFREE TIER
ElevenLabsfeatured
The reference standard for AI voice — TTS, cloning, dubbing and voice agents.
audio
freemium · free tier (10k credits/mo); Starter $6/mo
APIFREE TIER
Murf AI
Business-focused voiceover studio for presentations and e-learning.
audio
freemium · free tier; Creator Lite $29/mo or $23/mo billed yearly
APIFREE TIER