AUDIO & SPEECH
ASR Speed Test
15 min read
Whisper Large-v3-Turbo vs. whisper.cpp: Transcribing 100 Hours of Audio in 8 Minutes
JC
Jutt AI Engineering Lab
Principal Systems & AI Security Architect
September 2026
Jutt Cyber Tech™
Detailed Engineering Index
OpenAI's Whisper large-v3-turbo pruned the decoder down from 32 layers to just 4 layers while retaining full 32-layer encoder acoustic representations. When paired with whisper.cpp, we processed a 100-hour multi-speaker audio dataset in just 8 minutes and 14 seconds on a single RTX 4090.
Domain: #AUDIO&SPEECH #JuttCyberTech #AIInfrastructure