WhisperSpeechBSD-4-ClauseOct 2023

WhisperX

WhisperX wraps Whisper Large v3 with faster-whisper batched inference, wav2vec2 forced alignment for word-level timestamps, and pyannote diarization for speaker identification — achieving up to 70x real-time on an RTX 4090. It's the recommended pipeline for subtitle generation, podcast transcription, and meeting recording analysis. Trade-off: more dependencies and a slightly more complex install than vanilla whisper.cpp.

Parameters
1.55B
dense
Context
448
tokens
Min VRAM (Q4)
0.9 GB
weights only
Run locally?
YES
fits ≤48 GB GPU
Quick answer

What hardware do I need to run WhisperX?

WhisperX needs at minimum 0.9 GB of VRAM at Q4_K_M quantization (3.4 GB at FP16). The cheapest GPU that comfortably fits with KV-cache headroom is the Intel Arc A380 6GB (6 GB VRAM, $139 MSRP). Measured throughput hits 210 x realtime on NVIDIA B200 192GB. This model fits a single consumer GPU under 48 GB, so a one-card build works.

Source: MyAIHardware model card: WhisperX (Whisper, 1.55B params)As of 2023-10-09

TL;DR, what to buy

Recommended GPU
Intel Arc A380 6GB
6 GB VRAM · $139 MSRP
Min VRAM at Q4_K_M
0.9 GB
+ ~20-30% headroom for KV cache
Best measured speed
210x RT
on NVIDIA B200 192GB

VRAM requirements by quantization

Weights only. Add ~20-30% for KV cache at typical context lengths.

QuantVRAMQualityNotes
FP163.4 GBReferenceTraining-precision weights
Q8_01.7 GBNear-lossless8-bit, ~0.1% perplexity hit
Q6_K1.4 GBVery high6-bit, near Q8 quality
Q5_K_M1.2 GBHighStrong middle ground
Q4_K_M0.9 GBBalanced (recommended)Default for local deployments
Q4_01.0 GBLegacy 4-bitOlder GGUF, kept for compatibility
Q3_K_M0.7 GBLossyWhen VRAM is very tight
Q2_K0.5 GBExtremeRescue option, quality degrades visibly
Open the VRAM calculator with KV cache + batch size

GPUs that fit this model

Filtered from MyAIHardware's GPU database. FP16 needs 3.4 GB; Q4_K_M needs 0.9 GB.

At full precision (FP16)

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269

At Q4_K_M quantization

  • Intel Arc A380 6GB6 GB · $139
  • NVIDIA GeForce RTX 3050 (6 GB)6 GB · $169
  • Intel Arc A580 8GB8 GB · $179
  • Intel Arc A750 8GB8 GB · $199
  • Intel Arc B57010 GB · $219
  • Intel Arc A580 12GB (variant) [VERIFY]12 GB · $219
  • NVIDIA GeForce GTX 1660 SUPER6 GB · $229
  • Intel Arc B58012 GB · $249
  • NVIDIA GeForce RTX 3050 (8 GB)8 GB · $249
  • AMD RX 76008 GB · $269
  • NVIDIA RTX 5060 8GB8 GB · $299
  • NVIDIA RTX 4060 8GB8 GB · $299

Community benchmarks

33 measurement(s) for this model from MyAIHardware's benchmark database.

DeviceSpeedQuantContextPower
NVIDIA B200 192GB210x RTFP1601000W
NVIDIA H100 SXM5 80GB145x RTFP160700W
NVIDIA GeForce RTX 5090 32GB128x RTFP160575W
NVIDIA GeForce RTX 5090 32GB105x RTFP160575W
NVIDIA GeForce RTX 5080 16GB95x RTFP160360W
AMD Instinct MI300X 192GB95x RTFP160750W
NVIDIA GeForce RTX 4080 Super 16GB82x RTFP160320W
NVIDIA GeForce RTX 4090 24GB78x RTFP160450W
NVIDIA GeForce RTX 5090 32GB78x RTFP160575W
NVIDIA GeForce RTX 4080 Super 16GB64x RTFP160320W
NVIDIA GeForce RTX 3090 24GB56x RTFP160350W
AMD Radeon RX 7900 XTX 24GB48x RTFP160355W
NVIDIA GeForce RTX 4080 Super 16GB48x RTFP160320W
NVIDIA GeForce RTX 4070 12GB42x RTFP160200W
Apple M4 Max (40c GPU, 128GB)32x RTFP16070W
NVIDIA GeForce RTX 3090 24GB32x RTFP160350W
NVIDIA GeForce RTX 4060 Ti 16GB30x RTFP160165W
Intel Arc B580 12GB28x RTINT80190W
Apple M3 Max (40c GPU, 64GB)26x RTFP16065W
Apple M4 Pro (20c GPU, 48GB)22x RTFP16050W

Where to download

Official weights + popular runtime tags.

Hugging Face
Official weights
Open on Hugging Face

Related tutorials

Step-by-step guides that use this model.

Compare with other open models

Models in a similar size or capability class.

Cite this model card

Use these in papers, blog posts, or internal docs.

BibTeX
@misc{whisperx-2023,
  title={WhisperX},
  author={OpenAI},
  year={2023},
  url={https://github.com/m-bain/whisperX}
}
APA
OpenAI (2023). WhisperX [Model card]. Hugging Face. https://github.com/m-bain/whisperX
Plain text
WhisperX (Whisper, OpenAI, 2023). Available at https://github.com/m-bain/whisperX.
Browse the full model catalog
30 deep-dive pages covering every major open model.
All 30 models