Transcription models that run on 256 GB
39 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Models that turn recorded speech into text. They are metered by the minute or the second of audio, so cost follows the length of the recording and not the difficulty of it. What separates them is languages covered, whether they mark who is speaking, and whether they run in real time or only on a finished file — a model that is excellent on a podcast may be unusable on a live call.
- Voxtral Small 24B 2507Mistral AI24.3B≈15.8 GB at 4-bit4 also selling it hosted
- VibeVoice-ASRMicrosoft8.7B≈5.6 GB at 4-bit
- VibeVoice-ASR-Streaming-7BMicrosoft8.7B≈5.6 GB at 4-bit
- granite-speech-3.3-8bIBM Granite8.0B≈5.2 GB at 4-bit
- Voxtral-Mini-4B-Realtime-2602Mistral AI4.4B≈2.9 GB at 4-bit
- Voxtral-Mini-3B-2507Mistral AI3.0B≈2.0 GB at 4-bit2 also selling it hosted
- canary-qwen-2.5bNVIDIA2.5B≈1.6 GB at 4-bit
- seamless-m4t-v2-largeAI at Meta2.3B≈1.5 GB at 4-bit
- GLM-ASR-Nano-2512Z.ai2.3B≈1.5 GB at 4-bit
- cohere-transcribe-arabic-07-2026Cohere Labs2.1B≈1.3 GB at 4-bit
- MOSS-Transcribe-preview-2BOpenMOSS-Team2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2bIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2b-narIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- Qwen3-ASR-1.7BAlibaba1.7B≈1.1 GB at 4-bit1 also selling it hosted
- CrisperWhispernyra labs1.6B≈1.0 GB at 4-bit
- Whisper Large v3OpenAI1.5B≈1.0 GB at 4-bit2 also selling it hosted
- whisper-largeOpenAI1.5B≈1.0 GB at 4-bit
- whisper-large-v2OpenAI1.5B≈1.0 GB at 4-bit
- parakeet-rnnt-1.1bNVIDIA1.1B≈0.7 GB at 4-bit
- canary-1bNVIDIA1.0B≈0.7 GB at 4-bit
- canary-1b-flashNVIDIA1.0B≈0.7 GB at 4-bit
- granite-4.0-1b-speechIBM watsonx.ai1.0B≈0.7 GB at 4-bit
- canary-1b-v2NVIDIA1.0B≈0.6 GB at 4-bit
- mms-1b-allfacebook1.0B≈0.6 GB at 4-bit
- distil-large-v3Whisper Distillation0.8B≈0.5 GB at 4-bit
- distil-large-v2Whisper Distillation0.8B≈0.5 GB at 4-bit
- NVIDIA Nemotron 3.5 ASR Streaming 0.6BNVIDIA0.6B≈0.4 GB at 4-bit
- Parakeet TDT 0.6B v3NVIDIA0.6B≈0.4 GB at 4-bit
- nemotron-speech-streaming-en-0.6bNVIDIA0.6B≈0.4 GB at 4-bit
- Qwen3-ASR-0.6BAlibaba0.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-ForcedAligner-0.6BAlibaba0.6B≈0.4 GB at 4-bit
- parakeet-tdt-0.6b-v2NVIDIA0.6B≈0.4 GB at 4-bit
- VibeVoice-ASR-BitNetMicrosoft0.3B≈0.2 GB at 4-bit
- wav2vec2-lg-xlsr-en-speech-emotion-recognitionehcalabres0.3B≈0.2 GB at 4-bit
- wav2vec2-large-xlsr-53-englishjonatasgrosman0.3B≈0.2 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- wav2vec2-base-960hAI at Meta0.1B≈0.1 GB at 4-bit
- ast-finetuned-audioset-10-10-0.4593Massachusetts Institute of Technology0.1B≈0.1 GB at 4-bit
- whisper-tinyOpenAI0.0B≈0.0 GB at 4-bit