Pass IndexThe State of AISign in

Transcription models that run on 64 GB

39 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.

Models that turn recorded speech into text. They are metered by the minute or the second of audio, so cost follows the length of the recording and not the difficulty of it. What separates them is languages covered, whether they mark who is speaking, and whether they run in real time or only on a finished file — a model that is excellent on a podcast may be unusable on a live call.

Wider

Transcription modelsModels that run on 64 GB