Pass IndexThe State of AISign in

Speech models that run on 24 GB

51 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.

Models that read text aloud. They are usually priced per character, which makes the arithmetic simple: a thousand characters is roughly a paragraph. The real choices are the voice itself, whether you may clone one, how much control you have over emotion and pacing, and latency — a model that sounds wonderful in a rendered file may be too slow to hold a conversation.

Wider

Speech modelsModels that run on 24 GB