Pass IndexThe State of AISign in

Writing code models that run on 64 GB

66 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.

Models and agents for reading a codebase and changing it, rather than writing a snippet. The boards that matter here are the ones that run against real repositories — SWE-bench, SWE-rebench, Terminal-Bench — because a model can write a plausible function and still fail to make a test pass. Note whether what you are looking at is a model or an agent: an agent brings the harness, the file access and the loop, and is priced for it.

Wider

Writing code modelsModels that run on 64 GB