llama.cpp
A small, dependency-free engine for running models on ordinary computers, including on a laptop's CPU.
how it works · the vocabulary
Ollamalocal inference
It made local models practical for people without a data-centre card, and its GGUF format is the de facto standard for distributing quantised weights. Ollama and most desktop apps are wrappers around it.