vLLM
The most widely used open server for running language models, notable for handling many concurrent users on one GPU efficiently.
how it works · the vocabulary
SGLangTGIinference server
If you self-host, this or something like it is what actually answers the request. Its scheduling is a large part of why a company can sell an open model for a fraction of what it costs you to run the same weights badly.