Cold start
The delay when a model has to be loaded onto a GPU before it can answer, because nobody has asked for it recently.
how it works · the vocabulary
scale to zero
Seconds to minutes, depending on the model's size. It is the hidden cost of serving a long tail of models, and the reason a rarely used model on a serverless platform can feel broken while a popular one on the same platform is instant.