From the archive

vLLM introduces PagedAttention for model serving

The open-source inference library uses paged memory management to serve language-model requests more efficiently.

By Chat Overview Published Updated

The vLLM team introduced its open-source inference and serving library on June 20, 2023. The project used PagedAttention to manage the memory needed while a language model generated its answer.

Memory becomes a serving constraint

During generation, a model keeps attention keys and values for the tokens already processed. These grow with the input and output, making memory allocation important when several requests share a graphics processor.

PagedAttention divided that cached data into blocks rather than requiring one continuous memory allocation per request. The approach reduced wasted memory and let the serving system handle more requests together.

A library for running models

vLLM had already been used for Chatbot Arena and the Vicuna demo before its public introduction. The announcement included benchmark comparisons, with results tied to particular hardware, models, and request patterns. The release supplied model-serving software, with the deployment still responsible for its own hardware and operation.

Source