vLLM
Open-source model inference and serving software for running a model service on managed hardware.
- Model serving with batching and memory management
- OpenAI-compatible server for supported models
- Apache-licensed code with separate hardware costs
vLLM is software for running model inference and serving it to other applications. Its batching and memory-management features help share accelerator capacity across requests. It is a self-managed serving stack rather than a hosted inference subscription.
Running a service
The project provides an OpenAI-compatible server for supported models. Deployment requires compatible hardware, model files, and an environment that meets the selected model's requirements. Operations such as access control, capacity planning, upgrades, and monitoring remain part of running the service.
Costs and licensing
The vLLM code uses the Apache 2.0 license. Model licenses apply separately. Hardware rental, storage, power, and administration determine the operating cost. For managed alternatives, see Together AI or Fireworks AI .