Ollama
Run models locally or use optional cloud models through the same model tooling.
- Local models with a command line and API
- Optional cloud models for larger workloads
- Local hardware costs differ from cloud usage credits
Ollama downloads and runs supported models on a local computer. It also offers cloud models, making it possible to use models that exceed the machine's available memory.
Local or cloud
Local execution uses the computer's processor, memory, and storage. After a model has been downloaded, local inference does not require a hosted model endpoint. Cloud models run remotely and require account access. Select the model and execution location explicitly when handling sensitive material.
How pricing works
Local inference does not incur an Ollama cloud inference bill, but hardware and power still have a cost. Cloud use consumes usage credits; subscriptions include different allowances. Downloaded models retain their own license terms. For a desktop interface with local model management, see LM Studio .
Sources
Checked .
From the archive
- Ollama adds a cloud-model preview 2025-09-19
- Alibaba releases Qwen3 with thinking and non-thinking modes 2025-04-29