From the archive
Ollama adds a cloud-model preview
Ollama adds hosted inference for larger models while retaining its local tools and API workflow.
Ollama introduced cloud models in preview, allowing its existing tools to use models that would not fit comfortably on a personal computer. The release adds hosted inference alongside local execution rather than replacing the local model runner.
Familiar tools, remote compute
Cloud models use the same command-line and application programming interface (API) workflow as local models. The launch included hosted versions of Qwen3-Coder, gpt-oss, and DeepSeek V3.1.
The important difference is where the work runs:
- Local models use the computer or server running Ollama.
- Cloud models use inference hardware on Ollama's service.
- Cloud access requires signing in, and applications can also call the cloud API directly.
That choice changes the hardware requirements and where prompts are processed. An installation of the local client does not make a cloud request an offline operation.
Choosing a deployment
The preview makes it easier to try a larger model without provisioning a model server. Local execution remains useful when the model fits the available hardware and the workload needs local processing. The Ollama profile covers the local/cloud distinction and current account and billing options.