From the archive

Ollama adds a cloud-model preview

Ollama adds hosted inference for larger models while retaining its local tools and API workflow.

By Chat Overview Published Updated

Ollama introduced cloud models in preview, allowing its existing tools to use models that would not fit comfortably on a personal computer. The release adds hosted inference alongside local execution rather than replacing the local model runner.

Familiar tools, remote compute

Cloud models use the same command-line and application programming interface (API) workflow as local models. The launch included hosted versions of Qwen3-Coder, gpt-oss, and DeepSeek V3.1.

The important difference is where the work runs:

  • Local models use the computer or server running Ollama.
  • Cloud models use inference hardware on Ollama's service.
  • Cloud access requires signing in, and applications can also call the cloud API directly.

That choice changes the hardware requirements and where prompts are processed. An installation of the local client does not make a cloud request an offline operation.

Choosing a deployment

The preview makes it easier to try a larger model without provisioning a model server. Local execution remains useful when the model fits the available hardware and the workload needs local processing. The Ollama profile covers the local/cloud distinction and current account and billing options.

Source