Active

Cerebras Inference

Hosted model inference for chat, code, and agent workloads, with dedicated capacity options.

  • Hosted inference without managing GPU servers
  • Developer API and dedicated deployment options
  • Trial, developer, and enterprise access

Cerebras serves language models through an inference API backed by its own hardware. It supports workloads such as interactive chat, code generation, and tools that make repeated model calls.

Access and capacity

Developer access offers a hosted API with account and model limits. Dedicated capacity is available for workloads that need a separate deployment arrangement. Preview models are intended for evaluation and may change; use a production-supported model for a deployed application.

How pricing works

Self-service model inference uses the published model rates and account credits. Trial access is limited. Enterprise capacity and support require a separate agreement, so include expected request volume and context sizes in a quote.

Sources