From the archive
Cerebras opens its inference service to developers
A hosted API starts with Llama 3.1 8B and 70B, using Cerebras hardware and a familiar chat-completions format.
Cerebras launched a hosted inference service with chat and application programming interface (API) access. Its first supported models were Meta's Llama 3.1 8B and 70B, running with their original 16-bit weights.
A different way to serve the same models
The service runs on Cerebras' wafer-scale processors rather than conventional graphics processors. The launch emphasized the speed of generating a response, a useful consideration for interactive applications and workflows that make several model calls in sequence.
For developers, the API uses the OpenAI Chat Completions request format. That gives existing chat applications a familiar starting point for testing the endpoint with a new provider and model.
Hosted access without buying hardware
Cerebras offered free, developer, and enterprise tiers, including a daily token allowance at launch. Applications could use its hardware through a service instead of operating a Cerebras system themselves. The Cerebras profile covers the hosted offering and current access routes.