Nebius Token Factory
Open-model inference APIs, dedicated endpoints, and customization on Nebius infrastructure.
- Hosted language models and embeddings
- Model variants with different pricing and latency
- Dedicated endpoints and batch inference
Nebius Token Factory provides hosted access to open models for chat, coding, retrieval, and related application tasks. Its public model catalogue lists the available models and serving options.
Choosing an endpoint
Supported models can have base and fast serving variants with different latency and token prices. Dedicated endpoints and batch inference support different traffic patterns from a shared interactive API. Customization options include post-training and deploying custom models.
How pricing works
The pricing dashboard requires sign-in. Input and output token rates depend on the exact model and serving variant. Dedicated capacity requires its own configuration and commercial terms. Compare input and output rates for the chosen variant, alongside capacity and regional availability.
Connection with Clarifai
Nebius acquired related inference patents and licensed technology while Clarifai's core engineering and research team joined it. That agreement excluded Clarifai's legacy computer-vision and US government and defense products. Existing Clarifai accounts and contracts need a separate discussion with support; this is not an automatic account migration.