Active

Together AI

Open-model inference with shared APIs, reserved throughput, and dedicated deployments.

  • Shared inference billed by model usage
  • Reserved throughput for predictable traffic
  • Dedicated endpoints and model customization

Together AI hosts open models for language, images, embeddings, and other supported tasks. It offers several deployment choices instead of a single billing model.

Deployment choices

  • Serverless inference runs models on a shared fleet and suits variable traffic.
  • Provisioned throughput reserves capacity for a supported model.
  • Dedicated inference assigns hardware to a model deployment.

Fine-tuning and GPU clusters are separate products. For a self-managed model server, see vLLM .

How pricing works

Serverless text rates separate model input and output; image workloads use the units listed for that model. Reserved capacity uses a capacity agreement, while dedicated hardware is billed for the deployment. Training and storage can add costs beyond inference.

Sources