Active
Together AI
Open-model inference with shared APIs, reserved throughput, and dedicated deployments.
- Shared inference billed by model usage
- Reserved throughput for predictable traffic
- Dedicated endpoints and model customization
Together AI hosts open models for language, images, embeddings, and other supported tasks. It offers several deployment choices instead of a single billing model.
Deployment choices
- Serverless inference runs models on a shared fleet and suits variable traffic.
- Provisioned throughput reserves capacity for a supported model.
- Dedicated inference assigns hardware to a model deployment.
Fine-tuning and GPU clusters are separate products. For a self-managed model server, see vLLM .
How pricing works
Serverless text rates separate model input and output; image workloads use the units listed for that model. Reserved capacity uses a capacity agreement, while dedicated hardware is billed for the deployment. Training and storage can add costs beyond inference.