From the archive
Together adds Turbo and Lite endpoints for Llama 3
New serving options offer different precision and cost choices alongside full-precision model endpoints.
Together AI introduced a new inference engine and Turbo and Lite endpoints, starting with Meta's Llama 3 models. The release adds serving choices around speed, numerical precision, and price.
Choose a serving option for the workload
The announced endpoint families include:
- Turbo, using eight-bit floating-point optimization.
- Lite, using four-bit integer optimization for a lower-cost route.
- Reference, retaining full 16-bit model weights.
Reducing numerical precision changes how a model is served, so an application can test the options against its own questions and outputs. The model name alone does not describe every part of that deployment choice.
Hosted inference with more than one price point
The release provides alternatives for applications with different cost and performance requirements. It also expands Together's optimized serving stack rather than requiring teams to build that stack themselves. The Together AI profile covers its serverless and dedicated model services.