Fireworks AI
Hosted open-model inference, dedicated deployments, and training tools for AI applications.
- Serverless language, vision, and embedding models
- Separate input, cached-input, and output rates
- Dedicated deployment and training options
Fireworks AI provides model APIs and deployment tools for applications built around open models. The hosted service handles model serving, while dedicated deployments allow more control over a particular workload.
Serving options
Serverless endpoints offer shared model access. Supported serving paths trade different capacity and latency options, and batch processing suits work that does not need an immediate answer. Customization and training have their own setup and charges.
How pricing works
Serverless text and vision calls separate input, cached input, and output tokens. Embedding calls charge for input tokens. Model choice and serving path affect the rate; a dedicated deployment has separate hardware costs.