Active

Fireworks AI

Hosted open-model inference, dedicated deployments, and training tools for AI applications.

  • Serverless language, vision, and embedding models
  • Separate input, cached-input, and output rates
  • Dedicated deployment and training options

Fireworks AI provides model APIs and deployment tools for applications built around open models. The hosted service handles model serving, while dedicated deployments allow more control over a particular workload.

Serving options

Serverless endpoints offer shared model access. Supported serving paths trade different capacity and latency options, and batch processing suits work that does not need an immediate answer. Customization and training have their own setup and charges.

How pricing works

Serverless text and vision calls separate input, cached input, and output tokens. Embedding calls charge for input tokens. Model choice and serving path affect the rate; a dedicated deployment has separate hardware costs.

Sources