From the archive

Fireworks introduces discounted batch inference

Asynchronous jobs process bulk model requests at half the standard serverless price, with a 24-hour turnaround.

By Chat Overview Published Updated

Fireworks introduced a Batch API for large collections of model requests that do not need an immediate response. At launch, the service offered a 50% discount against its usual serverless inference prices and a maximum turnaround of 24 hours.

Upload a job and collect the results

An application supplies its requests in a dataset, starts a batch job, and retrieves the results when processing finishes. The announced service supports the platform's open and fine-tuned models.

Useful workloads include:

  • Evaluating the same questions across model choices.
  • Classifying or extracting information from a daily document collection.
  • Generating or augmenting training examples in bulk.

A cost choice for work that can wait

Batch processing trades response time for a lower bill. It fits background jobs better than an interactive chat or a request waiting on a live web page. The launch limited input datasets to less than 500MB. The Fireworks AI profile covers its hosted inference and deployment options.

Source