From the archive
Fireworks introduces discounted batch inference
Asynchronous jobs process bulk model requests at half the standard serverless price, with a 24-hour turnaround.
Fireworks introduced a Batch API for large collections of model requests that do not need an immediate response. At launch, the service offered a 50% discount against its usual serverless inference prices and a maximum turnaround of 24 hours.
Upload a job and collect the results
An application supplies its requests in a dataset, starts a batch job, and retrieves the results when processing finishes. The announced service supports the platform's open and fine-tuned models.
Useful workloads include:
- Evaluating the same questions across model choices.
- Classifying or extracting information from a daily document collection.
- Generating or augmenting training examples in bulk.
A cost choice for work that can wait
Batch processing trades response time for a lower bill. It fits background jobs better than an interactive chat or a request waiting on a live web page. The launch limited input datasets to less than 500MB. The Fireworks AI profile covers its hosted inference and deployment options.