From the archive

AI21 releases Jamba with a 256K context window

Jamba combines Transformer and Mamba layers in a downloadable base model for long-text applications.

By Chat Overview Published Updated

AI21 released Jamba, a language model with downloadable weights and a 256,000-token context window. The release uses a mixture of Transformer and Mamba layers, an approach designed to reduce the memory and processing cost of long inputs.

A base for custom applications

Jamba arrived on Hugging Face under the Apache 2.0 license. It was a base model for training and adaptation, rather than a finished chat assistant. AI21 said an instruction-tuned version would follow through its hosted platform.

Long context makes room for substantial documents or several files in a single request. Running that workload still requires suitable hardware: the announcement discussed an 80GB graphics processor, rather than a typical laptop.

Choosing the access route

The downloadable release gave teams a way to adapt and run the model themselves. AI21 also announced plans to make Jamba available through NVIDIA's model catalog. The AI21 profile covers the company's model and application services.

Source