From the archive

MiniMax releases M1 for long-context reasoning

MiniMax-M1 brings downloadable weights, long inputs, and separate reasoning-output variants.

By Chat Overview Published Updated

MiniMax released M1, a reasoning model with a one-million-token input context window. The release included downloadable weights and two variants with different reasoning-output budgets, named M1-40k and M1-80k.

Longer input and longer reasoning

The context window allows an application to supply a large body of text in one request. The reasoning budget controls a different part of the workload: how much output the model can spend working through a task. These are separate limits and have different effects on response time and resource use.

MiniMax highlighted software engineering, tool use, and long-document tasks as applications for the release. Its launch material described hybrid attention as part of the model's approach to long inputs.

Hosted or self-managed

MiniMax offered M1 through its application and developer API, alongside the downloadable release. The announcement also identified vLLM as a deployment option. Hosted requests avoid managing inference hardware, while a self-managed deployment needs enough capacity for the weights and long-context workload.

The launch's hosted input pricing used different tiers for shorter and longer prompts. The MiniMax profile covers current access, API billing, and the distinction between subscriptions and usage-based requests.

Source