From the archive
MiniMax releases M1 for long-context reasoning
MiniMax-M1 brings downloadable weights, long inputs, and separate reasoning-output variants.
MiniMax released M1, a reasoning model with a one-million-token input context window. The release included downloadable weights and two variants with different reasoning-output budgets, named M1-40k and M1-80k.
Longer input and longer reasoning
The context window allows an application to supply a large body of text in one request. The reasoning budget controls a different part of the workload: how much output the model can spend working through a task. These are separate limits and have different effects on response time and resource use.
MiniMax highlighted software engineering, tool use, and long-document tasks as applications for the release. Its launch material described hybrid attention as part of the model's approach to long inputs.
Hosted or self-managed
MiniMax offered M1 through its application and developer API, alongside the downloadable release. The announcement also identified vLLM as a deployment option. Hosted requests avoid managing inference hardware, while a self-managed deployment needs enough capacity for the weights and long-context workload.
The launch's hosted input pricing used different tiers for shorter and longer prompts. The MiniMax profile covers current access, API billing, and the distinction between subscriptions and usage-based requests.