From the archive

Alibaba releases Qwen3 with thinking and non-thinking modes

Qwen3 adds switchable reasoning modes and downloadable models across a range of sizes.

By Chat Overview Published Updated

Alibaba's Qwen team released Qwen3 with two operating modes: longer reasoning for complex questions and faster responses for simpler tasks. The launch included six dense models and two mixture-of-experts models, making the family available at several hardware scales.

Choosing speed or longer reasoning

Thinking mode produces intermediate reasoning before the answer. Non-thinking mode avoids that longer process when a quick response is more useful. The choice affects response time and output length, so it can matter for both interactive applications and inference costs.

The team released the models under Apache 2.0 and made downloads available through model repositories. The Alibaba Model Studio / Qwen profile explains the model family and hosted access.

Local tools and server deployment

The launch announcement identified Ollama and LM Studio as options for local use, and vLLM for serving models. Smaller models fit different hardware budgets from the largest mixture-of-experts releases. The number of active parameters does not describe the full memory needed to store a model's weights. Choose a model size and serving configuration that fit the available hardware before treating the whole family as interchangeable.

Source