From the archive
Alibaba releases Qwen3 with thinking and non-thinking modes
Qwen3 adds switchable reasoning modes and downloadable models across a range of sizes.
Alibaba's Qwen team released Qwen3 with two operating modes: longer reasoning for complex questions and faster responses for simpler tasks. The launch included six dense models and two mixture-of-experts models, making the family available at several hardware scales.
Choosing speed or longer reasoning
Thinking mode produces intermediate reasoning before the answer. Non-thinking mode avoids that longer process when a quick response is more useful. The choice affects response time and output length, so it can matter for both interactive applications and inference costs.
The team released the models under Apache 2.0 and made downloads available through model repositories. The Alibaba Model Studio / Qwen profile explains the model family and hosted access.
Local tools and server deployment
The launch announcement identified Ollama and LM Studio as options for local use, and vLLM for serving models. Smaller models fit different hardware budgets from the largest mixture-of-experts releases. The number of active parameters does not describe the full memory needed to store a model's weights. Choose a model size and serving configuration that fit the available hardware before treating the whole family as interchangeable.