Self-hosted coding agents
Teams that want an open model behind a controlled CLI or IDE agent and can operate the required inference stack.
Independent tool overview
Qwen3-Coder-Next is an Apache-2.0 open-weight coding model built for local and self-hosted agents, with 80 billion total parameters, 3 billion activated per token and a native 256K-token context window.
Visit the official Qwen3-Coder-Next site ↗
Overview
Qwen3-Coder-Next is Qwen's efficiency-focused model for coding agents. It is designed to plan across repositories, call tools, recover from failed executions and work inside agent scaffolds such as Qwen Code, Claude Code, Cline, Kilo, Trae and Qoder.
Its hybrid mixture-of-experts architecture has 80 billion total parameters but activates about 3 billion for each token. That can reduce inference work compared with a dense 80B model, but it does not make the download, memory footprint or operational burden equivalent to a 3B model.
The weights and code are available under Apache 2.0, while deployment, storage and inference still carry real costs. Qwen's benchmark results are useful directional evidence, not a guarantee that the model will solve a particular codebase, and any agent using shell, network or repository tools needs strict permissions and human-reviewed changes.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Teams that want an open model behind a controlled CLI or IDE agent and can operate the required inference stack.
Code exploration, multi-file changes, test-driven iteration and tool-assisted debugging where a long context window is useful.
Engineering organizations comparing an Apache-licensed model with hosted coding assistants on their own private task set.
Capabilities
Qwen trained the model for executable tasks, environment interaction, tool use and recovery from failures rather than only code completion.
The model contains 80B parameters and activates 3B per token across 512 routed experts, 10 active experts and one shared expert.
A 262,144-token native window supports large code and tool histories; the repository also documents YaRN extension to 1M tokens.
The official materials describe use with Qwen Code, Claude Code, Cline, Kilo, Trae, Qoder and other compatible scaffolds.
vLLM and SGLang can expose the model through a familiar chat-completions endpoint with Qwen's tool-call parser.
Quantizations and integrations are available for Ollama, LM Studio, MLX-LM, llama.cpp and KTransformers.
The model card reports 44.3 on SWE-bench Pro, 70.6 on SWE-bench Verified and 36.2 on Terminal-Bench 2.0 under the cited evaluation setup.
Process
Step 1
Select real bug fixes, refactors and tests from your stack, then define success, latency, cost and security criteria before adopting the model.
Step 2
Start with a supported quantization or current Transformers, vLLM or SGLang release; budget for the full 80B weight set and reduce context if memory is constrained.
Step 3
Run in an isolated workspace with least-privilege file, shell, network and credential access, and require approval for destructive or production actions.
Step 4
Provide repository instructions, acceptance tests, relevant files and explicit commands instead of asking for an unconstrained autonomous rewrite.
Step 5
Inspect the diff, run formatting, linting, type checks, tests and security scans, and require a qualified engineer to approve dependencies, migrations and deployment.
Cost
The official Qwen3-Coder-Next weights and repository are free to use under Apache 2.0. There is no single official per-token price for the open model: total cost depends on local hardware, cloud GPUs, storage, traffic and any third-party inference provider.
Free under Apache 2.0
Download and modify the official model subject to the license terms.
Compute-dependent
Operate the model on owned or rented hardware using a supported serving stack.
Provider pricing
Use a managed inference provider when offered rather than operating the model yourself.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
A hosted, end-to-end coding environment for people who prefer managed infrastructure and integrated app deployment.
Explore Replit →Business Operations
A hosted general AI assistant with coding help and a broader consumer interface, without self-hosting the model weights.
Explore ChatGPT →Questions
It is an open-weight Qwen language model specialized for coding agents, local development, tool calling and long-horizon software tasks.
Its official weights and repository are published under Apache 2.0. That permits broad use and modification, but users still need to comply with the license and review licenses for dependencies and generated material.
No. It is an 80B-parameter mixture-of-experts model that activates about 3B parameters per token. The active count helps explain inference efficiency, while storage and memory planning must account for the full weights and chosen quantization.
The model card documents 262,144 tokens natively. Qwen's repository describes extension to 1M tokens with YaRN, but longer windows increase resource use and should be tested for accuracy and latency.
Yes, with sufficient hardware or an appropriate quantization. Qwen lists Transformers, Ollama, LM Studio, MLX-LM, llama.cpp and KTransformers, while vLLM and SGLang are documented for servers.
The model files are free under Apache 2.0. Actual cost comes from hardware, cloud GPU time, storage, operations or a third-party provider; there is no universal official per-token price for the open model.
Qwen says it is comparable on selected agentic coding and browser-use benchmarks. That is a vendor comparison, not a blanket equivalence; reproduce relevant tasks with the same tools, budget and review rules.
It should not receive unrestricted production access. Use a sandbox, least-privilege credentials, approval gates and a human-reviewed pull request, then run the project's full test and security checks.
Bottom line
Qwen3-Coder-Next is a compelling open model for teams willing to operate their own coding-agent stack: its Apache license, sparse architecture, long context and broad runtime support create real deployment flexibility. The tradeoff is operational weight and agent risk—its 3B active count understates the full 80B footprint, and strong benchmarks do not remove the need for sandboxing, tests and expert code review.
Visit Qwen3-Coder-Next website ↗
Codex App - OpenAI's new Mac app interface for managing agents with its Codex agentic coding assistant

GPT-5.3-Codex - OpenAI's new SOTA agentic coding model

SERA - AI2's open-source coding agents able to be cheaply trained on private codebases with native support for Claude Code.

Composer 1.5 - Cursor's updated in-house agentic coding model with increased performance and speed

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.