Agentic software development
Plan and execute multi-file coding work with terminal tools, repository context, and iterative problem solving.
Independent tool overview
GLM-4.7 is Z.AI's MIT-licensed 358B mixture-of-experts language model for agentic coding, tool use, reasoning, and bilingual English-Chinese work, available through an API or self-hosting.
Visit the official GLM-4.7 site ↗
Overview
GLM-4.7 is an open-weight text model Z.AI released in December 2025 with a strong emphasis on completing multi-step coding and agent tasks. It can reason before acting, call tools, produce structured output, stream responses, and work with a 200K-token context window. The full checkpoint contains about 358 billion parameters and uses a mixture-of-experts architecture, so local deployment is aimed at serious inference infrastructure rather than an ordinary laptop.
Developers can call GLM-4.7 through Z.AI's OpenAI-compatible API, connect it to coding agents such as Claude Code, Cline, Kilo Code, and Roo Code, or run the MIT-licensed weights with supported engines including vLLM and SGLang. The family also includes the lower-cost GLM-4.7-FlashX and free GLM-4.7-Flash models.
GLM-4.7 remains available, but it is no longer Z.AI's newest flagship. GLM-5 and later GLM releases have moved the product line forward, so new projects should benchmark GLM-4.7's lower API price and mature open checkpoint against the latest models before standardizing on it.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Plan and execute multi-file coding work with terminal tools, repository context, and iterative problem solving.
Use a capable reasoning and coding model at lower token prices than many frontier proprietary APIs.
Deploy MIT-licensed weights in controlled infrastructure when data location, customization, or provider independence matters.
Build coding, writing, translation, long-context, and agent workflows that operate across both supported languages.
Capabilities
Focuses on task completion, including requirement interpretation, decomposition, multi-technology integration, terminal work, and repository-scale edits.
Supports reasoning controls, including preserved thinking for multi-turn agent tasks where intermediate reasoning continuity matters.
Can invoke external tools and APIs for browsing, coding agents, and other multi-step workflows.
Handles long codebases and documents with up to 200K tokens of context and up to 128K output tokens.
Works with Z.AI's chat-completions endpoint and common OpenAI SDK patterns, reducing integration work.
The 358B-parameter checkpoint is published under the MIT license and supports self-hosting with engines such as vLLM and SGLang.
GLM-4.7-FlashX offers much lower token prices, while GLM-4.7-Flash is listed as free for lightweight workloads.
Process
Step 1
Use the Z.AI API for managed inference, a coding-plan endpoint for supported IDE agents, or self-host the open weights.
Step 2
Compare full GLM-4.7 for quality, FlashX for inexpensive throughput, and Flash for free lightweight access.
Step 3
Set thinking behavior, streaming, structured output, function definitions, and context limits for the workload.
Step 4
Measure repository edits, tool-call reliability, latency, token consumption, and regressions on your own evaluations.
Step 5
Track price and quality against newer GLM releases, then pin model identifiers and test before changing production versions.
Cost
Z.AI charges the managed API per token and separately offers coding subscriptions. GLM-4.7 costs $0.60 per million input tokens, $0.11 per million cached input tokens, and $2.20 per million output tokens. FlashX is cheaper and Flash is listed as free. Self-hosting has no model license fee under MIT, but the full 358B checkpoint has substantial infrastructure costs.
$0.60 input / $2.20 output per 1M tokens
Managed full-model inference through Z.AI's general API.
$0.07 input / $0.40 output per 1M tokens
Lower-cost, higher-speed member of the GLM-4.7 family.
Free
The lightweight free-tier model in the family.
No model license fee
Download and operate the MIT-licensed model on your own infrastructure.
$18/month
Subscription access for supported coding agents, with annual billing shown at an effective $12.60 per month.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
A newer Z.AI flagship for longer-horizon coding work and substantially larger usable context.
Explore GLM 5.2 →Coding
A smaller open-weight agentic coding model that can be easier and cheaper to self-host.
Explore Qwen3-Coder-Next →Business Operations
Another cost-focused model family with open releases and strong reasoning and coding options.
Explore DeepSeek →Coding
A polished coding-agent product for teams that prioritize the complete development workflow over operating a model directly.
Explore Claude Code →Questions
Its model weights are publicly available on Hugging Face under the MIT license. That permits broad commercial use and modification, but operating the full checkpoint still requires substantial infrastructure.
The official Hugging Face checkpoint is about 358 billion total parameters and uses a mixture-of-experts architecture. Independent vLLM deployment documentation describes roughly 32 billion active parameters per token.
Z.AI lists $0.60 per million input tokens, $0.11 per million cached input tokens, and $2.20 per million output tokens. GLM-4.7-FlashX is cheaper and GLM-4.7-Flash is listed as free.
Its main strengths are agentic coding, tool use, multi-step reasoning, frontend generation, long-context text work, and English-Chinese applications.
Yes. Z.AI documents GLM-4.7 for Claude Code and other coding tools through its OpenAI- or coding-plan-compatible endpoints. It is a Z.AI model, not an Anthropic model.
Use GLM-4.7 when its lower price, stable integration, or open checkpoint meets your needs. Test newer GLM-5-series models when you need stronger current reasoning, coding, or long-horizon agent performance.
No. The documented GLM-4.7 input and output modality is text. Z.AI maintains separate vision-language and image-generation models.
Bottom line
GLM-4.7 remains an attractive value model for agentic coding: it is inexpensive through Z.AI's API, has a generous context and output envelope, integrates with common coding agents, and offers MIT-licensed weights. Its biggest tradeoff is timing—newer GLM generations now lead the family—so teams should choose it for cost, deployment control, or proven workload fit rather than assume it is still Z.AI's highest-capability model.
Visit GLM-4.7 website ↗
MiniMax 2.1 - New model with powerful capabilities across a variety of programming languages and for mobile and web app development

NousCoder-14B - Nous Research's competitive olympiad programming model

GPT-5.2-Codex - Agentic coding model for professional software engineering and defensive cybersecurity.

SERA - AI2's open-source coding agents able to be cheaply trained on private codebases with native support for Claude Code.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.