Agentic software development
Long, multi-step tasks that require reading a repository, editing files, running tools and iterating on failures.
Independent tool overview
Kimi K2.7 Code is Moonshot AI's open-weight, multimodal coding model for long-horizon software tasks, tool use and agent workflows, available through Kimi Code, a pay-as-you-go API and self-hosted deployment.
Visit the official Kimi K2.7 Code site ↗
Overview
Kimi K2.7 Code is a coding-focused mixture-of-experts model built on Kimi K2.6. It has one trillion total parameters, activates 32 billion per token, supports a 256,000-token context and accepts text, image and video input in supported environments.
Moonshot positions the model for end-to-end software engineering and reports about 30% lower thinking-token usage than K2.6. Its model card shows strong gains on Moonshot's coding and tool-use evaluations, while several comparison benchmarks are internal or use different agent harnesses and reasoning settings. Treat those numbers as vendor evidence to reproduce on your own repositories.
There are three distinct ways to use it: Kimi Code subscriptions for terminal, VS Code and compatible coding clients; the Kimi API with token-based billing; or downloaded weights under a Modified MIT License. Self-hosting is technically possible but the trillion-parameter model is infrastructure-heavy—the official guide's standard examples use eight H200 GPUs.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Long, multi-step tasks that require reading a repository, editing files, running tools and iterating on failures.
Developers comparing capable coding models at published per-token prices.
Organizations with substantial GPU or heterogeneous inference infrastructure and a reason to control model deployment.
Capabilities
A 256K-token context supports large code excerpts, tool histories and multi-step agent sessions.
The model supports interleaved reasoning and multi-step tool calls for coding agents and MCP-style environments.
Image and video inputs can supply screenshots, diagrams and visual application context; video support is limited by the access path.
Prior reasoning content is retained across turns to improve continuity in long coding workflows, and the model card says this mode cannot be disabled.
Moonshot publishes native INT4 quantization and deployment guidance for vLLM, SGLang and KTransformers.
Use the model in the official terminal and VS Code tools or connect supported third-party coding clients with a membership key.
Official API and coding endpoints provide familiar interfaces for integrating the model into existing developer tooling.
Process
Step 1
Use Kimi Code for an integrated coding agent, the token API for an application, or self-hosting only when control justifies the infrastructure.
Step 2
Give the agent clear acceptance criteria, relevant tests and the smallest permissions needed to inspect and modify the code.
Step 3
Let the model inspect repository instructions and tool output, but compact or restart sessions when stale context begins to outweigh useful evidence.
Step 4
Inspect diffs, run automated tests, scan dependencies and verify security-sensitive logic before merging or deploying.
Step 5
Track cache hit rate, input and reasoning growth, output tokens, latency and successful task completion rather than comparing headline token prices alone.
Cost
Downloaded weights have no usage fee but require expensive infrastructure and license compliance. The official Kimi API bills cached input, uncached input and output separately. Kimi Code is also bundled into RMB-priced Kimi memberships with shared credits and rate limits.
Free to download
Self-host Kimi K2.7 Code under Moonshot's Modified MIT License.
$0.19 cache / $0.95 input / $4 output per 1M tokens
Pay-as-you-go official model inference.
¥49 per month
Entry Kimi membership with standard Kimi Code access.
¥99 per month
Higher shared credits and productivity features.
¥199 per month
Professional membership that unlocks K2.7 Code HighSpeed.
¥699 per month
The highest published individual Kimi membership.
Varies by provider
Hosted access is also listed through Hugging Face inference partners.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Choose Qwen3-Coder when comparing another open-weight mixture-of-experts family built specifically for agentic coding.
Explore Qwen3-Coder →Coding
Choose GLM-4.7 for a different open-weight reasoning and coding stack with its own hosting and API options.
Explore GLM-4.7 →Coding
Choose Claude Code when a mature managed coding-agent product matters more than downloadable model weights.
Explore Claude Code →Questions
Moonshot publishes the model weights and related code under a Modified MIT License. The modification adds a branding requirement for very large commercial products, so review the actual license before distribution.
The current official platform lists $0.19 per million cache-hit tokens, $0.95 per million uncached input tokens and $4 per million output tokens.
The weights can be self-hosted, but this is not a typical laptop model. Moonshot's straightforward vLLM and SGLang examples use tensor parallelism across eight H200 GPUs, with a more complex CPU-GPU route also documented.
Kimi K2.7 Code is the model. Kimi Code is Moonshot's coding-agent service and clients, which can use K2.7 Code or newer supported Kimi models through a membership subscription.
Yes, the model card lists both. Image input works in official examples, while video is currently an official-API feature and experimental or unavailable in some third-party deployments.
Moonshot says it uses about 30% fewer thinking tokens than K2.6. That is not the same as 30% lower total cost because input, output, cache behavior, task success and access pricing also matter.
Bottom line
Kimi K2.7 Code is a credible option for teams seeking a strong coding model with open weights and low published API rates. The managed API or Kimi Code service will be the practical route for most users; self-hosting only makes sense at serious scale. Evaluate it on private tasks, account for forced reasoning and verify every code change rather than buying from vendor benchmarks alone.
Visit Kimi K2.7 Code website ↗
⚙️ North Mini Code - Cohere's first open-source agentic coding model

GLM 5.2 - Z AI's new flagship coding model with usable 1M context

Devin Desktop - Cognition's Windsurf rebrand for managing all your coding agents in a single platform

eve - Vercel's open-source framework that turns a file directory into an agent

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.