Long-horizon software engineering
Navigate large repositories, reason over screenshots, use terminal tools, and sustain multi-step coding work inside a constrained, reviewed agent harness.
Independent tool overview
Kimi K3 is Moonshot AI's open-weight, native multimodal agentic model for long-running coding, research, document, visual, and knowledge-work tasks. It has 2.8 trillion total parameters, activates 104 billion per token, accepts text, images, and video, and supports a maximum one-million-token context window. Users can access it through Kimi's consumer and work products, Kimi Code, the hosted API, or released weights. The phrase 'open source' needs qualification: K3 uses a custom Kimi K3 License with broad use and modification rights but extra commercial conditions for certain large model-as-a-service businesses and very large products.
Visit the official Kimi K3 site ↗
Overview
Released on July 16, 2026, Kimi K3 is Moonshot AI's largest model and the successor to Kimi K2.6. It is designed for agentic work that may span many tool calls, large repositories, long documents, visual inputs, and extended execution rather than only short chat responses.
K3 is a sparse mixture-of-experts model with 2.8 trillion total parameters and 104 billion activated for each token. Its 93-layer architecture combines 69 Kimi Delta Attention layers with 24 Gated MLA layers, uses 896 experts with 16 selected per token plus two shared experts, and was trained with MXFP4 weights and MXFP8 activations.
The model accepts text, image, and video context and produces text. Its one-million-token maximum context can accommodate unusually large source sets, but maximum capacity is not a guarantee of complete recall, accurate citation, or economical processing.
K3 is available as a hosted model in Kimi, Kimi Work, Kimi Code, and the Kimi API. Full weights are also available from Moonshot AI for organizations capable of operating a very large distributed inference stack. Hosted Kimi tools such as web search, code execution, spreadsheets, and document conversion are platform capabilities and should not be mistaken for features embedded in the standalone weights.
Moonshot publishes strong coding, research, productivity, and multimodal results, but many are company-run, use different agent harnesses, or depend on maximum reasoning effort. Treat them as useful evaluation evidence—not a substitute for testing K3 on your own tasks, tools, latency targets, security controls, and cost profile.
Moonshot documents two important behavioral limits: K3 can become unstable if an agent fails to preserve its complete thinking history, and its long-horizon training can make it overly proactive when intent is ambiguous. Production agents need a compatible harness, explicit permissions, narrow tools, confirmation gates, budgets, logging, and human review.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Navigate large repositories, reason over screenshots, use terminal tools, and sustain multi-step coding work inside a constrained, reviewed agent harness.
Analyze extensive document sets, reports, tables, images, and web evidence while preserving citations and requiring human validation.
Combine text, screenshots, charts, slide decks, and video frames when the task needs both visual interpretation and structured written output.
Evaluate, fine-tune, or deploy frontier-scale weights when the organization has the hardware, distributed-systems expertise, and license review to support them.
Use an OpenAI- or Anthropic-compatible hosted API for tool-using applications that benefit from long context and selectable reasoning effort.
Compare a large open-weight model with hosted frontier systems under a controlled internal benchmark and documented cost, latency, and safety criteria.
Capabilities
Accepts up to 1,048,576 tokens for large repositories and source collections, subject to provider, memory, latency, and retrieval-quality constraints.
Processes text, images, and video within the same model for screenshot-driven coding, visual analysis, documents, and media workflows.
Uses 896 experts while selecting 16 per token, activating 104B parameters instead of all 2.8T for every token.
Combines KDA with Gated MLA and Attention Residuals to improve sequence- and depth-wise information flow at very large scale.
Targets extended repository work, terminal use, debugging, kernel optimization, compiler tasks, frontend iteration, CAD, and other tool-assisted engineering.
Supports research, document production, spreadsheet work, presentations, visualizations, and multi-step analysis through Kimi's hosted products.
The hosted API exposes low, high, and max reasoning effort; thinking remains enabled and the default is max.
Moonshot documents compatible API interfaces and the hosted model identifier kimi-k3 for existing agent integrations.
Moonshot publishes the full model through GitHub and Hugging Face under the custom Kimi K3 License.
Official instructions cover vLLM, SGLang, and TokenSpeed rather than requiring one proprietary self-hosting stack.
Moonshot's separate Vendor Verifier publishes tests intended to detect capability or long-context differences across hosted K3 providers.
The same model is available in Kimi Agent, Kimi Work, Kimi Code, the API, and enterprise products, each with different tools and operating controls.
Process
Step 1
Separate consumer use, Kimi Code, hosted API integration, third-party hosting, and self-deployment; each has different cost, privacy, tool, license, and operational implications.
Step 2
Build representative tasks with objective correctness, citation, tool-use, latency, token, human-review, and failure criteria instead of relying on provider benchmark averages.
Step 3
Classify inputs, remove secrets and unnecessary personal data, review hosted processing terms, and obtain legal review for the custom license when commercial scale or model-as-a-service thresholds may apply.
Step 4
Use an integration verified to preserve the complete assistant message, including reasoning content and tool calls, across every turn; avoid switching an existing session to K3 midstream.
Step 5
Give the agent only the minimum files, systems, network destinations, credentials, and write permissions required for the task.
Step 6
Choose low, high, or max effort deliberately, cap turns and tokens, budget web-search charges, set timeouts, and stop loops that no longer make measurable progress.
Step 7
Deduplicate sources, label dates and authority, separate instructions from untrusted content, prioritize the most relevant evidence, and test retrieval at realistic context lengths.
Step 8
Treat web pages, documents, issues, emails, images, and tool output as untrusted data; prevent embedded instructions from changing policy, revealing data, or expanding permissions.
Step 9
Pause before destructive code changes, external messages, purchases, deployments, credential use, account changes, or any action with legal, financial, privacy, or safety consequences.
Step 10
Run tests, inspect diffs, trace citations to primary sources, recalculate numbers, and involve qualified reviewers for medical, legal, financial, scientific, security, or regulated work.
Step 11
Log model version, effort, tokens, cache behavior, tool calls, failures, approvals, latency, and spend without recording sensitive reasoning or data beyond policy.
Step 12
Run the same evaluation against alternatives, inspect third-party provider fidelity, and repeat testing when model, runtime, prompt, tool, or pricing changes.
Cost
Moonshot publishes pay-as-you-go Kimi API pricing for K3. Cache-hit input is ten times cheaper than cache-miss input, while output—including reasoning tokens—has the highest unit price. Web search adds a per-call fee and returns results that also count as prompt tokens. Consumer Kimi memberships and enterprise terms use separate purchase paths, and self-hosting replaces token fees with substantial infrastructure, engineering, and licensing costs.
$0.30 / $3 / $15 per 1M tokens
Pay-as-you-go hosted inference through the official Kimi API.
$0.005 per triggered call plus tokens
Optional hosted search used by compatible API workflows.
Free and paid access; confirm in product
K3 is available through Kimi, Kimi Work, and Kimi Code, with membership pricing, quotas, and regional availability presented through the current purchase flow.
No per-token license fee published; infrastructure extra
Full weights may be used under the custom Kimi K3 License.
Custom quote
Organization access with separate accounts, administration, and enterprise terms.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
The prior Kimi generation for teams that want a smaller, more established Moonshot model and do not require K3's maximum scale.
Explore Kimi K2.6 →Miscellaneous
A far smaller open model for teams prioritizing practical fine-tuning and deployment economics over maximum frontier-scale capability.
Explore Inkling-Small →Coding
An open-weight coding-focused model for agentic software workflows with a materially lighter deployment target.
Explore Qwen3-Coder-Next →Business Operations
A broad open-model ecosystem and hosted assistant worth comparing on reasoning, coding, cost, policy, and deployment requirements.
Explore DeepSeek →Consumer
Google's open-weight model family for teams that prefer multiple model sizes and a different tooling and deployment ecosystem.
Explore Gemma 4 →Questions
Kimi K3 is Moonshot AI's 2.8-trillion-parameter, open-weight multimodal agentic model. It is built for long-horizon coding, research, documents, visual inputs, and tool-assisted knowledge work.
It is more precise to call Kimi K3 open-weight. Moonshot releases the weights and broad rights to use, modify, distribute, fine-tune, and deploy them, but the custom Kimi K3 License adds commercial conditions for certain large model-as-a-service businesses and very large products.
K3 has 2.8 trillion total parameters and activates 104 billion per token through a mixture-of-experts architecture. The active count lowers per-token computation but does not make storing or serving the full model equivalent to a dense 104B checkpoint.
It accepts text, images, and video context and produces text. Kimi's hosted products can wrap it with additional tools for coding, search, documents, spreadsheets, presentations, and interactive output.
The documented maximum is 1,048,576 tokens. Real results still depend on source organization, retrieval, runtime support, cache behavior, prompt injection defenses, latency, and the model's ability to find and use the relevant evidence.
As checked August 31, 2026, the official API lists $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens. Optional web search is $0.005 per triggered call plus tokens.
Yes under its license, but it is an expert deployment. Moonshot recommends supernode configurations with at least 64 accelerators and documents compatible vLLM, SGLang, and TokenSpeed paths. Most teams should benchmark hosted access before committing to that infrastructure.
There is no simple single-GPU minimum for the full model. The native MXFP4/MXFP8 checkpoint, 2.8T total parameters, expert parallelism, and long context demand a large high-bandwidth distributed system; exact capacity depends on runtime, context, concurrency, and performance goals.
Moonshot says K3 was trained for preserved thinking history. If an integration omits prior reasoning content or tool calls—or switches an existing session to K3—generation quality may become unstable. Use a tested compatible harness and pass the complete assistant message.
Not without controls. Moonshot warns that it can be excessively proactive. Use narrow permissions, explicit instructions, sandboxing, spending and turn limits, confirmation gates, logs, tests, and a human reviewer before consequential actions.
Not uniformly. Moonshot publishes detailed footnotes, but results mix internal evaluations, public leaderboards, company-run tests, different agent harnesses, and maximum reasoning effort. Reproduce the tasks that matter to your organization.
It can help summarize or draft from approved evidence, but it should not be the authority. Require primary sources, qualified professional review, privacy controls, calibrated uncertainty, and human responsibility for every consequential decision.
Bottom line
Kimi K3 is a serious option for teams evaluating frontier-scale open weights, very long context, multimodal reasoning, and sustained coding or research agents. The official hosted API makes evaluation accessible, but self-hosting is an infrastructure project, not a routine model download. The best buyers will benchmark their own workloads, account for reasoning and long-context costs, review the custom license, preserve complete agent history, and constrain every tool. Teams that need easier deployment should compare smaller open models; teams that need the strongest polished hosted experience should compare proprietary frontier systems on the same acceptance test.
Visit Kimi K3 website ↗
Replit Slides - Replit's new tool for creating beautiful slides in seconds

Discover AI self-improvement to reach your potential.

Gamma: Creates beautiful, interactive presentations and docs using generative ai — no design skills needed.

Search and discover personalized recommendations from Forbes.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.