Custom model development
Fine-tune a strong multimodal base on proprietary tasks, preferences, or domain data and export the resulting weights.
Independent tool overview
Inkling is Thinking Machines Lab's open-weight multimodal model for text, images, and audio, built for adjustable reasoning, coding, tool use, and customization through Tinker.
Visit the official Inkling site ↗
Overview
Inkling is a general-purpose, open-weight model from Thinking Machines Lab. It is a 975-billion-parameter mixture-of-experts model with 41 billion active parameters and native reasoning across text, images, and audio.
Its defining feature is customization. The full Apache-2.0-licensed weights are available for self-hosting, while Thinking Machines' Tinker platform provides playground access, serverless inference, fine-tuning, and downloadable custom checkpoints. Reasoning effort can be adjusted to trade speed and cost against task performance.
Inkling-Small is now a full member of the family rather than a preview. It uses 276 billion total parameters and 12 billion active parameters, retains the same multimodal and controllable-effort approach, and is the more practical starting point when latency, hosted cost, or deployment scale matters.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Fine-tune a strong multimodal base on proprietary tasks, preferences, or domain data and export the resulting weights.
Build coding assistants and agents that combine long reasoning, tool calls, and large working contexts.
Create applications that reason jointly over text, images, and audio rather than relying on separate single-modality models.
Inspect, host, customize, and evaluate a permissively licensed model without depending entirely on one hosted API.
Capabilities
Inkling accepts text, image, and audio inputs in a shared model and generates text outputs.
Developers can vary the amount of reasoning to balance quality, latency, and generated-token cost.
The model is trained for coding, tool use, instruction following, and long iterative workflows.
The open model supports up to 1 million tokens, while Tinker currently offers 64K and 256K configurations.
Tinker supports supervised and reinforcement fine-tuning, sampling, a playground, and downloadable checkpoints.
A 276B-total, 12B-active sibling offers similar modalities and reasoning controls with lower compute and hosted prices.
Process
Step 1
Benchmark Inkling-Small first, then move to full Inkling only when its quality gains justify higher inference, training, and deployment costs.
Step 2
Use the shortest viable context and lowest reasoning effort that meets the task's quality target.
Step 3
Measure factuality, task completion, tool use, latency, safety, and cost on representative examples before fine-tuning.
Step 4
Train against domain data or a verifiable reward, compare the checkpoint with the base model, and retain a holdout set.
Step 5
Use Tinker's beta serverless inference for testing, a supported partner for production, or self-host when control justifies the infrastructure.
Cost
Tinker prices sampling, fine-tuning, and beta serverless inference per million tokens. Inkling and Inkling-Small currently carry limited-time 50% fine-tuning discounts, so the listed promotional rates can change.
$1.00 input / $0.17 cached / $4.05 output per 1M
256K-context hosted inference for evaluation and lighter use.
$0.30 input / $0.06 cached / $1.20 output per 1M
Lower-cost 256K-context inference for the smaller model.
$1.87 input / $4.68 sample / $5.61 train per 1M
Limited-time discounted 64K pricing for sampling and fine-tuning full Inkling.
$0.58 input / $1.44 sample / $1.73 train per 1M
Limited-time discounted 64K pricing for the smaller model.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Business Operations
A lower-cost open-weight model family with 1M context and mature hosted APIs.
Explore DeepSeek →Agents
Another large open-weight multimodal agent model with strong coding and tool-use capabilities.
Explore Kimi K2.5 →Agents
A broader open-weight platform with multiple model sizes and enterprise deployment options.
Explore Mistral AI →Project Management
A managed alternative for teams that value polished reasoning and coding tools over weight access.
Explore Claude →Questions
Inkling's full model weights are published under the Apache 2.0 license. Open weights provide broad use and modification rights, though the complete Tinker hosted platform is a separate commercial service.
Inkling accepts text, image, and audio inputs and produces text. It is trained for general reasoning, coding, tool use, instruction following, and multimodal tasks.
Full Inkling has 975B total and 41B active parameters. Inkling-Small has 276B total and 12B active parameters, making it less expensive and easier to serve while retaining the family's main modalities and effort controls.
The open model supports up to 1 million tokens. Tinker currently provides 64K and 256K configurations, so the usable limit depends on the deployment.
Tinker's beta serverless Inkling pricing is $1.00 per million input tokens, $0.17 for cached input, and $4.05 for output. Fine-tuning and sampling have separate usage-based rates, with a limited-time 50% discount currently shown.
Yes. Tinker supports fine-tuning and lets users download trained checkpoints. The open weights also allow customization through supported third-party and self-hosted stacks.
Bottom line
Inkling is most compelling as a customizable multimodal foundation model, not simply another chat assistant. Its open weights, audio support, controllable reasoning, and Tinker workflow give developers unusual flexibility. Inkling-Small should be the default evaluation candidate for cost and deployment practicality, with full Inkling reserved for workloads where measured quality gains justify the much larger footprint.
Visit Inkling website ↗
Bonsai 27B - PrismML's compressed 27B model small enough for iPhones

Hint - Martha Stewart's AI home app for maintenance, repairs, and fair quotes

ChatGPT Work- OpenAI's Codex-powered agent for everyday non-coding work
.webp)
Muse Glimmer - Meta's open-weights local model for on-device agents

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.