The Rundown AI homepage

Independent tool overview

Inkling at a glance

Inkling is Thinking Machines Lab's open-weight multimodal model for text, images, and audio, built for adjustable reasoning, coding, tool use, and customization through Tinker.

Visit the official Inkling site ↗
Inkling product preview
Developer
Thinking Machines Lab
Architecture
975B MoE; 41B active parameters
Context
Up to 1M tokens; 64K or 256K on Tinker
Inputs
Text, images, and audio
License
Apache 2.0

Overview

What Inkling is

Inkling is a general-purpose, open-weight model from Thinking Machines Lab. It is a 975-billion-parameter mixture-of-experts model with 41 billion active parameters and native reasoning across text, images, and audio.

Its defining feature is customization. The full Apache-2.0-licensed weights are available for self-hosting, while Thinking Machines' Tinker platform provides playground access, serverless inference, fine-tuning, and downloadable custom checkpoints. Reasoning effort can be adjusted to trade speed and cost against task performance.

Inkling-Small is now a full member of the family rather than a preview. It uses 276 billion total parameters and 12 billion active parameters, retains the same multimodal and controllable-effort approach, and is the more practical starting point when latency, hosted cost, or deployment scale matters.

Use cases

Who Inkling is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Custom model development

Fine-tune a strong multimodal base on proprietary tasks, preferences, or domain data and export the resulting weights.

Agentic coding and tools

Build coding assistants and agents that combine long reasoning, tool calls, and large working contexts.

Audio and image understanding

Create applications that reason jointly over text, images, and audio rather than relying on separate single-modality models.

Open-weight research

Inspect, host, customize, and evaluate a permissively licensed model without depending entirely on one hosted API.

Capabilities

Core Inkling features

1

Native multimodality

Inkling accepts text, image, and audio inputs in a shared model and generates text outputs.

2

Controllable thinking effort

Developers can vary the amount of reasoning to balance quality, latency, and generated-token cost.

3

Agentic capabilities

The model is trained for coding, tool use, instruction following, and long iterative workflows.

4

One-million-token context

The open model supports up to 1 million tokens, while Tinker currently offers 64K and 256K configurations.

5

Tinker customization

Tinker supports supervised and reinforcement fine-tuning, sampling, a playground, and downloadable checkpoints.

6

Inkling-Small

A 276B-total, 12B-active sibling offers similar modalities and reasoning controls with lower compute and hosted prices.

Process

How the Inkling workflow works

  1. Step 1

    Choose the model size

    Benchmark Inkling-Small first, then move to full Inkling only when its quality gains justify higher inference, training, and deployment costs.

  2. Step 2

    Set context and effort

    Use the shortest viable context and lowest reasoning effort that meets the task's quality target.

  3. Step 3

    Build an evaluation set

    Measure factuality, task completion, tool use, latency, safety, and cost on representative examples before fine-tuning.

  4. Step 4

    Customize on Tinker

    Train against domain data or a verifiable reward, compare the checkpoint with the base model, and retain a holdout set.

  5. Step 5

    Choose deployment

    Use Tinker's beta serverless inference for testing, a supported partner for production, or self-host when control justifies the infrastructure.

Cost

Inkling pricing and free plan

Tinker prices sampling, fine-tuning, and beta serverless inference per million tokens. Inkling and Inkling-Small currently carry limited-time 50% fine-tuning discounts, so the listed promotional rates can change.

Inkling serverless beta

$1.00 input / $0.17 cached / $4.05 output per 1M

256K-context hosted inference for evaluation and lighter use.

  • Serverless beta
  • Not recommended for intensive production yet
  • Usage based

Inkling-Small serverless beta

$0.30 input / $0.06 cached / $1.20 output per 1M

Lower-cost 256K-context inference for the smaller model.

  • Text, image, and audio
  • Serverless beta
  • Usage based

Inkling 64K on Tinker

$1.87 input / $4.68 sample / $5.61 train per 1M

Limited-time discounted 64K pricing for sampling and fine-tuning full Inkling.

  • Cached input: $0.374 per 1M
  • 50% promotional discount
  • Checkpoint storage extra

Inkling-Small 64K on Tinker

$0.58 input / $1.44 sample / $1.73 train per 1M

Limited-time discounted 64K pricing for the smaller model.

  • Cached input: $0.116 per 1M
  • 50% promotional discount
  • Checkpoint storage: $0.10/GB-month

Pricing checked . Check current pricing at the source ↗

Assessment

Inkling strengths and limitations

Where it stands out

  • Open weights and an Apache 2.0 license support customization and deployment flexibility
  • Text, image, and audio inputs are handled natively in one model
  • Adjustable reasoning effort creates a direct quality, latency, and cost control
  • Tinker connects experimentation, fine-tuning, evaluation, and checkpoint export
  • Inkling-Small offers a more deployable option without leaving the model family

What to consider

  • The full Inkling checkpoint is roughly 1.9 TB and self-hosting requires substantial specialized infrastructure
  • Tinker's hosted context options are 64K and 256K rather than the model's full 1M context
  • Tinker serverless inference is beta and is not recommended by the vendor for intensive production use
  • Current Tinker fine-tuning prices include a limited-time discount and may rise
  • Thinking Machines says Inkling is not the strongest model overall, so its value depends on customization, modalities, and openness
  • Published benchmark and capability results are primarily vendor-reported and should be reproduced on target workloads

Compare

Inkling alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Business Operations

DeepSeek

A lower-cost open-weight model family with 1M context and mature hosted APIs.

Explore DeepSeek

Agents

Kimi K2.5

Another large open-weight multimodal agent model with strong coding and tool-use capabilities.

Explore Kimi K2.5

Agents

Mistral AI

A broader open-weight platform with multiple model sizes and enterprise deployment options.

Explore Mistral AI

Project Management

Claude

A managed alternative for teams that value polished reasoning and coding tools over weight access.

Explore Claude

Questions

Inkling FAQs

Is Inkling open source?

Inkling's full model weights are published under the Apache 2.0 license. Open weights provide broad use and modification rights, though the complete Tinker hosted platform is a separate commercial service.

What can Inkling understand?

Inkling accepts text, image, and audio inputs and produces text. It is trained for general reasoning, coding, tool use, instruction following, and multimodal tasks.

What is the difference between Inkling and Inkling-Small?

Full Inkling has 975B total and 41B active parameters. Inkling-Small has 276B total and 12B active parameters, making it less expensive and easier to serve while retaining the family's main modalities and effort controls.

Can Inkling use a 1M-token context window?

The open model supports up to 1 million tokens. Tinker currently provides 64K and 256K configurations, so the usable limit depends on the deployment.

How much does Inkling cost?

Tinker's beta serverless Inkling pricing is $1.00 per million input tokens, $0.17 for cached input, and $4.05 for output. Fine-tuning and sampling have separate usage-based rates, with a limited-time 50% discount currently shown.

Can I fine-tune Inkling?

Yes. Tinker supports fine-tuning and lets users download trained checkpoints. The open weights also allow customization through supported third-party and self-hosted stacks.

Bottom line

Our Inkling verdict

Inkling is most compelling as a customizable multimodal foundation model, not simply another chat assistant. Its open weights, audio support, controllable reasoning, and Tinker workflow give developers unusual flexibility. Inkling-Small should be the default evaluation candidate for cost and deployment practicality, with full Inkling reserved for workloads where measured quality gains justify the much larger footprint.

Visit Inkling website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.