The Rundown AI homepage

Independent tool overview

Hunyuan Vision 1.5 Thinking at a glance

Hunyuan Vision 1.5 Thinking is Tencent's deep-reasoning vision-language model for image question answering, visual grounding, OCR, charts, STEM problems, and image-based creative tasks. It remains available through Tencent Cloud TokenHub under the model ID hunyuan-t1-vision-20250916.

Visit the official Hunyuan Vision 1.5 Thinking site ↗
Hunyuan Vision 1.5 Thinking product preview
Developer
Tencent Hunyuan
Model ID
hunyuan-t1-vision-20250916
Access
Tencent Cloud TokenHub
Context window
40K tokens
Maximum output
24K tokens
Media support
One image per request; no video

Overview

What Hunyuan Vision 1.5 Thinking is

Hunyuan Vision 1.5 Thinking analyzes an image together with a text instruction and produces a reasoned text response. Its intended workloads include locating objects, reading text, interpreting diagrams and charts, solving photographed problems, and answering questions that require several visual reasoning steps.

Tencent describes the model as a Mamba-Transformer hybrid with a thinking-on-images approach. The research framing includes visual reflection operations such as cropping, zooming, and drawing points, lines, or boxes to inspect an image more deliberately.

The serving path changed in 2026. Tencent retired the old Hunyuan multimodal endpoint on June 22 and moved HY-Vision-1.5-Thinking to TokenHub, which exposes an OpenAI-compatible API. The model is active, but integrations must use the new endpoint.

Use cases

Who Hunyuan Vision 1.5 Thinking is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

OCR and document images

Extract and reason over text in screenshots, photographed pages, forms, diagrams, and other image-based documents.

Charts and visual data

Interpret chart structure, compare plotted values, explain visual patterns, and answer questions grounded in an image.

STEM problem solving

Analyze photographed exercises, geometry, formulas, and other visual problems that benefit from deep reasoning.

Multilingual applications

Build image-understanding workflows for Chinese and more than 30 listed languages through Tencent TokenHub.

Capabilities

Core Hunyuan Vision 1.5 Thinking features

1

Deep visual reasoning

Uses a thinking-oriented response path for questions that require several reasoning steps instead of only a short caption.

2

Visual grounding

Identifies and reasons about where an object or relevant region appears within an input image.

3

OCR and chart understanding

Reads visible text and interprets structured visual information such as tables, plots, diagrams, and screenshots.

4

Image question answering

Answers general questions, analyzes scene content, and supports multi-turn conversations grounded in an uploaded image.

5

Multilingual interaction

Tencent lists more than 30 major languages and highlights improved English and smaller-language performance.

6

OpenAI-compatible API

Calls the model through TokenHub's chat-completions-compatible endpoint using the OpenAI SDK or another compatible client.

Process

How the Hunyuan Vision 1.5 Thinking workflow works

  1. Step 1

    Create a TokenHub service

    Use Tencent Cloud TokenHub rather than the retired Hunyuan endpoint, then enable the model's free trial or pay-as-you-go service.

  2. Step 2

    Configure the client

    Point an OpenAI-compatible SDK to the TokenHub base URL and select hunyuan-t1-vision-20250916 as the model.

  3. Step 3

    Prepare one clear image

    Crop to the relevant content and preserve readable resolution because HY-Vision accepts one image per request.

  4. Step 4

    Ask a grounded question

    Specify the OCR, localization, comparison, chart, STEM, or explanation task and request evidence tied to visible regions.

  5. Step 5

    Validate the response

    Check extracted text, labels, units, coordinates, calculations, and conclusions against the source image.

  6. Step 6

    Monitor the integration

    Track token usage, confirm the TokenHub endpoint in every environment, and compare HY-Vision-2.0-Instruct when speed matters more than deep reasoning.

Cost

Hunyuan Vision 1.5 Thinking pricing and free plan

TokenHub bills Hunyuan Vision 1.5 Thinking by input and output tokens. Tencent lists ¥3 per million input tokens and ¥9 per million output tokens. New TokenHub accounts can claim a one-million-token multimodal trial allowance valid for one year; the current trial campaign is listed through December 31, 2026.

New-user trial

1M tokens free

One-time TokenHub trial allowance for multimodal-understanding models.

  • Valid for one year after claiming
  • Available once per primary account
  • Current campaign listed through December 31, 2026

Pay-as-you-go input

¥3 per 1M tokens

Postpaid price for prompt text and encoded image input.

  • Billed hourly
  • Free allowance is consumed first
  • Postpaid billing must be enabled after the trial

Pay-as-you-go output

¥9 per 1M tokens

Postpaid price for generated response tokens.

  • Long reasoning uses more output tokens
  • Usage is visible in TokenHub
  • Prices are in Chinese yuan

Pricing checked . Check current pricing at the source ↗

Assessment

Hunyuan Vision 1.5 Thinking strengths and limitations

Where it stands out

  • Deep reasoning for image-grounded questions
  • Covers OCR, charts, visual localization, STEM problems, and image analysis
  • Supports more than 30 listed languages
  • OpenAI-compatible endpoint reduces client integration work
  • 40K context window and up to 24K output tokens
  • Low published token prices and a new-user trial

What to consider

  • Accepts one image per request and does not support video input
  • The old Hunyuan multimodal endpoint was retired, so older integrations require migration
  • A newer HY-Vision-2.0-Instruct model is available for fast-thinking image tasks
  • The GitHub repository still marks the technical report and downloadable checkpoints as forthcoming
  • Launch benchmark claims are creator-reported and date to 2025
  • OCR, chart values, spatial claims, and reasoning steps need source-image verification

Compare

Hunyuan Vision 1.5 Thinking alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Consumer

Gemini 3.1 Pro

Choose Gemini 3.1 Pro when a newer broadly multimodal flagship and Google's developer ecosystem are a better fit.

Explore Gemini 3.1 Pro

Consumer

Qwen3.5-Omni

Choose Qwen 3.5 Omni when one model must understand text, images, audio, and video across a broader language set.

Explore Qwen3.5-Omni

Project Management

Claude

Choose Claude when document-heavy visual analysis and Anthropic's general assistant workflow matter more than Tencent deployment.

Explore Claude

Questions

Hunyuan Vision 1.5 Thinking FAQs

What is Hunyuan Vision 1.5 Thinking?

It is Tencent's deep-reasoning vision-language model for image question answering, visual grounding, OCR, charts, photographed problems, and image-based creative analysis.

Is Hunyuan Vision 1.5 Thinking still available?

Yes. The old Hunyuan multimodal service ended in June 2026, but Tencent moved the model to TokenHub, where it remains listed under hunyuan-t1-vision-20250916.

How much does Hunyuan Vision 1.5 Thinking cost?

Tencent TokenHub currently lists ¥3 per million input tokens and ¥9 per million output tokens. Eligible new accounts can claim a one-million-token trial allowance valid for one year.

Can it analyze video?

No. Tencent lists image understanding but not video understanding for this model. HY-Vision-Video and YT-VITA are the relevant Tencent options for video inputs.

Is Hunyuan Vision 1.5 Thinking open source?

Not currently. Tencent announced an open-source plan in 2025, but the official repository still shows the report and checkpoints as forthcoming. The usable release is a hosted TokenHub API.

Bottom line

Our Hunyuan Vision 1.5 Thinking verdict

Hunyuan Vision 1.5 Thinking is a practical, inexpensive choice for Tencent Cloud users who need deliberate OCR, chart, STEM, and image-grounded reasoning. Compare it with HY-Vision-2.0-Instruct and broader multimodal models when multiple images, video, or global platform access are required.

Visit Hunyuan Vision 1.5 Thinking website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.