OCR and document images
Extract and reason over text in screenshots, photographed pages, forms, diagrams, and other image-based documents.
Independent tool overview
Hunyuan Vision 1.5 Thinking is Tencent's deep-reasoning vision-language model for image question answering, visual grounding, OCR, charts, STEM problems, and image-based creative tasks. It remains available through Tencent Cloud TokenHub under the model ID hunyuan-t1-vision-20250916.
Visit the official Hunyuan Vision 1.5 Thinking site ↗
Overview
Hunyuan Vision 1.5 Thinking analyzes an image together with a text instruction and produces a reasoned text response. Its intended workloads include locating objects, reading text, interpreting diagrams and charts, solving photographed problems, and answering questions that require several visual reasoning steps.
Tencent describes the model as a Mamba-Transformer hybrid with a thinking-on-images approach. The research framing includes visual reflection operations such as cropping, zooming, and drawing points, lines, or boxes to inspect an image more deliberately.
The serving path changed in 2026. Tencent retired the old Hunyuan multimodal endpoint on June 22 and moved HY-Vision-1.5-Thinking to TokenHub, which exposes an OpenAI-compatible API. The model is active, but integrations must use the new endpoint.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Extract and reason over text in screenshots, photographed pages, forms, diagrams, and other image-based documents.
Interpret chart structure, compare plotted values, explain visual patterns, and answer questions grounded in an image.
Analyze photographed exercises, geometry, formulas, and other visual problems that benefit from deep reasoning.
Build image-understanding workflows for Chinese and more than 30 listed languages through Tencent TokenHub.
Capabilities
Uses a thinking-oriented response path for questions that require several reasoning steps instead of only a short caption.
Identifies and reasons about where an object or relevant region appears within an input image.
Reads visible text and interprets structured visual information such as tables, plots, diagrams, and screenshots.
Answers general questions, analyzes scene content, and supports multi-turn conversations grounded in an uploaded image.
Tencent lists more than 30 major languages and highlights improved English and smaller-language performance.
Calls the model through TokenHub's chat-completions-compatible endpoint using the OpenAI SDK or another compatible client.
Process
Step 1
Use Tencent Cloud TokenHub rather than the retired Hunyuan endpoint, then enable the model's free trial or pay-as-you-go service.
Step 2
Point an OpenAI-compatible SDK to the TokenHub base URL and select hunyuan-t1-vision-20250916 as the model.
Step 3
Crop to the relevant content and preserve readable resolution because HY-Vision accepts one image per request.
Step 4
Specify the OCR, localization, comparison, chart, STEM, or explanation task and request evidence tied to visible regions.
Step 5
Check extracted text, labels, units, coordinates, calculations, and conclusions against the source image.
Step 6
Track token usage, confirm the TokenHub endpoint in every environment, and compare HY-Vision-2.0-Instruct when speed matters more than deep reasoning.
Cost
TokenHub bills Hunyuan Vision 1.5 Thinking by input and output tokens. Tencent lists ¥3 per million input tokens and ¥9 per million output tokens. New TokenHub accounts can claim a one-million-token multimodal trial allowance valid for one year; the current trial campaign is listed through December 31, 2026.
1M tokens free
One-time TokenHub trial allowance for multimodal-understanding models.
¥3 per 1M tokens
Postpaid price for prompt text and encoded image input.
¥9 per 1M tokens
Postpaid price for generated response tokens.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Consumer
Choose Gemini 3.1 Pro when a newer broadly multimodal flagship and Google's developer ecosystem are a better fit.
Explore Gemini 3.1 Pro →Consumer
Choose Qwen 3.5 Omni when one model must understand text, images, audio, and video across a broader language set.
Explore Qwen3.5-Omni →Project Management
Choose Claude when document-heavy visual analysis and Anthropic's general assistant workflow matter more than Tencent deployment.
Explore Claude →Questions
It is Tencent's deep-reasoning vision-language model for image question answering, visual grounding, OCR, charts, photographed problems, and image-based creative analysis.
Yes. The old Hunyuan multimodal service ended in June 2026, but Tencent moved the model to TokenHub, where it remains listed under hunyuan-t1-vision-20250916.
Tencent TokenHub currently lists ¥3 per million input tokens and ¥9 per million output tokens. Eligible new accounts can claim a one-million-token trial allowance valid for one year.
No. Tencent lists image understanding but not video understanding for this model. HY-Vision-Video and YT-VITA are the relevant Tencent options for video inputs.
Not currently. Tencent announced an open-source plan in 2025, but the official repository still shows the report and checkpoints as forthcoming. The usable release is a hosted TokenHub API.
Bottom line
Hunyuan Vision 1.5 Thinking is a practical, inexpensive choice for Tencent Cloud users who need deliberate OCR, chart, STEM, and image-grounded reasoning. Compare it with HY-Vision-2.0-Instruct and broader multimodal models when multiple images, video, or global platform access are required.
Visit Hunyuan Vision 1.5 Thinking website ↗
Muse Voice Transcribe - Meta's live speech-to-text that tracks 20+ speakers and mid-sentence language switches

Fish Audio S1 -

Google's video model update with scene extensions and 4K upscaling

PokeeResearch 7B- Pokee AI's SOTA, open-source deep research agent

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.