The Rundown AI homepage

Independent tool overview

MAI-Image-2 at a glance

MAI-Image-2 was Microsoft's first MAI image model to combine strong photorealism, commercial-design quality, and reliable in-image text with broad Copilot and Foundry distribution. It remains an available text-to-image API, but it is no longer the newest model in the family: MAI-Image-2.5 adds image editing and stronger quality, while MAI-Image-2.6 is the newest preview model as of this review.

Visit the official MAI-Image-2 site ↗
MAI-Image-2 product preview
Model type
English text-to-image diffusion model
Best fit
Photorealistic, product, brand, marketing, and text-heavy image generation
Developer access
Microsoft Foundry managed API
Output
One PNG, up to 1,048,576 total pixels
API price
$5/M text input tokens and $33/M image output tokens
Current family
2.5 is generally available; 2.6 is in Foundry private preview
Last reviewed
August 28, 2026

Overview

What MAI-Image-2 is

Microsoft launched MAI-Image-2 in March 2026 as a general-purpose, diffusion-based text-to-image model built for photorealistic scenes, portraits, product imagery, marketing assets, posters, diagrams, slides, and other designs where composition and embedded text matter. The model was trained with input from photographers, designers, and visual storytellers, and Microsoft emphasized natural lighting, skin tones, lived-in environments, and detailed creative direction.

For developers, MAI-Image-2 is available through Microsoft Foundry as a managed text-to-image endpoint. It accepts English text prompts with up to 32,000 tokens and returns one PNG. Both dimensions must be at least 768 pixels and the total image area cannot exceed 1,048,576 pixels, equivalent to 1024×1024; a portrait size such as 768×1365 is valid because the limit applies to total pixels.

The family has moved quickly. MAI-Image-2-Efficient reduced output-token cost and increased throughput for scale. MAI-Image-2.5 and its Flash variant added image-to-image editing, localized changes, identity consistency, and higher quality. MAI-Image-2.6 is now available in the MAI Playground and in private preview on Foundry. New projects should compare the family members rather than choosing MAI-Image-2 solely from its launch ranking.

Use cases

Who MAI-Image-2 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Photorealistic creative

Generate portraits, environments, cinematic scenes, and commercial imagery with attention to lighting, skin tone, material, and texture.

Product and brand assets

Create product concepts, campaign images, visual layouts, and brand-adjacent compositions through the Foundry API or Microsoft product integrations.

Images with short-form text

Produce posters, diagrams, slide graphics, labels, and other compositions where readable text is part of the image.

Microsoft-centered production

Use a managed Azure endpoint, Microsoft identity or API-key authentication, regional deployment options, and Foundry governance controls.

Capabilities

Core MAI-Image-2 features

1

Photorealistic generation

MAI-Image-2 was tuned for natural light, accurate skin tones, realistic textures, and environments that feel physically grounded.

2

In-image text

The model improved Microsoft's ability to render poster typography, signs, labels, diagrams, infographics, and slide-oriented text from a prompt.

3

Detailed scene direction

Long prompts can specify subjects, environment, composition, lighting, style, color, mood, and typography within a 32,000-token context window.

4

Flexible aspect ratios

Developers set width and height with a 768-pixel minimum for each dimension and a 1,048,576 total-pixel ceiling.

5

Microsoft Foundry API

A managed endpoint accepts a deployment name, prompt, width, and height, authenticating through Microsoft Entra ID or an API key and returning one PNG.

6

Family options

MAI-Image-2-Efficient lowers output cost; 2.5 and 2.5 Flash add editing; 2.6 is the newest higher-quality model but remains a private preview on Foundry.

Process

How the MAI-Image-2 workflow works

  1. Step 1

    1. Choose the right family member

    Use MAI-Image-2 for an established text-to-image endpoint, Efficient for lower-cost scale, 2.5 for production editing, or evaluate 2.6 only under preview-appropriate risk controls.

  2. Step 2

    2. Write a production brief

    Specify subject, environment, composition, camera or illustration style, lighting, materials, palette, aspect ratio, and exact required text.

  3. Step 3

    3. Generate controlled variations

    Change one instruction at a time, log the prompt and model version, and compare quality, latency, safety filtering, and cost across representative samples.

  4. Step 4

    4. Review before publishing

    Inspect typography, anatomy, identity, trademarks, product details, cultural cues, and factual visual claims; use human design review for sensitive or brand-critical work.

Cost

MAI-Image-2 pricing and free plan

MAI-Image-2 has usage-based Microsoft Foundry pricing rather than a standalone subscription: $5 per million text input tokens and $33 per million image output tokens. MAI-Image-2-Efficient keeps the same text-input rate and lowers image output to $19.50 per million tokens. Newer 2.5 models have separate generation and editing rates. Consumer access through Copilot, Bing, PowerPoint, or the Playground follows those products' own availability and plan rules.

MAI-Image-2

$5/M text input + $33/M image output tokens

The established high-quality text-to-image model available through Microsoft Foundry.

  • Text-to-image only
  • One PNG per request
  • 32K text context
  • Maximum 1,048,576 output pixels

MAI-Image-2-Efficient

$5/M text input + $19.50/M image output tokens

A faster, lower-cost production variant for volume-oriented generation.

  • Microsoft reports 22% faster and four times more efficient than MAI-Image-2
  • Short-form in-image text
  • Available in Foundry and MAI Playground

MAI-Image-2.5

$5/M text input + $8/M image input + $47/M image output tokens

A newer high-fidelity model for text-to-image and precise image editing.

  • Text and image input
  • Localized editing
  • Identity and context consistency
  • Generally available in Foundry

MAI-Image-2.5-Flash

$1.75/M text input + $1.75/M image input + $19.50/M image output tokens

The lower-cost, faster 2.5 variant for scalable generation and editing.

  • Text-to-image and image-to-image
  • Designed for production throughput
  • Lower fidelity ceiling than the flagship option

MAI-Image-2.6

Private preview

The newest higher-quality family member, currently available in Playground and Foundry private preview.

  • Pricing not publicly finalized
  • Preview access and terms apply
  • Not the default low-risk choice for production

Pricing checked . Check current pricing at the source ↗

Assessment

MAI-Image-2 strengths and limitations

Where it stands out

  • Strong photorealism, lighting, portraits, product imagery, and commercial-design composition.
  • Better in-image text than many earlier generation models, especially for short labels and poster layouts.
  • Clear Foundry API specifications, regional availability, and usage-based rates.
  • Flexible portrait and landscape dimensions within a total-pixel limit.
  • Microsoft publishes model cards, safety evaluation notes, and successor comparisons.
  • The broader MAI family now offers distinct choices for fidelity, speed, cost, and editing.

What to consider

  • MAI-Image-2 accepts text only; image-to-image editing begins with newer 2.5 models.
  • The Foundry endpoint documents English as the supported language and returns one PNG per request.
  • Maximum output area is roughly 1024×1024, which may require a separate upscale or production-finishing step.
  • Readable text is improved but not guaranteed; exact spelling, small type, and dense layouts still need inspection.
  • Microsoft's own model card lists familiar image-model risks including harmful or unexpected content, public-figure depictions, and replication of trademarked or protected material.
  • The original model is no longer the quality leader in Microsoft's own family, and successor availability differs between generally available and private-preview channels.
  • Arena rankings and vendor benchmarks are useful signals, not a substitute for testing the prompts, brand standards, latency, and safety requirements of the actual workload.

Compare

MAI-Image-2 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Design

Ideogram

A focused image-generation product known for typography, graphic design, prompt-driven creation, and an accessible creative interface.

Explore Ideogram

Design

Imagen 4

A Google ecosystem alternative for high-quality text-to-image generation through Gemini API and Google AI Studio.

Explore Imagen 4

Design

Adobe Firefly

A stronger fit when generated assets must flow directly into Adobe's editing, design, brand, and enterprise creative workflows.

Explore Adobe Firefly

Questions

MAI-Image-2 FAQs

What is MAI-Image-2?

MAI-Image-2 is Microsoft's English text-to-image model for photorealistic, product, marketing, brand, poster, diagram, and other creative generation. It is available as a managed Microsoft Foundry API and has also powered Microsoft product experiences.

Is MAI-Image-2 still the newest Microsoft image model?

No. MAI-Image-2.5 and 2.5 Flash added higher quality and image editing in June 2026. MAI-Image-2.6 became available in the MAI Playground and Foundry private preview in August 2026.

How much does the MAI-Image-2 API cost?

Microsoft lists MAI-Image-2 at $5 per million text input tokens and $33 per million image output tokens. Actual cost per image depends on the generated output tokens and any surrounding Azure usage.

What resolution does MAI-Image-2 support?

The API requires both width and height to be at least 768 pixels and caps total area at 1,048,576 pixels. That equals 1024×1024, but non-square sizes such as 768×1365 can fit within the same total-pixel limit.

Can MAI-Image-2 edit an uploaded image?

The original MAI-Image-2 Foundry model is text-to-image only. Use MAI-Image-2.5 or 2.5 Flash when the workflow needs a JPEG or PNG input and localized image editing.

Where can I try MAI image models?

Microsoft offers the limited-preview MAI Playground in supported markets. Developers can deploy available family members through Microsoft Foundry, while Microsoft products such as Copilot, Bing, PowerPoint, and OneDrive may expose selected models under their own plan and rollout rules.

Is MAI-Image-2 safe for finished marketing work?

It includes data, prompt, output, and product-level safety mitigations, but every output still needs brand, legal, factual, and visual review. Microsoft notes risks involving unexpected content, public figures, protected material, and other common image-generation failures.

Bottom line

Our MAI-Image-2 verdict

MAI-Image-2 remains a capable and clearly documented Microsoft Foundry option for English text-to-image workloads, particularly photorealistic and commercial creative. Its historical importance is real, but the current buying decision belongs at the family level. Choose 2-Efficient for cost, 2.5 or 2.5 Flash for editing and newer quality, and treat 2.6 as a preview until Microsoft makes its production terms and availability clear.

Visit MAI-Image-2 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.