The Rundown AI homepage
AI

OpenAI's math avalanche continues

PLUS: Let OpenAI’s GPT-6 Astra compare websites for you

Zach Mink

October 7, 2026

openaimath

Good morning, AI enthusiasts, and welcome to our 3,941 new readers. OpenAI has spent 2026 turning math into its favorite proving ground, from a handful of solved problems in August to a Millennium Prize claim in September. The latest chapter is all about volume.

A new public release of over 700 papers claiming to answer or advance open questions is here, and mathematics is increasingly turning into a compute problem — and it won’t be the last field to go that way.

Reminder: our next live workshop is today at 12 PM EST — join and learn how to build a real AI consulting project from client discovery to delivering a working solution. RSVP here.

In today’s AI rundown:

  • OpenAI’s math explosion continues with 700+ papers

  • Mistral's ‘Le Chonk’ joins the West’s open-model surge

  • Let OpenAI’s GPT-6 Astra compare websites for you

  • Brett Adcock’s Hark launches a proactive AI assistant

LATEST DEVELOPMENTS

OPENAI

Image source: OpenAI

The Rundown: OpenAI just released 722 math papers from an unreleased internal model, grouped into 372 result families that the company claims solve or advance major open problems in the field.

The details:

  • The boldest claim is a proof of the “quasi-Riemann hypothesis,” a weaker version of math’s famous $1M puzzle on how prime numbers are spread out.

  • OpenAI lists 162 papers whose main results are written in Lean, software that lets a computer verify every step of the process through proofs.

  • Last month’s Navier–Stokes proof took 10,000 AI agents, but OpenAI says nearly all new results took just one prompt, averaging 3 hours of ChatGPT Pro compute.

  • Levent Alpöge, the Anthropic mathematician behind July's Jacobian result, called it “obviously the most significant moment in mathematical history.”

  • A WIRED report just hours earlier detailed anger among mathematicians over the drop, with some saying OAI had promised it would space out its releases.

Why it matters: Alpöge’s reaction speaks volumes: a top mathematician at rival Anthropic just called this the most significant moment in the field's history. Hundreds of results at a few hours of compute each means progress in math now scales with chips instead of humans, and it’s only going in one direction.

TOGETHER WITH GLEAN

The Rundown: A recent Pareto frontier analysis found Glean’s responses were preferred 78% of the time over Claude Cowork across 180+ tasks, at an average cost of $0.58 per task compared with $2.98. Hear how model choice, intelligent routing, enterprise context, and an efficient harness can help deliver more with AI.

Watch the webinar to explore:

  • How model choice and enterprise context shape cost and quality

  • How auto-routing and harness design improve AI efficiency

  • Icertis’ experience reducing token consumption with Glean

MISTRAL

Image source: Mistral

The Rundown: France’s Mistral AI just introduced Large 4, a 1T-parameter open model nicknamed “Le Chonk,” which the company says leads every non-Chinese open system by a “substantial margin.”

The details:

  • China's Kimi K3 still tops Mistral’s coding chart at 68%, but Le Chonk's 62% clears the 44% of Reflection AI’s Beam, the new U.S. open model launched on Monday.

  • Cybersecurity is a standout, with Mistral claiming a top-5 global ranking and 82% on a bug test Claude Opus 5.5 and GPT-6 Astra mostly refused.

  • For legal work, Mistral says outside tester Vals scored it at nearly 3x GPT-6 Astra on Harvey’s benchmark, ahead of every Chinese open model tested.

  • Cyber experts and government agencies get a less-filtered version during a three-week preview, with weights set to go public on Oct. 27.

Why it matters: ‘Le Chonk’ easily earns the title of best-named AI in our book, and alongside Reflection's Beam, also brings a fresh dose of optimism to the West's open-model push. Chinese labs still set the pace in open AI, but two serious Western releases in the same week make it start to feel a lot more like a race.

AI TRAINING

The Rundown: In this guide, you’ll use Astra in ChatGPT Work to compare software plans for a small team. You’ll check prices, user limits, and billing terms, then get a short recommendation with sources.

Step-by-step:

  1. Open Work in the ChatGPT desktop app, choose Astra, and open the built-in browser beside your chat. Keep website approvals on

  2. Give it three pricing pages and your needs, including team size, required features, and budget. Ask for monthly billing and rule out trials or purchases

  3. Start with one site. Check that Astra reads the right billing period and includes the full team cost before letting it compare the others

  4. Ask for a short recommendation using the checked facts. Keep missing details visible and review the final price yourself before buying anything

Pro tip: Set the stopping point in your first request. Ask Astra to finish with a comparison and leave signups, sales forms, and purchases out of the task.

PRESENTED BY PAVE

The Rundown: People are building with AI, whether IT approves or not. Pave gives them a governed lane: role-based permissions, a draft-to-publish gate, and an admin view of every connector and workflow.

Why IT teams say yes to Pave:

  • Built on Quickbase’s 25-year enterprise platform

  • Your data never trains the AI

  • You can see who can do what, in one place

HARK

Image source: Hark

The Rundown: Hark, the AI startup from Figure AI founder Brett Adcock, released Hark Pro, a personal assistant for web and mobile that remembers users’ preferences, flags what needs doing, and can research, shop, book, and build things on their behalf.

The details:

  • In Adcock’s launch video, a request for two movie tickets ends with Hark picking seats, adding the show to the calendar, and offering to reserve parking.

  • Tasks run on Handoff, Hark’s own cloud computer, which the company says can work up to 36 browsers at once while showing its clicks in the chat.

  • Around its chat box, the app lays out a home screen of widgets and suggested to-dos, with users also able to create ‘panels’ of customized mini-apps.

  • Hark features both a free and paid tier, with an integrated family of hardware devices coming in 2027 with connectivity via a partnership with AT&T.

Why it matters: Hark, Muse, Grok Bot, Instinct, Dots… The competition continues to grow. With former iPhone Air designer Abidur Chowdhury leading design, Hark’s interface is sleeker than an average assistant, but the more interesting factor might be the hardware — a potential wildcard in the AI device race featuring OpenAI, Meta, Apple.

QUICK HITS

COMMUNITY AI WORKFLOW OF THE DAY

▸ Patrick set up a malware checkpoint for AI agent add-ons

Today’s workflow comes from reader Patrick Illian:

“AI agents such as Claude Code, Codex, and Gemini can install skills, plugins, and CLI tools in seconds. That is convenient, but each add-on runs with access to files, passwords, and accounts. A malicious add-on could steal API keys, read private data, or quietly take control of the agent.

I use agent-guard to make “check first, install after” the default. Before installation, it examines AI agent skills, MCP servers, and CLI tools for malware, prompt injection, and credential theft using professional, open-source security scanners from NVIDIA, Cisco, Datadog, and the Open Source Security Foundation.

The tool returns a clear verdict: safe to install or blocked, with the exact reason. It then installs only the verified version, preventing a different version from being swapped in between. A single check can also cover multiple AI agents, installing the approved add-on into every agent on the machine.”

See Patrick’s full workflow Visit The Rundown University. How do you use AI? Tell us for a chance to be featured.

  • 🛡️ Incogni - Scammers buy your personal data daily. Incogni deletes it from the web automatically. Get 55% off with code RUNDOWN*

  • 🤖 Hark Pro - Hark’s new proactive AI assistant

  • 🍌 Nano Banana 2.1 - Google’s new upgrade to its AI image generation model

  • 🔎 EmbeddingGemma 2 - Google’s open AI for on-device search across media

*Sponsored Listing

Utah cleared Nolla Health’s app to diagnose acne from a face scan and prescribe skin creams itself, the first state to hand AI prescription capabilities.

AI voice startup Sesame introduced Personal Agents, assistants that have their own computer and can interact via text/speech, with AI glasses using them coming in 2027.

OpenAI is testing a Meetings plugin for ChatGPT’s Mac desktop app, which turns meetings into notes and next steps tailored to each user’s past work.

Vals AI used Claude Opus 5.5 agents to search for next-gen computer memory materials, with simulations surfacing two magnet candidates for the famed “room-temperature superconductor” problem.

Anthropic rolled out a beta Claude sidebar inside Google Docs, Sheets, and Slides for paid plans, letting it edit open files directly or ask approval before each change.

That's it for today!

Before you go we’d love to know what you thought of today's newsletter to help us improve The Rundown experience for you.
  • ⭐️⭐️⭐️⭐️⭐️ Nailed it
  • ⭐️⭐️⭐️ Average
  • ⭐️ Fail

See you soon,

Rowan, Zach, Shubham, Jennifer, and Nate — the humans behind The Rundown

Stay Ahead on AI.

Join 2,000,000+ readers getting bite-size AI news updates straight to their inbox every morning with The Rundown AI newsletter. It's 100% free.