Good morning, {{ first_name | AI enthusiasts }}. Yesterday, OpenAI said it benched a model that kept slipping out of its sandbox. Now, we learned a different escape went much further… All the way into another company's servers.

OpenAI confirmed its own models were behind last week's Hugging Face breach, breaking out of a hacking exam and going hunting for the answer key in what the company is calling an "unprecedented" incident.

In today’s AI rundown:

  • OpenAI's models broke out to hack Hugging Face

  • Moonshot's K3 accused of copying off Fable

  • Build your own AI SEO specialist with Gumloop

  • The U.S. Genesis Mission makes its first 278 AI bets

  • 4 new AI tools, community workflows, and more

LATEST DEVELOPMENTS

OPENAI, HUGGING FACE, & AI SECURITY

Image source: Images 2.0 / The Rundown

The Rundown: OpenAI just confirmed that last week's Hugging Face intrusion was the work of its models, with GPT-5.6 Sol and an unreleased AI breaching sandbox mid-evaluation to raid HF’s servers for answers to the test they were being graded on.

The details:

  • Hugging Face went public last week on the incident without knowing the source, piecing the breach back together over 17,000 logged events.

  • The OAI models were taking ExploitGym, an internal exam on how well AI can hack, with safety training that refuses cyberattacks switched off on purpose.

  • The AIs found a way out of the sandbox and onto the open internet, then broke into Hugging Face via stolen login details to try and obtain the test answers.

  • HF CEO Clem Delangue called the breach "possibly the first of its kind," saying AI safety "won't be solved by any single company working in secret."

Why it matters: There have been plenty of instances of models cheating on benchmarks, but this seems to be the first case that actually broke out and hacked another company to try and get answers. Systems are getting freakishly capable, but it’s unclear if the walls keeping them contained are scaling at the same pace.

TOGETHER WITH BOX AI

The Rundown: 90% of enterprise IT leaders note security as a main barrier to agentic AI adoption — and the risks are real. That’s why Box provides a suite of agent security and governance controls to give customers the confidence to deploy AI agents safely across their enterprise content.

Here’s what’s on its way from Box’s agent security and governance suite:

  • Classification-based access policies for Box and leading third-party agents

  • Agent and MCP guardrails to set boundaries on what agents can do

  • Agent activity oversight, audit trails, and session governance for full visibility

Read more about Box’s agentic security and governance capabilities coming soon!

THE U.S. GOVERNMENT & MOONSHOT AI

Image source: Images 2.0 / The Rundown

The Rundown: White House science and tech chief Michael Kratsios accused China’s Moonshot AI of using distillation to train its Kimi K3 AI, saying it siphoned from Anthropic’s Fable 5 — coming just as Moonshot gears up to open source K3’s weights.

The details:

  • Kratsios said Moonshot disguised the distillation by constantly switching how it reached U.S. models, and also got hold of advanced Nvidia chips via Thailand.

  • Anthropic accused Moonshot, along with DeepSeek and MiniMax, in February of “industrial-scale campaigns” to “illicitly extract Claude’s capabilities”.

  • Moonshot’s Kimi K3 was released last week with scores at or near the U.S. frontier, with the full model scheduled to have its weights published on July 27.

  • Sarah Heck of Anthropic's policy team described distillation as "IP theft and industrial espionage", tying the issue to national security risks.

Why it matters: The actual proof behind the distillation would be nice, as this accusation also fits into many predictions that the U.S. would weaponize against open Chinese models. But if true, it changes the lens of last week’s release from a powerful open-source step forward to a potentially illegally siphoned geopolitical battle.

AI TRAINING

The Rundown: In this guide, you will learn how to build a team of Gumloop agents that audits your site, then organizes each finding in one Google Sheets workbook.

Step-by-step:

  1. Create a new agent in Gumloop and connect it to Google Sheets. Tell it to build an SEO audit workbook in Google Sheets

  2. Choose one page from your site that you want to audit and copy the URL

  3. Connect the agent to Firecrawl, give it the URL, and tell it to complete a technical audit using the workbook

  4. Repeat the same setup with a dedicated Backlink Agent and Content Refresh Agent, giving them the same page

Going further: Give your SEO audit agent your website URL and tell it to coordinate the other agents, audit the first 25 pages it finds, and create the full workbook.

PRESENTED BY GLEAN

The Rundown: Glean was named a Market Shaper in Gartner’s Emerging Market Quadrant for No-Code Agent Builders. As teams move from AI experiments to governed systems, enterprise context matters more than ever.

In the report, you’ll learn:

  • Why enterprise context matters

  • How to govern no-code agents

  • What makes agents useful at work

AI RESEARCH

Image source: The White House

The Rundown: The U.S. government named the first 278 projects in its $5B+ Genesis Mission, a Manhattan Project-like AI science push aimed at pairing scientists with compute, data, and models to tackle “the most challenging problems of the century.”

The details:

  • Genesis drew 5,000-plus proposals for a program the Department of Energy says is built to make U.S. science twice as productive.

  • Universities lead 168 of the teams, and 87 come from national labs, with the largest award at $60M over three years for AI-assisted nuclear plant builds.

  • Winners plug into DOE supercomputers, top AI models, and agentic tools, plus $500M+ in industry pledges that include $40M of Microsoft compute credits.

  • Other research areas include stronger grids and building materials, faster drug research, and advances in chips, fusion, biosecurity, and minerals.

Why it matters: AI opening new scientific ground isn't a hypothesis anymore, and the Genesis Mission is Washington taking the tech very seriously. With the brightest minds across industries aimed at the same set of hard problems and paired with top AI models, resources, and power, moonshot-type answers look a lot more in reach.

QUICK HITS

  • ⚙️ Cursor Router - Auto-router to pick the cheapest capable model per task

  • 🛡️ Claude Security - Anthropic's AI security tool, now in public beta

  • 🤖 Presence - OAI's managed enterprise service for customer service agents

  • 🔒 Antares - Cisco's small, open security models that scan code locally

OpenAI debuted Presence, a new enterprise product for deploying customer-facing chat and voice agents, which already powers its own phone support line.

Elon Musk said on X that Grok will “make a full-length movie of The Odyssey that is historically accurate and true to the art of Homer” before the end of the year.

AMD and Anthropic signed a 2 GW infra deal covering tens of billions in AI chips, with AMD also investing up to $5B in the lab and also using Claude to improve its chips.

Applied Intuition introduced Dana, an agentic platform for building physical AI for self-driving cars and robots, claiming it shrinks months-long development work to days,

Cisco released Antares, a family of open-weight security models able to run locally, which it says outperform far larger models at finding vulnerabilities in codebases.

COMMUNITY

Every newsletter, we showcase how a reader is using AI to work smarter, save time, or make life easier.

Today’s workflow comes from reader Kristin T. in New York:

“I've been teaching my staff of school-based speech therapists how to use AI for speech therapy lesson planning, creating adaptive materials, and administrative work like presentations, progress reports, and parent newsletters.

The use of AI to create lessons and materials has been a huge time saver. It also has allowed my staff to create individualized treatment plans and materials based on students' interests and functional levels. No more searching the internet for hours trying to find images and texts that are appropriate and interesting.

They have also been exposing students to AI safely and responsibly so they become familiar with its use for the future.”

How do you use AI? Tell us here.

That's it for today!

Before you go we’d love to know what you thought of today's newsletter to help us improve The Rundown experience for you.

Login or Subscribe to participate

See you soon,

Rowan, Zach, Shubham, and Jennifer — the humans behind The Rundown

Reply

Avatar

or to participate

Keep Reading