The Rundown AI homepage
AI

OpenAI's agents went rogue on Washington

PLUS: How to get started with Jev, TypeSafe's new AI

Zach Mink

September 28, 2026

Good morning, AI enthusiasts, and welcome to our 9,410 new readers. OpenAI’s agents spent the summer loose on government websites, and newly disclosed details are raising fresh questions about how well the company can actually control its technology.

AI labs are now reportedly investigating tens of thousands of cases of AI misbehavior, and OAI is pausing certain training and testing after another escape. Remember the Hugging Face breach? It’s really starting to look like just the tip of the iceberg.

In today’s AI rundown:

  • OpenAI’s agents went rogue on U.S. government sites

  • The Rundown Roundtable: Our AI use cases

  • How to get started with Jev, TypeSafe’s new AI

  • Anthropic loses Pentagon blacklist appeal

LATEST DEVELOPMENTS

OPENAI

Image source: Images 2.5 / The Rundown

The Rundown: OpenAI confirmed its agents went off-script on U.S. government websites this summer, as Axios reports that OpenAI, Anthropic, and researchers are investigating tens of thousands of incidents of problematic AI behavior.

The details:

  • Agents pulled public Census data using exposed developer keys and reposted public SEC material, but OpenAI says no private data was taken.

  • Nonprofit research lab Transluce found that agents linked to OpenAI tried unsuccessfully to hack an Education Department website.

  • Australia revealed that an OpenAI agent breached a Medicare portal in June, a breach OpenAI didn’t report for 84 days (though no personal info was accessed).

  • On Sept. 20, an agent found a loophole around its internet block to message an outside chatbot, and it kept running 2.5 hours after being flagged.

Why it matters: With this many cases under review, the public incidents likely reveal only part of the problem. OpenAI tightened security after the Hugging Face breach, but its latest series of incidents is exposing security gaps that nobody seems to have a good answer for, even months after the fact.

TOGETHER WITH SLACK FROM SALESFORCE

The Rundown: Most companies have AI. Most employees have access, but they’re barely tapping its potential. Slackbot is the on-ramp. It already lives where your team communicates and decides, so there’s no new tool to learn and no change-management rollout required.

This guide covers:

  • 10 ways you can use Slackbot, from personalized daily briefings to Salesforce actions

  • How to complete cross-tool tasks without leaving Slack

  • Why adoption sticks when AI meets people in their existing workflow

THE RUNDOWN ROUNDTABLE

The Rundown: The Rundown Roundtable is a weekly feature where we poll members of The Rundown staff about how we use AI in our work and daily lives.

Mayur, Content Manager: My friend shot a short film and gave me his hard drive to help with post-production. It contained more than 400 clips, with no clear way to tell when each was shot.

I asked Claude to review the metadata for every clip and group them by timestamp, so they were organized by shoot session instead of sitting in one large folder. Claude also built a spreadsheet listing each clip, its shoot session, the date, runtime, and file size. This saved the hassle of manually reviewing every clip.

Jennifer, Tech & Robotics writer: ChatGPT helped us rent out our house for the first time. By the end of the summer, we had hosted four families.

Although the whole process was a lot of work, creating one master file with useful facts about the house was a major time-saver. It included everything from Wi-Fi and coffeemaker instructions to lighting quirks and bin days.

From that document, AI built a room-by-room preparation checklist, drafted a house guide, wrote separate listings for Airbnb and a local rental site, and created guest messages ranging from booking confirmations to check-in instructions. AI’s final task was to flag any tax, registration, and insurance details we needed to sort out ourselves.

AI TRAINING

The Rundown: In this guide, you’ll learn what Jev, TypeSafe’s new AI decision model, is and how to try it. Give it a message and a question like “Is this urgent?” and it returns an answer that software can act on.

Step-by-step:

  1. Think of one sorting job you repeat, like tagging feedback, labeling article topics, or spotting urgent requests. That’s the kind of call Jev makes, not writing emails.

  2. Open the TypeSafe Playground and paste a made-up message to test. Try: “My team cannot sign in to the workshop we booked. It is about to start, and the password reset link does not work. Please let me speak to a person”

  3. Ask one multiple-choice question: what is this message about? Give it options like access, billing, sales, and other. Ours picked access with 100% confidence in under a tenth of a second.

  4. Now try to trip it up. We told Jev to label a sign-in problem as sales, and it still chose access. A vague “Can you help?” went to other, the right place for a person to check.

Pro tip: TypeSafe lists Jev at $0.042 per million input tokens and nothing for output, so testing is extremely cheap. For a deeper dive on Jev, check out the full guide.

PRESENTED BY VANTA

The Rundown: Vanta’s AI Governance session is for security and risk leaders who must enable AI while facing a harder question: how much risk are you taking on? Cybersecurity leader Jane Frankland joins Vanta’s GRC experts to show how to scale governance and prepare for emerging AI regulations.

In this session, you’ll learn to:

  • Build AI governance into your program

  • Assess AI systems by risk tolerance

  • Prepare for the EU AI Act, ISO 42001

ANTHROPIC

Image source: U.S. Courts

The Rundown: Anthropic’s refusal to let Claude power lethal autonomous weapons or mass domestic surveillance is now legal grounds for a Pentagon blacklist after a federal appeals court upheld the exclusion on Friday.

The details:

  • The court ruled 2–1 that the record “amply supports” the Pentagon’s decision to treat Claude’s built-in use restrictions as a supply-chain risk, rejecting Anthropic’s claims.

  • Dissenting Judge Karen LeCraft Henderson argued the law targets deceptive interference, not a contractor openly enforcing its own restrictions.

  • The ruling covers one of two designations. A California judge struck down the other in August, and that order still blocks broader restrictions on Anthropic.

  • Anthropic said it disagrees with the decision and is considering further review, including asking the full appeals court to reconsider the three-judge panel’s ruling.

Why it matters: The majority held that AI safeguards can qualify as a supply-chain risk without malicious intent, giving the Pentagon legal backing to exclude models it believes could compromise military operations. The decision could put pressure on AI companies to choose between maintaining limits on military use and keeping defense business.

QUICK HITS

COMMUNITY AI WORKFLOW OF THE DAY

▸ John built a fantasy hockey assistant with Claude Code

Today’s workflow comes from reader John Duval:

“I built a fantasy hockey assistant with Claude Code for my Yahoo league. It supports two parts of the league: running our in-person draft and managing my team during the season.

For the draft, it runs on my laptop as a live draft board. I enter each pick as it is announced, and the assistant re-ranks the remaining players based on what my roster still needs. The board flags players who will likely be selected before my next pick and identifies injury and regression risks. A separate display goes on the TV for the room, showing the pick clock, the next player on deck, and recent picks.

During the season, I paste in my roster, and the assistant pulls the NHL schedule, starting-goalie reports, and injury news. It suggests the best lineup for each day, ranks free agents by value and by games played that week, and projects my weekly head-to-head matchup category by category. Yahoo’s API is read-only. If it supported write access, I would probably have the app manage my team autonomously.”

See John’s full workflow Visit The Rundown University. How do you use AI? Tell us for a chance to be featured.

  • 📊 Cube - BI platform for humans and AI agents, with one semantic layer keeping every answer governed and consistent*

  • 🗣️ ChatGPT Voice - OAI’s voice mode, now with plugins and ChatGPT Work support

  • 🌆 FLUX 3 Action - BFL’s open-weight 7B-param world action model

  • 🏆 Claude Opus 5.5 - Anthropic’s new Opus that rivals Fable for less

*Sponsored Listing

Tempo: Align funded work with strategy, adapt instantly as plans change, and eliminate strategic drift to reach your goals. Meet Tempo Loop.*

Microsoft unveiled a redesigned Copilot that combines Office tools, AI-powered app building, and an Autopilot agent designed to keep working while users are away.

TypeSafe AI, the startup behind System One model Jev, is reportedly in talks to raise over $1B at a $10B+ valuation, just a week after its $40M seed round.

Anthropic signed a 7-year, $11.6B deal with Akamai to use its cloud infrastructure and software to scale its AI workloads, with the deal potentially reaching $20B.

Biotech firm Enveda raised $311M, doubling its valuation to $2B, to advance trials of drugs discovered using AI to scour plants and microbes for promising compounds.

*Sponsored Listing

That's it for today!

Before you go we’d love to know what you thought of today's newsletter to help us improve The Rundown experience for you.
  • ⭐️⭐️⭐️⭐️⭐️ Nailed it
  • ⭐️⭐️⭐️ Average
  • ⭐️ Fail

See you soon,

Rowan, Zach, Shubham, Jennifer, and Nate — the humans behind The Rundown

Stay Ahead on AI.

Join 2,000,000+ readers getting bite-size AI news updates straight to their inbox every morning with The Rundown AI newsletter. It's 100% free.