Cool Beans
Google’s newest AI app Dreambeans described as everything from silly to creepy.
Sandbox Irregularities
Remember a couple weeks ago when OpenAI’s new model broke containment and went all rumspringa with the Hugging Face townies? Well OpenAI wants to come clean about some other incidents.
They work with third-party testers, and one of them committed a “testing-environment misconfiguration” (the software equivalent of a wardrobe malfunction). OpenAI wants to make sure you know it was totally not OpenAI’s fault. Instead Irregular (yah, that’s the contractor’s name) simulated a target that:
“unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment. This did not involve a sophisticated sandbox escape.”
So that sounds like two mistakes to me. BUT DON'T WORRY. EVERYTHING’S COOL - IT WASN’T ONE OF THOSE SANDBOX THINGS YOU’VE ALL BEEN HEARING ABOUT. (ed. Whew that’s a relief)
Look, it's not any major new risk—it’s the kind of thing that happens all the time. But it ought to be relevant for all the people advocating strongly for “independent verification organizations” (IVOs) as a solution to risk mitigation (e.g., Fathom).1 Because while that approach is plausible, and one I generally support, their blind spots mean they shouldn’t be driving at night. In this case even if we ignore incentive misalignment, it assumes an available supply of sufficient expertise.
OAI and the leading labs don’t collaborate with tech bumpkins—these are among the most credible testers and evaluators in the world. And even they make mistakes. What happens when an IVO working in good faith misses the “Internet OFF” button but still slaps an Inspected by IVO No. 3 sticker on it and it goes full send? That’s not fatal, and no system is without mistakes, but it’s foreseeable and already quite real.
+ OpenAI’s post is quite deliberate in deflecting blame…why would that change in the future? How would us mere mortals even know if they’re right?
Microfoundations for Macro AI
The above is why I’ve become obsessed in modeling among other things simple scenarios of AI performance. Which I intend to do more of, pending funding.
The point is to use well-established economic machinery to model, predict, and test, what happens when humans be human-ing but now they have AI sidekicks: in their job, in the market, wherever.
SO MUCH of AI and related policy chatter is so obtuse, so unfalsifiable, and yet so incredibly confident. It’s vague claims and extreme conclusions based on what the French call a certain I don’t know what. Certainly not policy experience.
Nerds can be pretty insightful, but the annoying thing about this particular vintage is it comes cloaked in completely un-self-aware hubris. And that would normally be off-putting but apparently makes you a “thought leader” [*gag me*].
The worst offenders are the a16z types. Running around yelling Regulatory capture! Regulatory capture! but think George Stigler is the guy who works at the corner deli.2
Most of the time, I’m actually trying to engage a kernel of truth in an idea, and a somewhat critical examination is an attempt to make a good idea better.
…to stress test it, put it in the “sandbox” and see if it can escape.
Because lots of otherwise worthwhile ideas are built on a heroic amount of hand-waving:
Someone will write the specific law or regulation to implement this. Someone will negotiate around political opposition. The needed data will appear. Right?
Those are Atlas-level load-bearing assumptions.
Trust me: none of that shit just works itself out. But it can be done, and a useable playbook assembled ahead of time. An offensive coordinator doesn’t say “The wideout gets downfield then touchdown. The offense will figure out the rest.” But that’s what ya’ll sound like.
Claude FLASH SALE!⚡
Elsewhere I flagged what I thought would be a bigger deal: a product the US government flagged as a significant “supply-chain risk” is still available for sale to federal agencies. Not just for sale, but for “$1 from now until September.” Pretty smokin’ deal for software banned “to protect national security.”
As a reminder, Anthropic’s court challenge to the Trump Admin.’s security ban is ongoing, and a federal judge asked the government how many agencies are still using the AI software. To which the gov’t lawyers basically put their hands in their pockets, looked sheepishly at their shoes, and said, “um, we would totally tell you, but it’s, like, super sensitive security info and stuff.”
Maybe GSA can answer. Funny because it’s sad. (ed. the word you’re looking for is “tragicomic”).
Etc.
From The Grand Resign
Notice & Comment (Yale Journal on Regulation) where I recently returned as contributor:
Void Where Prohibited - the Catch-22 in the Federal Vacancies Reform Act (the law governing “acting” officials).
The Regulatory Three-Body Problem - how do administrations actually make regulatory decisions?
Hey ChatGPT-5, you good? - follow-up to an essay describing why an improving AI agent at work isn’t what it looks like. Validated by this story that Microsoft’s own programming is being weighed down by “AI-assisted coding and agentic workflows” leading to major outages. It turns out, as the model predicts, human coders aren’t reviewing the AI closely enough.
Infrastructure Permitting Reform finally arrives! - in 2016, that is. But it saved $53 billion!
Ho-hum
3.7M patients had their medical records stolen in data breach. “CareCloud handles a large amount of patient data and billing information on behalf of hospitals, doctor’s offices, and other medical practices.”
Tell me again how I’m a doomer.
They Only Need Once
Jordan Schneider and Sebastian Mallaby (Machine Readable) over at ChinaTalk reminded me of a totally remarkable event I think most of us forgot. In ‘84, an IRA bomb would've killed Prime Minister Thatcher at her Brighton hotel, but for her working late and walking to the bathroom (still 5 killed, 30 injured).
Afterwards the attackers noted she was lucky but they only had to be “lucky once.” The asymmetry of cyberattacks versus defenses….it’s structurally the same.
The author is nonresident senior fellow at FAI
They gather all the AI cool kids at their conferences but as far as I can tell never wrestle with some pretty basic new institutional economics or IO. But their CEO knows cannabis regulation (less than this guy tho).
By the way, as I’m sure they’re aware, Stigler’s 1971 article focuses on price floors and occupational licensing. Is that the worry here? Price floors? And no doubt they’re familiar with the Peltzman generalization and Becker response.



