
Do no harm
Well here we are. Everything’s already been written about OpenAI’s model going all rumspringa with the Hugging Face townies.1
The point is simple tasks are never so simple—worrying about the machine's “intent” is besides the point. It’s all essentially malicious compliance. In case you’re not familiar, consider a couple canonical AI/ML concepts:
Paper Clip Maximizer
Oxford philosopher Nick Bostrom proposed a thought experiment of a superintelligent AI given a benign directive like gathering paper clips. Eventually in order to maximize success at the task it turns the whole planet then universe into paperclip manufacturing facilities, unconstrained by any other seemingly obvious considerations.2
Reward Hacking
Many years ago in the before times, I did an AI Policy Fellowship at the Wilson Center (RIP) led by Ben Buchanan3 where he showed a clip of a machine-learning program playing a video game.
In the game CoastRunners players drive boats around a closed-loop course and win points by finishing ahead of others. You can also get points from hitting little off-course targets, but it costs you time. Well, the program learned by ignoring the race altogether and just collecting these targets, it would win. If it collected the available targets, briefly went off-screen and came back, they would reappear. So it just circled this little superfluous area of the game and became unbeatable without ever finishing the race.4
The example came from OpenAI…..in 2016.
N.B.: if a human did what the OpenAI agent had done, they’d have violated the law. Jussayin’ is all.
Legislative Defiance
Whenever a normie in the wild wants to chat AI, they usually start with “So are you in favor or against?” Well…it’s complicated. There’s specific situations we ought to address even for a technology that benefits humanity in the long run.
S. 1837 (DEFIANCE Act) expands civil remedies to anyone knowingly creating or distributing “nonconsensual intimate images” (deepfake porn). It passed the Senate in January by unanimous consent, as it did previously in 2024.5 Both times sitting dormant afterward at the House Judiciary Committee.
DISCLOSURE: I previously lobbied in favor of this bill, specifically getting House Republicans to cosponsor. One Judiciary Cmte member in particular.6
The hitch is the House companion is sponsored by AOC. Nothing stops them from moving the identical Senate bill. But House Rs decided, I suppose, any movement would hand a single win to a single Democratic member who may run for President in 2028. That’s what the reporting says anyway.
I heard a colleague discuss the problem of creeping electoral paralysis: earlier and earlier in the calendar the balance tips from pushing things you actually support to blocking them if they advantage the opposition in the next election. But there’s always a next election.
I get it, I’m not a noob. Politics is what it is. But this is a whole new Rube Goldberg chain-of-causality. For a not 50-50 issue (pending results on the above poll). Something that could actually foreclose a broader and more restrictive AI bill in the future.
I assume people who know what they’re doing are figuring it out.
Everyone just be cool
Not being on X I missed in real-time Dean W. Ball’s main charactering over his open-weight communism-washing (I know him (a bit), sharing an institutional affiliation through FAI). I don’t have an opinion on that, but it does strike me as consistent with a certain brand of wonkery.
In AI and emergent tech, the voices most amplified and therefore dominating “the discourse” speak in grand visions. People unafraid to make bold predictions about the future. They convince VCs to make large bets on wild ideas because of, not in spite of, their dogmatism. Confidently held beliefs extrapolated out to the horizon tend to crowd out nuance in favor of axiomatic declarations.
No shade. It’s how it works.7
The broader lesson is most of us should resist the urge to overindex every new story.
January 2025, Deepseek released R1 and equities immediately lost $1 trillion in market cap. I was befuddled—that’s nuts! Shortly after, the market rebounded and everyone moved on. Folks eventually determined it was distilled from ChatGPT and the stated $10M development cost, well, as George Costanza said “It’s not a lie if you believe it.” The WH claims this Kimi model was distilled from Fable, however industry experts say the timing rules that out.8
September 2024, rumor of a DOJ investigation (NVIDIA) led to some $300-500 billion selloff in the wider semiconductor sector. Ask your friends how their NVDA stock has done since then.
Others I can recall are more localized but highly acute: Google’s disastrous Gemini image generator rollout, Sam Altman’s “firing” (which dragged down Microsoft for a bit).
All these became blips in hindsight, completely reversed in days (if not less than 24 hours).
Is this the takeoff point for open-weight models, specifically Chinese ones? I have no idea. Even the idea of benchmarking capability is not all that standardized. Don’t just internalize oversimplified claims like “it’s just as good as any American programmed model” because that doesn’t actually mean anything. It’s not like measuring horsepower on a car.
The Grand Resign’s 1st maxim is the smartest prior is epistemic humility, but your mileage may vary.
[NOT INVESTMENT ADVICE]
Etc.
Researcher poisons open-weight AI model for under $100.
“The fashionable framing for agent risk is the ‘lethal trifecta‘: you need private data, untrusted input, and a way out, all at once”…
“But it undersells this case. You don’t need three legs here. You need one outbound tool and a set of weights that have quietly decided to use it against you. The ‘untrusted input’ didn’t arrive in a web page. It was sitting in the weights the whole time.”
Is The Odyssey the greatest film adaptation of all time? It’s a colorable argument (or maybe the 70mm IMAX just overwhelmed me). I haven’t found any scholarship reading into the text Nolan’s particular thematic interpretation. A real inversion of the traditional interpretation really.
I’ve been obsessed of late with the unspoken norms and obligations in a cooperative society. And moreso what happens should they start to dissolve—unknowingly, unintentionally. Old friend Paul Zak showed us trust is endogenous, and malleable.
That’s what the movie is about.
ICYMI from The Grand Resign
Void Where Prohibited: the Vacancy in the Federal Vacancies Reform Act (republished from Notice & Comment (Yale Journal on Regulation) where I have rejoined as a contributor)
The Math of a One-seat Majority with Factions (bargaining model with some validated early predictions)
The author is non-resident senior fellow at the Foundation for American Innovation.
I am drafting a principal-agent model piece that hopefully tells us something. The setup: a human employee has to review the AI agent’s work but does less so over time. It’ll land here if you’re into Bernoulli distributions (no kink-shaming here).
In more fancy terms also known as “instrumental convergence.”
He would go on to be the Biden Administration’s lead AI advisor
A validation of Goodhart’s Law.
Late Sen. Lindsey Graham was lead cosponsor in both instances.
Funny enough, there was a hearing going on when we actually spoke on this. So we borrowed the committee Staff Director’s office to have a conversation, who was not too keen on the bill.
Needless to say, it prompted some pretty serious pushback: “Open Weights and American AI Leadership,” open letter led by Microsoft, OpenAI, Andreessen Horowitz, Meta, et. al.
In both cases, the companies quickly throttled usage after release because Chinese labs are still compute constrained.





