Preach
My friend Brad Carson1 recently joined the Machine Learning Street Talk podcast. A few (of many) thoughtful framings:
“Human-in-the-loop” as meaningful agency in warfare is a fiction: If the AI identifies an enemy combatant with 0.82 confidence, the soldier pulls the trigger every time. Without AI assistance, a soldier may make a mistake but we have a long history of military and other norms to know how to treat that. The AI targeting inevitably displaces whatever human judgement wrestled with uncertainty without it.
The “It’s just a tool, neither good or bad” argument: in some product categories, manufacturers assume strict liability. If some knucklehead blows up messing with a toaster or something, there are scenarios where the manufacturer pays up. Others where liability applies jointly. The point is liability risk isn't limited to “uses.”2
I say it a lot but he’s more eloquent: whatever the eventual beneficial impacts of AI (and I count myself among believers) if people broadly reject it, it doesn’t really matter. Political incentives are not put on pause while “the optimal” future unfolds, because democracy. The tech sector hasn’t cultivated a friendly playing field here (even if some of them get to watch from the owner’s box). Our moral technology is stubborn and slow to change.
Concession
[UPDATE: LaTeX not rendering properly for some reason]
I reactively dust AI Utopians for dropping irrefutable claims. The future to be fair, it really do be like that sometimes.
In spirit of fairness, one of my own similarly vulnerable claims is:
Clear, sensible rules encourage adoption.
(ed. “Sensible” is doing a lot of work there)
I’m asserting regulation is not always suppressive to emerging technologies, by (1) lowering transaction costs, (2) facilitating adjunct markets (e.g., insurance, specialized training), and (3) speeding interoperability (read: recombination of ideas).
Think of it as a modified Romer Growth Model where a portion of R&D labor is diverted to reducing “translational drag”: reconciling differing standards, bespoke due diligence, equalizing metrics, etc.3
So it’s not totally irrefutable; there’s an empirical door we can walk through. Though if aviation and GAAP standards are supporting examples, there’s plenty that go the other way—and I should acknowledge that logical weakness.
Paradox
Speaking of adoption frictions…my new car led me to conclude:
For any added “smart,” “automatic,” or “self-adjusting” product feature, a high but not perfect performance rate is worse than not having the feature at all.
Above some threshold, habituation/attention decay takes hold and you forget the embedded efficiency gain.
That is, once in a while the feature to which you’ve now adapted slaps you upside your head. Ex: My car unlocks as I approach with my phone. So I stop carrying keys. But then…
This G%d@MN car has me running back into my F*@K>’N house, where did I leave…? (“WE’RE ALREADY LATE!!!”)[hands full]…GAH!!!
The closer a feature gets to 100% execution the harder it is to fully trust.4 Is it possible this feature has on net provided me some utility? Sure….but:
↑ E(Reliability) → ↑ PsychicCost(Failure)
This article circulating widely says it's exactly what happened at a bank. An autonomous AI agent:
“…approved a $1.4 million commercial line of credit. No human reviewed it. No human knew it had happened until the morning standup six hours later, when a junior analyst flagged it. The decision was correct—the borrower met every credit criterion. She flagged it because nobody was sure exactly who was accountable for it.
Six weeks later, the agent was rolled back to a lower autonomy tier. Not for a model reason or for a regulatory one. The decision was operational: The people supervising the agent had stopped looking at the screen.”
IOW the human oversight begins diligently, but after awhile
“supervisors are trusting the agent on the obvious decisions… [Eventually] they're skimming. They open each case but spend less than 15 seconds on it. The override rate, which started at maybe 8%, has dropped under 2%, and the institution reads this as the agent improving. It isn't.”
That Escalated Quickly
A former colleague very fulsomely investigates (11.000 words!) every nook and cranny of the Anthropic/DoW dispute. I can’t add anything. But it edifies my early (hesitant) take: most people early on were completely missing the point. Like this absolutely confounding piece by a respected national security expert [complimentary] (ed. though telling a policy story in such Manichean terms…tsk tsk5).
My simplified conclusion is this. It’s not a matter of: Who should set rules of war? Who is the moral actor? What should the DoW be allowed to do? and so on.
There was an agreed-upon contract. Now they disagree on the meaning of the terms.
Happens.
All.
The.
Time.
(just not usually this high-profile)
Arguing otherwise, the earlier piece imagines “a company that makes a specialized bolt for the F-35 demanding a veto over how the aircraft is used in war.”* lol. The Joint Strike Fighter program is not your friend here:
50-90% cost overruns
doubling of the delivery timeline
violated federal law (forcing Congress to amend)
were still retrofitting* IOC (testing) phase lot units while actively buying new ones6
That’s what veto power looks like. Specifically over how a thing is used in war. Not delivering the thing after getting paid is the ultimate use restriction.
*The analogy, presented as a reductio ad absurdum, is in fact what happened!
Lockheed truncated the TR-3 platform by disabling honest-to-god combat capabilities.
The contractors can holdup the buyer - having already already spent $Billions - forcing a renegotiation. Such is the business. The downgrade + retrofit was a client concession meant to speed up production.
She uses the phrase “held hostage.” Let me introduce you to Lt. Sole Prime Contractor; he certainly knows from holding hostages.7
It’s just so smug I can’t even Anthropic is not equivalent to a bolt subcontractor. There are 4-6 leading AI labs in the world. It’s one of them.8 I'm NOT defending their position or behavior. I have no idea about and don’t particularly care to weigh in on that.
But if you care at all about U.S. defense capacity you should care whether contract terms are unilaterally changeable ex post or not. Because if so, one or both things result:
(1) costs go up to reflect contract risk
(2) fewer suppliers come to your party
(and right now that’s looking like an issue).
BTW, contractors set de facto use restrictions all the time.
F-35 maker: You can’t go above 50,000 ft.
DoW: You don’t tell me what to do!
F-35 maker: Well, you can, but it’ll probably fall apart. We don’t know how to build for that. But hey, knock yourself out!
Etc.
N.B.: If you hated Disclosure Day, consider trying again and not thinking of it as an alien movie or a MacGuffin thriller. But rather as Spielberg fully and unironically conversing with his career, and treating the audience like they’re in on it.9 See if it lands better.
Regardless, Emily Blunt puts in a “Jalen Brunson in the NBA Finals” level performance.
Preview: Upcoming essay on pop culture as solution to Wittgenstein’s language games.
Recipient of the “Highest impressiveness-to-modesty ratio of anyone I know” Award.
Some examples. Not implying these are equivalent to AI: nuclear (42 U.S.C. § 2210(n)); ship seaworthiness (Mitchell v. Trawler Racer, Inc., 362 U.S. 539 (1960) - common law); some vaccines (42 U.S.C. § 247b(j)–(l)). Numerous others, varying among states. Note: if you’re keeping a wild animal, you’re on the hook wherever you are.
Start with standard Romer growth model:
Instead of R&D labor (LA) going entirely to production of the knowledge stock, some is peeled off into resolving knowledge frictions. The diverted R&D labor (LR) then is described as:
bounded by: θ(R) ∈ [0,1]. At R = 0, we have maximally inefficient drag, and as R→1 we’re at hyperregulation.
At some optimal R* we can’t reduce additional drag without regulatory compliance greater than the gain. I’m skipping some steps but think of θ(R) = 1 - (T(R) + C(R)) as the tradeoff between translational drag and regulatory (compliance) drag. Effective share of R&D labor contributing to true new knowledge production is θ(R)LA.
Through some comparative statics we get an optimum where -∂T/∂R (marginal benefit of regulation/reducing drag) = ∂C/∂R (marginal cost of regulation).
Not strictly monotonic, methinks, and certainly a different gradient for different tasks (ex: high vs. low saliency/value).
In other words, before the military could fully demo and validate systems on the initial units delivered, they were so far behind they had to push ahead to ordering more. This in spite of those first planes needing significant upgrades before they could be considered “complete.” Which they agreed to as a way to speed up delivery.
It’s not what you want.
Although the 1968 CX-5 looks pretty cool.
Putting a finer point on it: the specialized bolt maker indeed can “veto” war applications. Not in those terms, but they may later determine “under certain conditions while executing a G-Warm / Break Turn the bolt becomes structurally compromised and risks engine failure.” It's not (nor should it be) in moral terms, and not formally a veto. They're saying “this doesn’t have the ability” not “you shouldn't be allowed.”
All of this is premised on uncertainty. The Pentagon might run it's own tests concluding it's within tolerable margin of error - nothing is 100%. But a single O-ring failure can blow up a space shuttle so capabilities and robustness are the criteria for “vetoes.” Anthropic could be wrong about the ability to accurately perform some security tasks. It's irrelevant whether they put it in civil liberties framing or not, the true answer is an empirical one and neither party knows with any certainty at this point.
It’s also important that there is in fact more than one. It means no one of them has the holdup leverage of a Lockheed vis-a-vis the JSF program. Again, the software provider could very well be wrong, but critical capacity is not constrained by lack of substitutes (xAI, OpenAI, Gemini…). It didn’t help that DoW’s breathless rhetoric of “this is essential for our national security mission” turned into “this is essentially a threat to our national security mission.” QED
It's not overstatement to say many of the elements of film we consider cliche are because he created them.
That NBC anchor though, c’mon! Felt like an arrival moment for her.






