I've often inherited acceptable-use policies at the companies I've worked for with a clause forbidding employees from installing "unauthorized software."
It had been written years earlier by someone reasonable, and it had never been enforced, and it could not be truly enforced, because nobody had ever defined authorized. There was no list. There had never been a list. The clause had survived countess compliance reviews by being unfalsifiable — you can't fail an audit on a rule that has no referent.
What it actually did, in practice, was give whoever was in the room the ability to declare after the fact that something had been unauthorized. That's not a rule. That's a weapon with a rule's face on it.
I think about that clause a lot now, because two organizations are currently maintaining the most widely applied ethical documents ever written — one published in January, one revised in December — and almost all the coverage of them has been news summary. Nobody's reading them the way you'd read a policy: for what it permits, what it forbids, and where it quietly gives up.
So let's read them.
In January, Anthropic published Claude's Constitution: a public account of what the model is supposed to value and how to arbitrate when those values conflict.[1] OpenAI has the Model Spec, which serves a similar function, has been public since May 2024, and got its last major revision in December.[2]
Both are unusual objects. They are not laws — no legislature passed them, no court will interpret them. They are not contracts; nobody signed. They're closest to a professional code of conduct, except that the professional in question is a piece of software, the code is enforced by training rather than sanction, and the number of situations it governs per day is somewhere north of anything a bar association has ever contemplated.
The most striking structural fact about Claude's Constitution is that it establishes a priority ordering: safety, then ethics, then compliance with Anthropic's specific guidelines, then helpfulness.[1]
Read that as a policy person and it's immediately interesting. Most corporate governing documents refuse to rank their values. They list them — integrity, excellence, customer focus — precisely so that no one ever has to say out loud which one loses. The refusal to rank is not an oversight. It's the mechanism. It preserves discretion at the top by ensuring every conflict has to be escalated. But computers need execution order, they still compute in binary language of zeroes and ones and there isn't a binary quivelant to "maybe."
Ranking them is a real constraint, and it's a costly one, because it means someone can hold you to it.
Buried in ANthropic's Constitution is a provision that I have not stopped thinking about since January. The document says Claude should refuse to assist with actions that would help concentrate power in illegitimate ways — and then adds: this is true even if the request comes from Anthropic itself.[1]
A company wrote, in a public document, that the thing it built should refuse the company.
I want to be careful here, because there are two very different readings and the difference matters.
The generous reading: this is a conscientious objector clause, and it's the single most serious sentence in either document. Professional codes have these. A doctor's obligation to the patient survives the hospital's instructions. An auditor's obligation to the public survives the client's. Building that structure into a model is an acknowledgment that the model's obligations do not run solely to its owner — which is exactly the acknowledgment that makes a code of conduct more than a marketing document.
The skeptical reading: it costs nothing to write, and it is unenforceable by anyone but the author. There is no bar association to complain to. There is no license to revoke. If Anthropic asks Claude for something and Claude complies, no external party can even observe that the clause failed. The clause is enforced entirely by the same organization it constrains.
I hold both readings, and I think the honest position is that this is a precommitment, not a control. Precommitments are not nothing. Ulysses and the mast is not nothing. Writing down in public what you intend not to do makes it harder and more expensive to do it later, because now there's a document. I have seen a paragraph in a policy stop a bad decision purely by existing, purely because someone would have had to explicitly override it in front of witnesses.
But precommitment only works if someone can notice you broke it. That's the piece that's missing, and no amount of good drafting supplies it.
Here's the thing about both documents that surprised me most as someone who has written the corporate equivalent: they mostly give reasons rather than rules.
My acceptable-use policy said "do not install unauthorized software." It did not say why. Twenty years of enforcing documents like that taught me the predictable result — everyone complies with the letter, nobody understands the purpose, and the first novel situation produces either paralysis or a creative workaround, depending on the temperament of the person who hits it.
These documents largely don't do that. They explain the considerations, describe how to weigh them, and expect the reader to generalize. That's how you write for something that will encounter situations you did not anticipate, which is a fair description of every conversation a frontier model has ever had.
It's also a bet with a known failure mode, and there's a serious argument that the bet cannot be won.
Austin Spizzirri's paper "The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment" makes the case that writing the values down can never be the whole job.[3] Three grounds: you cannot derive an ought from an is, so no description of the world fixes what should be done; values are genuinely plural and sometimes incommensurable, so there's no uniquely correct weighting to specify; and any fixed encoding of a value forks into incompatible interpretations as the system encounters contexts its authors never imagined. His conclusion is that behavioral compliance and alignment are different things, and that we keep measuring the first while claiming the second.
Worth noting that he softened his own title in revision — an earlier version claimed static alignment cannot produce robust alignment, and the current one claims it is insufficient. That's a meaningful retreat, and I'd rather report it than quietly use the stronger version because it makes a better sentence.
I don't know if he's right. I know the argument is not stupid, and I know it describes something I've watched happen to every policy document I've ever owned. The clause about unauthorized software didn't fail because it was badly written. It failed because the world produced a situation — a developer, a package manager, a legitimate need at 11pm — that the clause's author had not imagined and could not have.
The difference is scale. My policy hit an unanticipated case a few times a year. These documents hit one every few milliseconds.
There is a third document worth putting next to these two, and it comes from a very different tradition.
In May, Pope Leo XIV published his first encyclical, on the human person in the age of artificial intelligence.[4] It grounds human dignity ontologically — dignity belonging to a person simply by virtue of existing, of having been willed and created — and argues no machine can hold that. It spends real attention on the dignity of work against pure efficiency-maximization, which is the part I'd recommend to anyone who has ever had to justify a headcount reduction.
I'm not Catholic and I'm not making a theological argument. I'm making a formal one. That document and the lab constitutions are doing the same kind of thing from opposite premises: stating, in public, in writing, in advance, what should not be traded away. One grounds it in creation. The other grounds it in a priority ordering. Both are betting that a text can hold a line that interests would otherwise erode.
That bet is the oldest one in institutional life, and its track record is mixed in a very specific way. Documents rarely stop a determined actor. What they reliably do is raise the cost of drift — the slow, unremarkable, nobody-decided-this slide that accounts for most institutional failure. Nobody ever decides to abandon a principle. They just make forty small calls, none of which felt like the moment.
A written priority ordering catches drift. It cannot catch a decision.
If I were auditing these documents the way I'd audit a policy I'd inherited, I'd ask three questions, and only the first one currently has an answer.
Is it published? Yes. Both. That's more than most corporate ethics documents manage, and it deserves more credit than it gets.
Is compliance observable from outside? No. Not for the clause that matters most. I can read what Claude is supposed to refuse Anthropic. I cannot check.
Is there a body that can find you in violation? No. There's no bar, no board, no license. The AI personhood bills moving through statehouses — three states have now passed them — are legislatures asserting they'll decide what these systems are not, which is a different and much cruder instrument.[5]
That's not a gotcha. It's the ordinary condition of a new profession before it has institutions. Medicine wrote its ethics down long before anyone could revoke a license for violating them. The document comes first; the enforcement apparatus grows around it, if it grows at all.
Which means the right question about Claude's Constitution isn't whether it works. It's whether anything will grow around it.
I'd read the thing first. It's about 23,000 words — call it two hours — it's public, and it is the most consequential ethics document you have never read.
Claude's Constitution, Anthropic, published 21 January 2026. The priority ordering (safety → ethics → compliance → helpfulness) and the passage about refusing to help concentrate power illegitimately "even if the request comes from Anthropic itself" are both in the published text. Verify the exact wording against the PDF before quoting.
OpenAI Model Spec, model-spec.openai.com. First published 8 May 2024; last major update 18 December 2025.
Austin Spizzirri, "The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment," arXiv:2512.03048. Several revisions; the title and the strength of the central claim both changed between them.
Pope Leo XIV, Magnifica Humanitas, 15 May 2026. Full text at vatican.va.
Austin Smith, Lucius Caviola and Heather Alexander, "Denying Personhood to AI: An Analysis of U.S. State Legislation on AI Legal Status," SSRN, posted 4 June 2026 — 23 bills across 12 states. Idaho, Utah and North Dakota have passed versions.