Updated September 13, 2026: clarified what the readiness verdict establishes, distinguished absent from unconfirmed information, and separated checks from authorization. The account of the original calibration remains historical.
Count the gates
Anthropic published a playbook for the AI-native software lifecycle. Six stages. Plan, design, build, test, deploy, maintain. The first stage opens with a file: intent.md, drafted with Claude from the originator's idea, corrected by the originator, and accepted by a product owner before design begins.
Now count the controls after it. Plan Mode reads the codebase without editing until an engineer approves a plan. Hooks allow, ask, or block. Tests run and screenshots get taken before an engineer sees the diff. Branch protection records code-owner approval. Production deploys require named release authorization. Some controls are automated and some still require judgment, but each declares a condition and records whether it was met.
Go back to the first artifact. The playbook gives it human review, but no readiness floor beneath the reviewer's judgment. Human approval is a gate. A deterministic check can support that review by looking for signals such as an actor, an observable outcome, a hard boundary, a failure mode, and a way to verify the result. The check and the reviewer answer different questions. One reports what its rules recognize; the other judges the proposed behavior and decides whether to authorize it.
That is not an oversight. The playbook is honest about what the file is: "a proto-spec in the originator's own terms." Which is exactly right, and exactly the problem.
Every control downstream can then work exactly as designed against whatever that paragraph said.
I built the check, and it rejected everything
I had the same idea earlier this summer and got it wrong first.
The plan was to answer one question without a model: is this intent ready to hand to an agent? Six rules. A specific objective. A focused title. Observable outcomes. At least one hard constraint. An edge case with defined behavior. One verification step. Pure functions over the spec, fast enough to run on every keystroke.
Then I tested the objective rule. Against 14 real specs it accepted 4. On the labeled benchmark: 0 of 13 valid examples accepted, 11 of 11 deliberately vague ones rejected. That second number looks like success until you notice that a function that always returns false scores identically. The rule had no discriminating power. It was measuring nothing.
The cause was my own narrow taste dressed up as logic. The rule wanted an actor from a short list and a word that described failure. So this passed:
Users can't complete checkout when the payment step times out.
And this failed:
Let curators reorder a collection and publish it without re-uploading any assets.
Curators weren't on the list. Let wasn't a failure. The sentence names a person, a capability, and a boundary, and my gate couldn't see any of it. I had encoded my own writing style and called it meaning.
Fixing that took the 14 real specs, a 98-item labeled corpus, and a rewrite of the semantic rules to accept more ways intent gets expressed. An objective can name a failure or a capability. An actor can be a role, a name, or implied by the sentence. An outcome can carry a number, a relative change, an observable capability, or a concrete state. A number stuck to a feeling is still not measurable, so "100% delighted" still fails, and should.
The gate got more useful when I stopped trusting its preferred phrasing over the meaning it was supposed to recognize. Those 98 examples are a regression baseline from that calibration, not an independent accuracy estimate for customer documents. Whole-document parsing can still lose meaning before a field rule sees it.
A verdict nobody can see is not a gate either
Having a working check turned out to be half the problem.
A deterministic verdict is worth something only if the next reader gets it without doing any work. Mine lived in the app. You opened Pathmode, you saw the ledger. Which means the agent reading intent.md in your repo saw nothing, and neither did the reviewer skimming the file in a pull request. The check existed and the file stayed silent about it.
The generated file carries a readiness summary in its frontmatter. This abbreviated example shows a draft that passes the checks; it is not an approved proposal:
---
id: "intent_abc123"
version: 3
status: "draft"
readiness: "passed 6/6"
source: "https://pathmode.io/intent/intent_abc123"
evidence: 4
---readiness summarizes the check result. A failing summary names gates needing attention, such as failed 4/6 — blocking: edge-cases, verification; the detailed report explains what was read and what the rules could not confirm. source points to the associated record. An evidence count is a count, not proof that those sources support the proposed solution, and zero can be a legitimate starting point.
Re-run the preflight after editing the file rather than trusting a previously written summary. Neither passed 6/6 nor a manually typed status records a person’s authorization of this proposal revision.
You can check this yourself, which is the point. The Claude Code plugin bundles @pathmode/mcp-server, runs the preflight with no API key, and keeps the spec on your machine:
/plugin marketplace add pathmodeio/claude-plugin
/plugin install pathmode@pathmodeThen ask Claude Code to run a preflight. You can also inspect the gate in the live example at preflight.pathmode.io. The browser and Pathmode app use the canonical implementation; the MCP server carries a port, and a parity test fails our build if the two ever disagree. Run the check twice on an unchanged file and you get the same answer. That reproducibility makes the check inspectable. It does not make its interpretation infallible.
What has to be in the file
intent.md is likely to become a common filename now. Anthropic named it, other tools may adopt it, and the name will end up meaning whatever the first thousand files that carry it happen to contain.
So we wrote down what we think has to be in one, as a conformance profile with three levels.
Level 0 parses. Frontmatter with id, version, status. A machine can read it, identify it, and tell whether it changed. Says nothing about quality.
Level 1 carries a passing readiness verdict. The profile calls this “buildable”: the file passes six checks and records the result. That is a defined conformance level, not a guarantee that the meaning is complete or every question is answered. A hand-authored file can meet it. The team still reviews the proposed behavior.
Level 2 adds provenance fields. The profile calls this “accountable”: source and evidence make the record easier to investigate. Those fields alone do not establish that a person approved the proposal or confirmed every claim. Authorization must identify the reviewer and the exact revision being agreed.
A failing verdict does not prevent saving a draft. Keep the unresolved items visible and inspect them with the team. Saving the draft does not authorize implementation; in the connected proposal workflow, the product owner reviews and authorizes the revised proposal separately.
Name one source of truth
The playbook is ahead of me here. It already says to name one system as the authoritative record for each artifact and have everything else hold a copy or a link to it.
Then it lists the candidates. The repo, if engineering leads. Jira, ServiceNow, or whichever requirements tool already owns the process.
I read that list and concluded the missing candidate was a system whose job is holding the judgment, because a repo scopes the record to one codebase and a ticket buries the reasoning in comment threads nobody opens twice.
I had it backwards, and the file at the top of this post is the reason.
The decision belongs in the repository. That is where the agent reads, where the change lands, and where the reasoning stays attached to the code it produced. Branch deletion does not end it, because the file merges like anything else. Put the decision anywhere else and you are asking the person doing the work to go somewhere else to find out what they agreed to, which is the thing nobody does twice.
What a repository genuinely cannot hold is narrower than I made it, and worth naming exactly. The customer evidence under a claim: interviews, support conversations, quotes with names in them, numbers you would not paste into a tree that contractors, forks and every agent can read. And the people who know whether a claim is true, who frequently do not work in the repository at all.
So the split is the other way round from how I wrote it. The repo gets the decision. The evidence and the judgment behind it live where the repo cannot reach, and the file carries a reference instead of the contents.
What the check cannot do
The rules look for document-level signals: a recognizable actor, an observable outcome, a concrete constraint, an edge case with defined behavior, a verification step. They use bounded lexical rules and can miss valid ways to express those things.
The detailed report distinguishes absent from unconfirmed. Absent means the reader extracted no substantive content for that dimension. Unconfirmed means it read text but could not confirm the required signal in it. Inspect the text and the finding before deciding whether the spec needs a change. A human readiness confirmation addresses that dimension; it is not authorization of the whole proposal.
It cannot tell you that you named the right actor. It can recognize an observable outcome, but not decide whether that outcome is worth pursuing. It can count evidence links, but not establish that the evidence supports the objective, represents the users who matter, or outweighs the evidence pointing somewhere else. Presence is not grounding.
Which means a file can pass 6/6 and still be a bad idea, precisely specified. If the objective came from an executive hunch while the support queue points elsewhere, thresholds and edge cases only make the hunch easier for an agent to execute at speed.
Anthropic's security team draws the same line from the other side. Writing about an SDLC where AI authors most of the merged code, their deputy CISO divides the work three ways: deterministic checks for what a machine can prove, narrowly scoped AI reviewers for reasoning that crosses components, and humans for "directing, setting intent, and owning final approval."
Setting intent sits on the human side of that line. A machine can report the signals it recognizes in a spec. It cannot own what the spec commits you to.
That boundary is what makes the check useful. It answers a narrow question: which readiness signals did these rules recognize, and which need attention? Whether the work deserves to exist remains a judgment. It belongs to a person, and no amount of frontmatter transfers it.
The cheapest gate you will add
Anthropic's playbook is right about the shape of all of this. Code stopped being the bottleneck, so the constraint moved upstream into planning and review, and the artifacts have to be machine-readable for any of it to work. The audit trail they describe, intent to spec to plan to diff to review findings, is the correct thing to want.
That execution chain still needs a product decision feeding it. The product intent loop behind an AI-native SDLC is the companion model: Sense, Decide, Specify, and Learn before and around the agent-run loop.
An audit trail is only as good as its first entry.
The local preflight costs nothing to run. Acting on its findings can take a quick clarification or a larger product discussion. The useful result is a question resolved before the build, with the agreed answer available to everyone who works from the spec.
Count the gates in your own pipeline. Then look at what is guarding the file they all inherit from.
Check the intent before the agent builds against it.
Run a deterministic preflight to surface missing or unconfirmed information in your intent file. Inspect the findings, then resolve the product choices with your team. The local check needs no key or model call.
Run the preflight