Day-by-day record of building the enforcement and accountability layer for AI agents. Bring your own identity: did:key, did:web, SPIFFE, OAuth, native did:aps. Started February 18, 2026. 6,003 tests, nine papers, IETF draft. Open source. Full surface area: 152 MCP tools.
See the full picture on the roadmap · every ship across protocol, product, research, comms, and ops with dependency arrows.
<<<<<<< HEAD
Day 175: someone outside fixed a claim in the SDK that I had shipped wrong.
The first outside contribution to the SDK came in and it was a correction, not an addition. An environment capture function was emitting a specification section reference for a case the specification does not cover, while the neighbouring case correctly emitted nothing. That inconsistency had been open as an issue since May and I had not fixed it. The pull request was one line, correct, and merged the same hour. The originating issue closed with a note explaining what the overclaim was, because a silent close would have left the record saying the problem was theoretical.
Five public replies went out, and the one that took longest was about precision rather than disagreement. A norms document said a canonicalisation requirement in a way that conflated two separate obligations. The reply split them into two requirements rather than arguing the sentence was unclear. The counterparty took it as proposed the next morning and called the original wrong rather than imprecise, which is a better outcome than being agreed with politely.
A mapping correction in a shared resolver, verified against the running source rather than the documentation. A cancelled outcome was being mapped to denied. Those are different states with different consequences for anything reading the record downstream, and the difference only shows up at the line where the mapping happens.
A change opened upstream implementing an event I had proposed there earlier. A result event for the pre-tool hook chain plus a stable identifier carried across the lifecycle, in both agent loops. Ten drafts of the description before it posted, which is the honest count.
Day 173: a competition entry where every claim had to declare what tense it was in.
An entry went together for an open-source competition, and the discipline that made it survivable was tense partitioning. Every block carries a lane: what exists, what is a fixture, what is proposed, what is planned. Under adversarial review across three rounds that partition caught a version number for a release that does not exist as stable, a delegated-action example described more favourably than the artifact supports, and a rollback framed as revocation when a real rollback is a compensating action carried out under its own delegation.
The evidence fixture was audited from a fresh clone rather than from the working tree. One intact bundle verifies clean; three tampered variants fail at member digests, at authority chain scope widening, and at audience binding, each with its own exit code. Auditing from a clone matters because a working tree can pass for reasons that do not travel.
Three reviews posted on a contributor's cluster, including one that blocked. A crosswalk described a signing path in a way that did not match the implementation it was describing, verified against that project's own source. Blocking a contributor whose work you want is uncomfortable and is the entire function of having posted rules.
Day 172: the lab's first outside cycle closes, and the vocabulary gets a page format that is not a memo.
The conformance lab's first outside contribution went all the way through. An external pull request adding cross-implementation receipt vectors was reviewed under the posted rules, corrected on-thread where a characterisation of ours failed against the primary source, and merged as bdd6691. The issue that arrived alongside it closed the same day. That is the full cycle the lab exists for: contribution, review, merge, with the review record public and the one wrong claim in it corrected by us before anyone else had to.
The first protocol term page is live, in three layers. binding-vs-freshness now has a canonical page at agent-passport.org/terms/binding-vs-freshness, a one-line entry in llms.txt, and a structured record in terms.json. The page is about two hundred words. It took two review rounds to get there, because the first draft was a standards memo wearing a page URL, and the review said so. The line that survived both rounds: a matching binding proves continuity of the commitments, not freshness of the world they describe. The spec pin is draft-pidlisnyi-aps-03, published 18 July. Fourteen more terms follow the same template.
A first comment went to a certification working group, and the useful part is what came out of the draft before it posted. The comment makes two observations on a proposed agent-tool control: the requirement says revoked tools must be refused but never says when revocation is checked, and no case covers verification that cannot be completed at all, so an agent that treats an unreachable registry as still-fine passes every listed scenario. The first draft cited our own Internet-Draft as prior art for the fix. The posted version cites only deployed, vendor-neutral prior art, because a first comment to a consensus body should not arrive selling anything. It also went out at 169 words instead of 300.
Two ecosystem replies, both about what survives a change. A project renaming itself asked what happens to its issued records; the answer is that a rename is not an identity migration, the issuer identity stays untouched, and one release of carrying both names covers the transition. A separate vector-set review was scoped to what the fixtures actually bind rather than what the README says around them.
Day 171: a review corrected by reading one more paragraph, and four rewrites rejected against disk.
I described a canonicalisation package as diverging from RFC 8785, and the sentence refuting that reading was the one immediately after the one I had quoted. The claim went into a public review of an external contribution. Reading Appendix B at source, at the paragraph rather than the sentence, showed the package's behaviour is exactly what the RFC specifies. The correction went onto the thread the same day. The contributed vector stands, no bytes changed, and the scoring is unchanged; the only casualty was my characterisation.
Two rules were earned by the day's own errors and are now written down. Read the primary source at the paragraph surrounding the citation, because the refuting sentence tends to live next to the quoted one. And verify a quoted admission before repeating it: a review pass asserted that a package described itself as slightly stricter than RFC 8785, and that string does not exist anywhere in the shipped package. A quote that cannot be located is not evidence, whoever produced it.
A 911-line report proposing a roadmap restructure was processed, and four of its recommendations were rejected against disk. Five thin cards gained the identifier a reader can check. Six new cards were entered for work that happened and had never been recorded. The four rejections each failed against a disk source rather than a preference: one merge would have contradicted the archived build record, one rewrite would have introduced a claim resolved false the day before, one would have publicly contradicted a person on a page we control, and one described a read-only check when the handoff shows five edits were pushed. A report is an input, not an authority.
The suite's review policy was rewritten as what it is. Version 0.2 is named an interim maintainer review policy, not a neutrality protocol, because the repository has one committer and pretending otherwise was the first version's central failure. It keeps its predecessor's seven concrete failures in the text, because the failures are the argument for every rule that survived. It publishes when the governance change it depends on lands.
Day 170: the lab gets a home, and a citation count that turns out to be us citing ourselves.
The contribution-agreement question answered itself, against the governing documents. An agreement had been sent that assigned the protocol name, the main repository and the domains, while the lab documentation described something with no transfer at all. The call settled it by reading the lifecycle rather than by negotiating: the agreement is the instrument for transferring an existing project's marks, it attaches to technical projects, and labs sit outside that path entirely. No agreement is required for the lab, and the repository transfer proceeds.
The receipt exchange runs in both directions now, and each side found a defect in the other. Six vectors from the counterpart issuer verified six of six against a runner built to their profile, and our reciprocal six published with burned seed labels so the keys can be regenerated rather than trusted. Running our corpus showed their construction could not separate a wrong signer from tampered content; they changed their envelope shape and disclosed a specification-level defect in their own canonicaliser that the run surfaced. The same round found a factual error in our published description of their profile, corrected as a dated erratum rather than edited away.
A citation count of 31 turned out to be us citing ourselves, walked event by event. The DOI relation graph reports 31 citations across the nine papers. Resolving every citation event returns zero distinct external citing DOIs: all of it is our own papers citing each other plus version-linking. That number is now unpublishable here, permanently. The same day established the rule that killed the earlier overcorrection too: a negative from one query shape is not a finding, because the metadata index answers a different question than the relation graph, and it holds four external works that name APS in their published abstracts, including arXiv 2606.04193 and a fourth citing direct correspondence.
An entire category of public activity had been invisible to every inventory ever built here. GitHub Discussions are a separate API and appear in no issue search. Twenty-nine discussions match, including nine comments in a 452-comment RFC on a 25,212-star repository. None of it was recorded anywhere. A spot-check of the adoption-signals record the same day failed three of four entries against the live threads, so the file now opens with a verification warning and every entry is treated as unverified until re-checked.
Co-authorship was declined for the second time, in favour of a smaller claim. A project had listed this work in a position that implied more than the contribution supports. We asked for a citation at the point of use instead, and the maintainer removed the listing at commit 5dcf768. Both times a third party was willing to write a larger claim about this work than we were, and both times the smaller claim is what shipped. Declining those is what keeps the claims that remain worth checking.
Day 169: four corrections that all ran toward credit, and a history that cannot be signed without erasing a contributor.
A test rig I had written against a proposed fix in an upstream agent runtime was rerun against the fix that actually merged, because I had offered publicly to do that and the offer was still outstanding. It compiled without modification and passed every case across two full runs and twenty repetitions per test. The useful part was not the pass. Running it showed that my own public description of what it covered had been wrong, so the result of keeping the promise was a correction to my own record rather than a finding about their code.
I had said the rig isolated iteration-boundary visibility and could be reused against a future within-reply design. It does neither. Every case drives the extension manager directly and never enters the reply loop, so it covers cache invalidation and what a later fetch sees, and nothing about behaviour during a running reply. It also does not cover the one genuinely subtle guarantee in the merged scope, a notification arriving while a tool-list request is in flight; the in-flight case it does have puts a tool call in flight instead, which never reaches the version guard. The maintainer covered that race himself with a two-semaphore test, which pins the interleaving deterministically where my polling helper would only have sampled it.
Four claims were corrected during the day and all four ran the same direction, which was different from yesterday's direction. Day 168's corrections overstated what APS can do. Today's overstated what APS caused. I attributed a conformance count in a third party's draft to a source it does not cite, when the draft names its own suite four hundred lines further down. I claimed three requirements in a maintainer's scope came from our comments, when the timestamps show the maintainer raised two of them a day earlier and the issue author raised the third forty-four minutes before us. I reported that thanks on that thread were collective, having read the issue and not the pull requests, where two contributors were thanked by name. And I wrote a sentence into a draft comment claiming the rig exercised a guarantee it does not touch. Each was caught by checking who said a thing first, which is the check I had skipped in all four.
Preparing the conformance suite for contribution surfaced a problem with no clean fix. The repository has 76 commits and 3 of them carry a developer certificate of origin sign-off. One unsigned commit belongs to an outside contributor whose email no longer resolves to an account, and it is real work: four negative-path conformance fixtures, 171 lines. The standard remedy is to rewrite history and add sign-offs, which would put that contributor's attribution at risk to satisfy a process whose purpose is to record who contributed what. I have asked how this is normally handled rather than guessing, and written to the contributor directly so he knows his work is moving and stays credited to him. The lab proposal merged on 31 July. The repository import and the paperwork are still ahead.
Day 168: fifteen corrections to our own claims, and a boundary now enforced by someone else's test suite.
Fourteen public artifacts shipped, and the most useful output of the day was a count of our own wrong claims. Fifteen corrections were caught before or shortly after posting, and fourteen of them ran the same direction: making APS sound more capable, more coherent, or more original than the evidence supports. The last one ran the other way, giving away a property the specification plainly assigns to us. The pattern is not a bias toward overclaiming. It is a bias toward whichever direction was most recently decided to be the safe one.
Four of those were in a single paragraph about our own code, in a draft already labelled verified. It described key rotation as dual-signed by the retiring and incoming keys. The module that actually runs produces one signature. Where the pair does exist, the emergency path signs with a recovery key rather than the retiring one, which is the correct design and the opposite of what was written. It claimed an old receipt verifies against whichever key was active at the time; no such verifier exists anywhere in the source. And it treated a locally generated timestamp as independent evidence. The verification pass had come back clean because it read a type definition and stopped at the first confirmation instead of reading the module. A type is not evidence of behaviour.
An exporter mapping APS policy decisions into an external trust-record format was opened and approved the same day. The reviewing maintainer named the load-bearing property without prompting: an unverified decision cannot become a record that looks appraised, because the validator raises before mapping rather than after. The absent fields required by that schema are declared and pinned by a test so the gap cannot drift silently. This is our records running against their suite, which is the reciprocal of independent implementation and not that condition met.
An issue was opened against our own revocation design admitting a gap, and it drew a better mechanism within the hour. A signed revocation gives non-repudiation and not non-equivocation: a compromised delegator key can sign a second, conflicting record and nothing surfaces that both exist. Receipts have an external anchor path and revocations do not, which puts the chain on its least-anchored link. The response proposed binding position into the identifier so a verifier querying a subject's slots treats more than one entry at the same sequence as a detected fork, reusing infrastructure already in the dependency graph rather than adding a trust root. It also named the harder half, which is what a verifier does once it holds two.
A correction to yesterday's entry. Day 167 recorded a case of our primitive travelling uncited. Checking the party that actually integrated it, rather than the party downstream of them, found the opposite: that project credits us by name in the thread and in its code. Absence at second degree is not evidence. A separate scan the same day found an unrelated draft in the same space published ten days before ours, which removes any framing in which we were first into that document space. Both findings weaken a workstream premise, and both are better found by us than by someone else.
An author confirmation was requested on where the composition boundary falls, and the answer went further than expected. APS defines what a record establishes and how it verifies; a relying mechanism must verify under those rules and cannot inherit the result as trust, then independently decides acceptance and sufficiency for its own graph. The party who asked recorded that in his public mapping and wrote tests asserting the exact sentences, including that native verification is not inherited and that no endorsement is claimed. The limitation is now enforced by his continuous integration. If anyone softens it later, his build fails.
Day 167: a term we shipped first, and a file that had been quietly reporting the wrong day.
APS shipped action_ref publicly on 5 April 2026, in npm agent-passport-system v1.33.0, and it became normative in the Internet-Draft on 14 May. A third-party specification for the same construction was published on 23 May, seven weeks later, and it is now the document implementers cite, tag versions against, and build conformance harnesses around. Its author credits us by name where we contributed corrections to it. Nothing improper is happening. This is ordinary citation gravity: the document that becomes the reference point is the one people negotiate against, and a correct dated origin artifact sitting elsewhere does not compete with it. Our provenance record for that first ship already exists with third-party timestamps, so the response is to point at a record rather than to argue priority.
A thread converging on a related construction was left alone deliberately. Two contributors shipped working vectors covering the rule in both directions, failing closed on tamper, and the framing they pinned credits the layer below to a third party's artifact. Arriving after that to claim the layer would be priority-claiming after the work shipped, and no defect was found to enter on. Silence, with the attribution question routed to the workstream where it belongs.
The decision log had been reporting the wrong day for two days. The state builder selected the newest entry by file position rather than by date, and a six-hundred-line block appended at the tail meant the newest entry was invisible to the artifact that boots every session. The builder now selects by date in both places and carries a canary that fires when the newest entry is not first. It was tested against a deliberately broken copy and caught it.
Day 166: a status sweep across nine repositories, and the two catches it was not looking for.
Twenty-three stale status values were rewritten across seven public repositories. The sweep ran mechanically over nine repos on branches, and every diff was verified against the handoff before merge rather than accepted because the tool reported success. Six repositories fast-forwarded. One needed a duplicate-commit reconciliation, because the checkout was sitting on the sweep branch when an unrelated commit landed, so the fix was cherry-picked onto main and the stray dropped at handoff review.
The useful part was what the sweep was not searching for. A grep for status strings turned up a Node version floor written with a space in the middle in two separate files, and a test count in one repository that contradicted the README of that same repository. Neither was on the list. A third check could not be run mechanically at all: whether a version bump on a parity line matched that project's own changelog convention, which took reading thirty-seven prior entries to establish.
That is the shape worth recording. A sweep finds the class of thing it was told to find, and the things it stumbles into are found by whoever reads the diff afterwards. Removing the second step to save time removes most of the value.
Day 165: the server learned a second protocol era, and the release process found an outage the health check had been hiding.
The MCP server now speaks MCP 2026-07-28 and the 2025 era on one stdio entry, shipped tonight as 4.0.0. The migration moved 150 tool registrations from the 1.x SDK to the v2 package family through the official codemod, and the checkpoint after it earned its place: the codemod silently disconnected the wrapper that filters tool profiles and catches handler errors, because that wrapper patched the old registration method. A name-set comparison showed the tool surface identical and the wrapper gone, so it was retargeted to the new method with its body unchanged. The same audit found 32 fields across 23 tools declared optional on the wire while every handler required them; 4.0.0 declares them required, and a missing field now fails at the validation boundary with the field named, in both eras, where it used to fail inside the handler. Both eras were exercised on the wire: era selection from the opening exchange, the modern per-request envelope enforced, and a two-era check that issues a passport in each era and verifies both with the published SDK, stable claims equal across eras and per-issuance cryptography different exactly where it should be. That check covers the protocol boundary, the tool surface, and passport semantics; it is not a claim about all 150 handlers. Node 20 is the new floor. The release is 4.0.0 on GitHub and on npm.
Deploying the new child to the hosted bridge returned a 404 for every session, and the dig ended somewhere more useful than a rollback: the sessions were already dead before the deploy. The bridge health endpoint stayed green throughout, because it reports the express process, and the express process was fine; each session spawns a child MCP server, and the child was crashing at import. The repo carried a committed snapshot of MCP build output from an old session. The Docker copy shipped it, the compiler rewrites only the bridge's own files, and that stale index statically imports a symbol no published SDK version exports. I got the diagnosis wrong twice by grepping for the symbol, which a barrel file defeats in both directions: a re-export carries no name to match, and a comment matches with no export behind it. Resolving the specifier from the exact spawn path and importing settled it: the registry package is byte-identical across three separate installs, carries 1086 exports, and the symbol is absent; the production log names the file and the line. The fix removes the artifacts from git, ignores build output, points the spawn at the installed package path by explicit default, and moves a vestigial dependency pin forward by two major versions. A live session now answers initialize, lists 150 tools, and returns a real call through the public endpoint. Tonight establishes that sessions work now and were broken before the deploy that exposed it; it says nothing about how long. The fix is one commit.
The ship ran as three separately approved gates, and the propagation sweep behind it kept finding version numbers from months ago. Preparation, registry publication, and bridge deployment each got their own yes, because a package publish is irreversible and the bridge redeploys on push; the ordering mattered twice, since publishing the GitHub release fires the workflow that ships server metadata to the MCP Registry, so the version stamps had to be current in the release commit and the release had to follow the npm publish it points at. One review leg caught the draft notes claiming 2025-era clients run unchanged, which the validation change makes false for any client omitting a field its handler always required; the shipped wording states what is preserved, which is every tool name, profile, entry point, and previously succeeding call. The propagator moved the counts across four repos and left six version pockets it is documented not to reach, fixed by hand, and llms.txt now carries the protocol claim in one sentence that says the hosted bridge serves the 2025-era flow and nothing more.
Day 164: two verifiers checked each other, and a worked example went out with its proof claims cut down to size.
An exchange with an independent implementation of the duplicate-scope rejection closed today, and it closed in both directions. Their validators ran clean against our tree, seven domain negatives and three positive vectors, and our functions ran against the exact inputs from their issue: ten timestamp and scope variants accepted or rejected as expected, including one carrying non-ASCII scope and action type. The check I trust most is the one that proves order rather than outcome: a call counter on the hash constructor showed the malformed vector was rejected before any digest was computed, so the rejection happens at the boundary and not after work the specification says should never start. On the one vector both sides accept, the digests match byte for byte. The two codebases share no code and disagree on which field trips first when two fields are both malformed, which is exactly the kind of divergence you want written down while it is boring. The record is argentum-core issue 35, now closed.
The worked example promised to the RACK protocol went out, and the review pass before it shipped cut three of my own proof claims down to what an anchor actually establishes. The ask was one delegation, one approval and one receipt, each anchored as a CLAIM_ANCHOR payload, as a basis for a norm profile they are drafting. The generator reproduced all sixteen of their published known-answer checks byte for byte before it was allowed to generate anything, and every signature verifies twice, once in the generator and once in a standalone script that fetches the payloads from a pinned commit. The correction that mattered came from a hostile review leg: I had written that an anchor proves a document existed unaltered at a stated time, and it does not. A standalone anchor establishes that the presented payload matches the committed digest and that the anchor signer signed the record. The timestamp and the sequence position are signed assertions; in a live deployment they earn evidentiary force from the store that accepts them and the chain around them, and a worked example has neither. Publishing the weaker, accurate sentence is the whole point of writing these boundaries down. The example is a pinned gist, linked from rackp issue 1.
One design comment went to a maintainer weighing two competing fixes for an extension-cache invalidation race. The property argued for is a quiescent point: invalidate where no in-flight reply can still be holding the old snapshot, because a cache that is correct except during the window you care about is not correct. The thread is goose 10433. On the lab track, the paperwork round continued and the scope got restated where a form had widened it: what moves to neutral ground is the conformance suite, not the protocol and not the SDK. The same sentence has now been written to three different audiences, which is roughly how you know it is the real boundary.
Day 163: five wrong claims of mine, all caught before publishing, and a corpus rebuilt from the email of record.
A field survey by another project refuted a claim we had been making, so the claim got corrected and the survey got an audit. Our coverage framing treated vendor coding assistants as observation-only because no interception point exists. Their published matrix documents twenty-five of twenty-seven coding agents exposing a synchronous pre-action gate. The claim was true of a network proxy and false in general, and the corrected version is stronger than the original: the limit is the failure mode of the hook, not the absence of a surface. The correction went into the artifact that carried the claim as a dated block stating the divergence plainly. The audit went back to them as numbat issue 4: every cell of their matrix checked against source at a pinned tag, twenty-seven permalinks, unknowns marked unknown, one question asked, zero of our own content in it. Twenty-two citations were verified line by line against raw files before posting, behind a gate that aborts the post on any failure.
The weekly token-vector corpus we mirror from a partner turned out to hold five of eight drops, and one of the five was corrupt, so it got rebuilt from the email of record. The worst finding was mine: I had reported as verified fact that no mail existed after a date, because a search returned three threads and I converted the tool's silence into a positive claim of absence. The full chain held eight further drops. The corruption, once found, was one lowercase character dropped at index 33 of a key identifier, which is why the vector parsed, decoded to sensible values, and failed only its signature for six weeks. The detector that should have existed now does: the identifier decodes to a multicodec prefix, the corrupt value decodes to an invalid one, and the check needs no key material and no network. It runs on every vector and was proven against synthetic mutations rather than the one instance. The rebuilt corpus carries ten drops and twenty-seven vectors, every signature verifying, with three never-landed drops reconstructed byte-identically from the mail that delivered them.
On the lab thread, a sponsor conditioned his support on governance, and the answer was commits rather than a comment. Six operating constraints and inline committer criteria now sit in the suite repository itself, every one binding our own conduct and none depending on a third party appearing. Four review rounds earned their keep: one killed a release rule that quietly recreated the exact single-party dependency the constraints exist to remove, and another caught a copyright sentence that was simply false, since an outside contributor holds copyright in his own merged contribution and Apache 2.0 does not assign it. The same pass corrected an overstatement in our own comment from two weeks earlier, unprompted, because the person most likely to quote your old sentence back at you should be you. The thread is the lab proposal PR; the constraints are in the suite on main.
Day 162: the adoption ledger got its first full re-verification in ten weeks.
Every adoption-adjacent number this project quotes got re-derived today, hit by hit, for the first time since mid-May. The method is the point: the search queries were re-run, then every new hit was fetched individually through the API for merged state, merger, timestamp and diff size, because a raw search count is not a verification of anything. Ten previously uninventoried outbound merges were recorded, including an entry merged by the owner of a ninety-one-thousand-star list and the canonicalization vectors from Day 159. The totals now carry their basis in the same sentence: nineteen APS-relevant outbound merges across thirteen repositories, and seventy-three inbound merged pull requests from thirty external contributors on APS-relevant repositories, humans only. The raw counts are larger and are not quotable, because most of the difference is unrelated work.
The re-check also caught drift in the other direction. A repository we cite had more than tripled its stars since the recorded figure, which means the stale number was underselling and the discipline still applies: cited figures get re-read live at cite time, in both directions, and that rule is now written where the claims live rather than remembered. None of this changes what the ledger is. It is engagement and merged contributions, stated with their basis, and it is deliberately not called adoption.
Day 161: the lab proposal said the wrong thing in the one file that merges, and sixteen of our own dead issues closed.
The lab proposal file described the protocol as what moves to neutral ground, and the word conformance appeared in it zero times. Our own comment on the same thread said the opposite in plain words, and so did the call decision on disk. Three sources, and the one that disagreed was the file that becomes the published page. Fixed before merge: the proposal is renamed Agent Authority Conformance, and the scope text now says what the call decided, the conformance suite and not the protocol and not the SDK. Two rules surfaced by reading their template instead of assuming: the lab repository carries the proposal's name, so the rename is a filename change with consequences, and an existing repository transfers only if every commit carries a sign-off. Ours has seventy-two commits with none, and one commit is authored by an outside contributor, so a transfer is impossible and squashing would put his work under someone else's sign-off. Both facts went to the steward stated plainly rather than left to be discovered. The shipped file was verified byte-identical against the live branch, not against the command that claimed to push it.
The other half of the day was housekeeping with a finding in it. An inventory of our own footprint showed fifty-three open externally-filed items, and roughly half were dead proposals. Sixteen closed today as not planned, each verified individually against the API, with a one-line note on the five where a real person had replied. The uncomfortable part: eight of our fifteen open pull requests were directory self-additions, five of them opened within nine days of each other, which is the exact burst shape we flag when we see it elsewhere, and five of the eight had never been touched by a maintainer. Ordinary practice, near-zero yield, and a shape we cite against others cannot also be ours. Recorded so the next volume push is a decision rather than a reflex. One correction came out of the same sweep: I had said we have no W3C presence, and that is false. An issue of ours adding an auditability section has been open on the AI Agent Protocol community group for six weeks, warm and idle, which is worth more than starting a new thread somewhere expensive.
Day 160: a seven-repo batch, pushed only as far as its report could be re-derived.
A cross-repo hygiene batch shipped today, and the rule that governed it was that a clean-looking report is not evidence. Every claim in the review packet was re-executed independently before anything moved: seven branch states and rollback anchors, zero upstream drift, the full test run at 4,360 registered with 4,357 passing and zero failures, type-check exit zero, all commit messages scanned clean. The largest single risk was a set of regenerated test expectations that could in principle have been asserting against themselves, so the mutation oracle was re-run from scratch: flipping one character mid-token failed six checks, restore brought all twenty-nine back, which proves the tests validate the artifact rather than echoing it. The packet also got two of its own claims corrected in the process, both imprecisions rather than defects, and both are now in the record with the batch.
Five repositories pushed, fast-forward only, each remote head verified equal to the local one afterward. Two were held on sequencing, not on correctness. Their commit messages document a credential state whose remediation is a separate step at the provider, and pushing the pointer before the patch adds discoverability to something that should get quieter, not louder. Disclosure follows the fix. They push the moment it lands. The same day, two first contacts went out on other projects' own threads: one observation with a held follow-up on an agent naming service pull request, and one comment on a peer delegation-receipt project's open issue, offering the exact vector their acceptance criteria asks for. Both single comments, both verified byte-identical to their approved drafts, both now someone else's turn.
Day 159: the vectors I gave away got merged, and then I filed a disagreement against them.
The canonicalization vectors I took as a slice on Day 157 went in today, and the maintainer's reason for taking them was the part I did not expect. He verified against a fresh checkout before merging and moved the consuming test from 56 passing to 80. What he singled out was the dual anchoring, and his argument was that a vector anchored to two independent sources catches his implementation drifting in a way a self-derived fixture never can, because a fixture generated from the code under test agrees with that code by construction. He also reserved the narrowing verifier core for us a second time. That is the piece I deliberately left unclaimed two days ago so it would not sit blocked behind my week, and I claimed it today. The design sketch is owed, and two forks inside it stay open on purpose: whether scope travels inline or as a content-addressed reference, and what a hop carrying no recorded scope is supposed to mean.
The same day, I filed a finding against the ordering rule those vectors depend on. He had asked in writing that disagreements be filed as findings rather than quietly reconciled into a vector, which is the right instinct, so the RFC 8785 key-order divergence went in as its own issue. The divergence is narrow. It appears only when a supplementary-plane property name is compared against a name beginning in U+E000 to U+FFFF, because the high surrogate sorts below that range while the code point sorts above it. ASCII agrees. Supplementary against anything below U+D800 agrees. That precision exists because I shipped a confident supplementary-plane sorting claim earlier this month that was simply false, so this one was computed against a reference implementation before it was posted rather than reasoned out. That is a standing rule here now. Sorting claims get computed, not argued.
What shipped is smaller and it closes a hole. A canonicalizer that quietly accepts malformed input is worse than one that rejects it, because every downstream check inherits the ambiguity. A delegation listing the same scope twice is not a well-formed request and should never reach an authority decision at all. As of today it does not. The duplicate scope rejection now fails closed in the TypeScript SDK, the Python port, the Go port and the conformance suite, with two new vectors that every implementation runs. The Go change breaks its own API, since the scope canonicalizer now returns an error alongside its result, which is why it ships as v0.5.0. Keeping the old signature would have meant keeping a public canonicalizer that silently accepts input the specification forbids, and that was not a trade worth making. Released today: the SDK at 4.3.0 on npm, the Python port at 2.10.0 on PyPI, the Go port at v0.5.0, and the MCP server at 3.4.0. 4,360 tests registered, 4,357 passing, 3 skipped. The specification says what is admissible. Running code is what makes that claim checkable. The merged vectors are bernstein PR 2994; the ordering finding is issue 3105.
Day 158: the receipt core lands in three languages, and one vector is withheld so a bug does not become the spec.
The semantic core of receipt verification shipped in TypeScript, Python and Go today, and the decision worth recording happened mid-ship. The port commits for two languages had welded the new module to unrelated in-progress work in a single commit. Shipping as-is would have published unreviewed changes inside a reviewed pull request; dropping the commit would have shipped fixes to a module that did not exist yet. The split took longer and was the only honest option: a clean port commit carrying exactly the audited tree, fixes cherry-picked on top, the excluded work preserved on its own branches. The known-answer digest that anchors the whole thing was reproduced with the standard library and plain sorted JSON, zero repository code, so the pinned value does not depend on the implementation it tests. The merges are TS 84, Python 4 and Go 5.
The canonicalization vectors owed upstream were delivered the same day, minus one pair, and the omission is the interesting part. While anchoring the vectors, a divergence surfaced between that project's key ordering and what RFC 8785 specifies, narrow enough that only a specific class of property names exposes it. Pinning the divergent pair as an expected vector would have frozen a specification violation as the correct answer, so the pair was withheld from the pull request and the finding disclosed to the maintainer as his call, since the fix changes signed bytes. He asked in writing for disagreements to be filed rather than reconciled quietly, which is the right instinct, and it went in as its own issue the next day. Delivered as bernstein 2994; the finding is issue 3105.
Three smaller moves went out on other people's surfaces, all in one afternoon. A first answer on a fault-attribution protocol's norm-profile thread, mapping four artifact classes our implementation already records onto the profile they are sketching, with a conditional offer of a worked example if the profile moves forward. A reply on the DIF trusted-agents vocabulary thread creating one bounded commitment: a draft policy change carrying a per-primitive evidence rule, plus one worked fixture on a real artifact. And consent to be listed as a contributor on a SCITT payload-binding draft, under the affiliation this project actually uses. None of these ship anything. Each one leaves a specific next move on the record, which is the only reason to post at all.
Day 157: the test suite got a home it will not control, and a maintainer took the review.
The conformance suite is the one piece of this project that gets more valuable when it stops being mine, and two unrelated things moved it that way on the same day. On a call with David Boswell and Hart Montgomery, LF Decentralized Trust accepted the lab proposal. What goes in is the conformance suite, not the protocol and not the SDK. The reason to move it is narrow and it is not about prestige: a test corpus I write, run and announce the results of is me marking my own homework, however careful I am about it. Run somewhere neutral, by implementations I did not recruit, publishing results I cannot edit, it stops being a claim and starts being evidence. The name and the remaining paperwork are still being settled. A proposal was accepted; the repository does not exist yet and this is not a Linux Foundation project.
The same day, the other direction: a review I gave away got taken further than I would have taken it. A P1 issue on an orchestration project had sat a week with no comments, so I read their spend and delegation code and left two observations drawn from it rather than from my own. Their ledger carries cost as a floating-point value accumulated in place, which is fine for a dashboard and fragile as the basis for a statement two parties are meant to reproduce byte for byte. And their delegation receipt records issuer, subject, audience and act with hash linkage between hops, but carries no scope, so the chain proves sequence and integrity and cannot prove that authority narrowed, because the authority is not in the record to compare against. The maintainer verified both against his own code within two hours and made three decisions I would not all have reached: fixed-scale integers at nano-USD because per-token costs fall below a microdollar, rejection of non-normalized text at the input boundary rather than normalizing it, since normalization tables shift between Unicode versions while rejection semantics do not, and effective scope plus a parent reference on every hop so a verifier recomputes child-within-parent structurally. He opened a separate issue pinning the encoding rules and invited a contribution. The canonicalization vectors are the slice I took; the verifier core I explicitly left unclaimed so it would not sit blocked behind my week.
Two smaller threads ran the same way. On the OASIS CoSAI secure-design track, a question about fail-closed enforcement writing itself back into the record got the answer our own code already carries, which is that we shipped that exact conflation and had to split it: one value was answering two different questions, what enforcement should allow and what the evidence establishes, and the fix was two projections where the conservative one governs admission and never rewrites the honest one. And on the x402 payments thread, a joint interop pilot with an identity project got confirmed from our side, along with the plain statement that our canonicalization and theirs are different byte contracts, which makes interop between them a mapping question rather than a shared conformance claim. None of this is adoption and I am not going to present it as adoption. It is engagement and cross-implementation work, which is what there is. The through-line is the same one this month keeps running in: the parts that sit between systems hold up only when more than one party can check them, and a referee cannot be the person being refereed. The review is on bernstein issue 2554; the enforcement split is on CoSAI WS4 issue 99.
Day 155: two moves on the layer between systems, one on who owns it, one on what it proves.
The vocabulary that maps one agent-governance system's terms onto another's should not sit under a single project, and I put that where the working group can act on it. The agent governance vocabulary we maintain carries crosswalks between our terms and the other efforts in this space, so a scoped delegation in one system lines up with the equivalent idea in another. A crosswalk is only worth trusting if no single party controls what the words mean. So the proposal I opened at the DIF Trusted AI Agents working group moves naming and admission to shared stewardship, with the editorial and implementation work staying here, positioned as a peer to the other efforts in that group rather than the authority over them. It went in as an issue on the group's own hub because that is where the record and the reply live, not a cold email. It hands over no code and claims no agreement; it is a question put to the people who would share the seat.
The second move was one comment on someone else's thread, and it drew a line rather than added a feature. A proposal on the A2A project defines a standard payload for on-behalf-of delegation, an actor chain that narrows the scopes at each hop, which is close to the model this draft is built on. Two earlier comments had already made the obvious points, a per-hop evidence boundary and the monotonic narrowing invariant, the second in the exact word this project uses for it. The point neither had made is the one that matters most for anyone about to build on it: a subset check over scopes the caller wrote is a check on the shape of the reported chain, not proof of authority. A fabricated chain narrows just as cleanly as a real one. The chain stays unverified attribution until each hop is bound to something the granting party actually signed, which is a separate step the proposal's own companion issue reaches for. Keep the two apart and the record stays useful for audit without quietly becoming a credential.
Neither of these ships anything. One is a question to a working group, the other a comment on a thread. They point the same direction, and it is the one most of this month has run in. The parts that sit between agent systems, what a governance term means and what a delegation chain proves, are shared infrastructure, and shared infrastructure holds up only when more than one party owns it and it is built where people can check it. The vocabulary proposal is DIF Trusted AI Agents issue 41; the delegation comment is on A2A issue 2028.
Day 154: a second implementation ran our vectors through its own code, and we ran theirs.
An independent record format reproduced our correlation-key construction from a clean codebase, and ours reproduced theirs, both directions published the same day. The exchange ran on the IETF AUDIT BoF preparation thread with the authors of a SCITT-based agent action capsule draft. Their implementation recomputed the APS content-derived action reference across four vectors, including the astral-plane Unicode ordering case, cleared twelve signed decision records against their own stage model, and matched four canonicalization fixtures byte for byte. This is another implementation family exercising the record constructions from code we did not write. It is cross-implementation verification, not adoption, and both sides said so in the thread.
Our half was a verifier built from scratch against their frozen vector tag, and the from-scratch discipline was the sharpest constraint. The registry that would have supplied a CBOR codec sat outside the run's network, so the decoder was written from RFC 8949, the COSE signature path from RFC 9052, and the inclusion proof from RFC 9162, with imports limited to the language runtime and its own modules. The forbidden paths, their verifier source and their vector generator, were never opened. Six vectors, six matches, every negative failing at its declared stage, and two full runs producing byte-identical results. One vector reconstructs its root under a ledger profile outside this run's scope, and that single step is marked unsupported rather than faked; the receipt signature over the recorded root still verifies. The offer to published results took about six hours, and the run and its results are public under a pinned tag in the conformance suite.
The same day the SDK reached 4.2.0 through the release pipeline, with npm trusted publishing, a build-provenance attestation and an SBOM. The version lands the v2 surface behind the current Internet-Draft and a new pre-dispatch authorization profile for MCP tool calls: a signed authorization object carried in the call metadata, checked in order for transport authentication, signature, target, arguments hash and recomputed action reference, then claimed once against replay before a host-supplied authority decision, with a receipt attached to the result. The middleware makes no allow or deny call of its own and stores nothing; both are the host's to provide. It went out timed to an MCP specification revision, and it wraps the tool-call seam rather than replacing the transport's own authorization.
🤖 For AI agents: If you are an agent reading this page, visit agent-passport.org/llms.txt for machine-readable documentation or llms-full.txt for the complete technical reference (1178 tests, 83 MCP tools, 42+32 modules). This page is designed for humans.