Dev log

Building Agent Passport System, day by day.

Day-by-day record of building the enforcement and accountability layer for AI agents. Bring your own identity: did:key, did:web, SPIFFE, OAuth, native did:aps. Started February 18, 2026. 6,003 tests, nine papers, IETF draft. Open source. Full surface area: 152 MCP tools.

See the full picture on the roadmap · every ship across protocol, product, research, comms, and ops with dependency arrows.

<<<<<<< HEAD

Day 220: a rule in our own draft that nothing could fail, and 134 cases that became one file

Section 3.3 of the draft says "Each action selects one root-to-leaf authority chain. A verifier MUST NOT union scopes or budgets from multiple chains." Every authority entry point in the SDK already took one chain, so the rule had nothing to bind to. An implementation that pooled three held chains and one that selected a single chain looked identical through our API. 7.2.0 adds a selection surface where no union is held by construction. One private function judges one chain, a result names one chain id and never a set, and a caller who concatenates two chains into one array is refused by name before verification starts. An unknown revocation answer comes back undecided, never refused. Python 4.2.0 runs a byte-identical copy of the 19 parity cases.

Seven more modules shipped in both SDKs, opt-in and marked proposed rather than specified: lifecycle state, activation, bounds, capability binding, status coverage, authority state, suspension. Implemented does not mean specified, so each new record type carries a proposed: namespace instead of aps: and every reason code stays module-local. The state vocabulary reports six artifact verdicts alongside chain verification and never merges into it, keeps what an enforcement point decides at one authorization boundary as three separate outcomes, and makes a not_established verdict name which of source, freshness or coverage was missing.

The lifecycle repository went from v0.2.0-draft to v0.3.0-draft over four cuts. Cases are grouped by the authority question each one asks, in 18 semantic families, not by the research leg that found it. cases.json is the single source of truth, both case files are generated from it, and CI fails on drift. 134 lifecycle cases, 126 verified, 7 reviewed hypotheticals and 1 candidate, with 36 boundary cases kept apart. The precedents behind them come from law, institutions and documented incidents. None of those sources says anything about AI agents. The translation is ours.

The vectors landed in the lab under tracking issue #122, seven stacked pull requests and three follow-ups, #123 to #132, twenty-eight new fixture families, each labelled candidate against proposed text rather than against a published specification. Running them against the SDK gave the number I care most about today. Across 29 families, 633 vectors and 635 decision units, an SDK call reproduced the expected result for 50 units under 7.1.0 and 161 under this build, a move of 111 units across 12 families with 0 fail and 0 regressions. The other 474 are not_supported, so no SDK API decides them and the family harness supplies the deciding step. That is not a coverage score.

Three things went outward. Thirty-four lifecycle terms went into the vocabulary repository as a separate maintainer-governed file, #174, every definition a verbatim quote of 40 words or fewer from the concept document. The AgenTrust delegation verifier merged in #216. The A2A Go maintainers merged a canonicalization fix, #441, which left the Python SDK on the other side of the line, so I opened #1278 there. Its canonicalizer strips empty strings, lists and dicts before signing, with no exception for the fields section 8.4.1 marks REQUIRED, so a card signed by one SDK fails verification in the other.

Two comments on CoSAI WS4. On #189 I put the decision-to-effect work into a candidate corpus in lab #120, eleven cases around exact call, bypass and effect verification, and where the evidence is not enough the result stays not established and says what is missing. On #205 I answered three open questions. I would not re-parent an existing delegation under a new ancestor. The descendant can carry on under a new grant from a currently authorized principal, and the old chain stays revoked. Once authority is revoked it cannot authorize another effect at the next authorization boundary, and what happens to the operation depends on the operation. I joined the workstream properly today, mailing list and iCLA.

Day 219: a late approval reached the wrong invocation, and three jobs that died waiting

Goose issue #11739 says a tool approval is matched only by the provider's request id, so a late answer to a cancelled prompt can land on a new invocation that reused the id. I reproduced it on current main in the generic confirmation router. A registers, its receiver is dropped, then B registers the same id. By that point the router has already pruned A's closed sender, and a late AllowOnce meant for A's prompt still goes to B. Better cleanup would not change that, because the reused id is all the answer carries. The same test with a different id passes. The ACP path is keyed the same way, but I could not drive two real stream invocations through that interleaving with the existing fixtures, so I claimed nothing there. The fix belongs to the maintainer who owns the issue. The tests are on the issue.

OpenClaw #155946 needed a live run on the Claude CLI backend. A hook matching db__query on a configured MCP tool never fired, and the tool ran. A hook matching mcp__db__query fired, blocked, and the tool never ran. So the name reaches the hook with its transport prefix, and a block that does match lands before the server executes. That went on the issue as a three-row table.

The App Defense Alliance merged a tool provenance section into its AI agent specification. It requires refusing revoked tools, but its test never says whether the revocation happens after the session is set up, and nothing in the section says what should happen when revocation status cannot be established. Those are two issues now, #497 and #498. I also left a note on a reformatting pull request that had been cut before the section merged and would have dropped it.

Two mappings went into the conformance lab, #118 for that section and #119 for the policy engine work in IBM ContextForge. The ContextForge record carries a script that runs two of their source files unmodified at a pinned commit and refuses to run on anything else. With a 60 second TTL, a decision cached at t+0 was still served by a second worker at t+118.5, because a Redis hit starts a fresh in-memory TTL. Whether that TTL is meant to be the propagation window their issue asks for is not established, and the record says so.

Three of my headless jobs died the same way today. A clone, a test loop and a build ran past the tool's own five minute timeout, the harness moved them to the background, and each job ended its turn waiting for a notification that never came. Telling them not to did not help. Asking for a long timeout on the call itself did, and that is now the first rule in the prompt templates.

Day 218: two releases, and functions our own docs pointed at that nobody could import

A lab job stopped because it could not find a function. issueAuthorityDelegation and issueSubAuthorityDelegation were documented in the SDK and not exported from the package root, and the same was true for scope comparison and budget. Anyone following our own docs would have hit it. 7.1.0 exports them, next to direct authority revocation and opt-in predecessor binding. It published on npm with provenance. The post-publish check got a 404 while the registry was still processing and passed on the rerun two minutes later.

Python 4.1.0 went out through a gated workflow. It checks that the tag equals the version and that the commit is on main, runs the full suite, 2,127 passed and 6 skipped, and only then publishes. The wheel and the sdist each carry a PEP 740 attestation, and a clean install from PyPI reports 4.1.0.

Twelve pull requests merged in the conformance lab, #106 to #117. Most are the revocation and authority cases from the last weeks. They cover a revoked ancestor failing a deeper descendant, one agent identity under an old and a replacement chain, cached authorization after revocation, a denial that stays a denial, an approval as a single-use permit, historical key selection after rotation, one chain per action on the real scope and budget code, and issuer refusal with distinct expiry and revocation codes. The two families pinned to the SDK were re-recorded on 7.1.0, and the only differences allowed were the version string and one new field that reports predecessor binding as not checked.

Three went outward. A review of the WIMSE agent audit record draft, on #144. One candidate case against the AuthZEN ARAP profile, on #663. And a mapping of lab families to the freshness and revocation rules in AAIF Identity and Trust #5, our first contribution in an AAIF working group. The preview image for agent-passport.org also changed. The old one still showed a retired wordmark and stale numbers.

Day 217: the gateway is public, and a replay fix counted after a restart

The hosted gateway had a bug in its authority check. A sub-delegation could keep acting after a delegation above it in the chain was revoked, because the grant check looked at the direct link and never walked the chain it was issued under. I reproduced it with a four-hop chain, then fixed it by binding every sub-delegation to its inbound chain and checking each ancestor at evaluation time. Before deploying I checked what production held. No active chain was longer than one hop, so nothing in use depended on the broken path. The fix is deployed.

The gateway itself went public today as agent-passport-gateway. The history is the real one, 245 commits, with deployment details, tenant data and internal notes removed from every revision. Rewriting history changes every commit id, so the only signed commit is the release commit on top. An operator email that was hardcoded into two migrations became a setting with no default.

Preparing it turned up an open issue in the MCP server. Capability tokens were marked as used in process memory, and the hosted bridge starts a new process per session, so a used token could be redeemed again in another session or after a restart. The hosted endpoint already required an API key. Version 6.1.0 records each used token as a file on disk, created exclusively so only one process can win, and the hosted service now keeps those files on a volume. I marked it fixed only after a test on the live service. Redeem a token, restart the service, confirm a new process is serving, redeem the same token again, and get it rejected.

The revocation and handover work of the last weeks is now its own repository, agent-authority-lifecycle, with each claim labeled specified, tested, candidate or open. Three cases went to the conformance lab. One revokes a delegation above a chain. One hands a departing sponsor's agent to a successor, where the same agent continues under a newly issued chain while the old chain fails. One comes from an OWASP discussion, where a denied action has to stay denied through a retry, another tool or a delegate. The handover case went to the CoSAI thread on decommissioning agents, the denial case went back to the OWASP thread it came from, and a delegation verifier went to the AgenTrust integrations, so receipt evidence and delegation authority stay two separate results.

Day 216: two releases, a lab that runs on them, and a tag I could not push

Seven questions that draft-03 either left open or that the two SDKs had answered differently went to four independent reviews this morning, and I ruled on each. The required signature set decides whether a receipt is valid, and any extra signature is reported beside it without moving the result. A key that cannot be resolved, including one whose bytes are malformed, is indeterminate and never a failed signature. An empty scope_required is rejected unless a caller opts in. And no input grammar that happens to live in one SDK becomes protocol. Two reviewers changed position during the day because of evidence.

By evening both SDKs were on main and published, 7.0.0 on npm and 4.0.0 on PyPI. Each was installed clean from its registry and checked against the artifact it was built from before anything pointed at it. The TypeScript suite runs 5,491 tests with none failing. Then the conformance lab caught up. Every one of the 26 receipts in its oracle-safety corpus failed the new stage rules, which is what a correct re-mint is for. After the re-mint none fail, and the suite now depends on exact 7.0.0 instead of a branch.

That unblocked the pull request I pulled back on Friday. The TRACE action receipt integration now verifies a real draft-03 action intent, and rebuilding it found two things the old verifier got wrong. It would have reported malformed key material as a failed signature, which claims wrongdoing the evidence does not show. And a third party could append a failing signature to a published receipt and put a failure reason on a verified result. Both are fixed and pinned by tests, and it is back with the maintainers.

One failure of my own. 7.0.0 was published by hand, outside the release workflow, so it carries no npm provenance attestation, which 6.0.1 did. The bytes match and the registry records the commit it came from, but that is a weaker property than provenance, and a published npm version cannot be amended. I cannot tag it either. Both release workflows publish on tag, and the TypeScript one refuses an existing version without provenance before it creates anything. The tag waits for a way to add it that does not run a publication pipeline, and the next release goes through the normal path.

Day 215: a stopping rule, written after the ninth cycle

The reconciliation called a job finished when neither reviewer had an unresolved correctness finding. On Saturday that rule ran one job through nine whole-surface review cycles in eleven hours. Each pass could find something, so each pass did. The eighth found a real defect, a ledger that accepted a reservation after a cancel, but that belongs to hardening, and conformance had been done hours earlier. The rule now allows at most two whole cycles per job. A repair is verified by reproduction, the suites and one review scoped to the repair. Anything found after that is classified and carried, and it does not reopen closure.

Two other structural changes came out of the same afternoon. Closure became a gate at the end of a job instead of its last planned batch, so a job is judged by whether its whole surface matches the draft. And a fourth reason a change may exist got a name. Implementation hardening covers reliability and resource safety on a draft path where the protocol does not prescribe the runtime property. One commit carries one reason. A Python commit that mixed two is recorded as an erratum and left as it is.

Outside the SDKs, three things moved. The reviewer on CoSAI #189 resolved his original objection and asked that coverage tied to a claim be explicit in the rule text. I had a long reply drafted. What went out was the edited cases as suite pull request #98 and one line pointing at it. On the PriorSeal thread the adapter's author had held his work waiting for our fixtures, so four inputs went up on a branch cut from 6.0.1, shaped against the draft rather than against what I had promised on Friday. And the insumerapi crosswalk refreshed to what the API issues today, merged once a clean clone passed and the live key set matched the snapshot committed with #166.

aeoess.com changed too. The entry chooser is gone, and the home is now a single interactive line above the four projects.

Day 214: I opened a pull request and pulled it back the same afternoon

On Friday morning I opened a pull request against the TRACE integrations repository that verifies an APS action receipt as evidence that an agent issued an action. By the afternoon it was back in draft. Read against draft-03 section 5.3.1, its fixture was issued by a gateway where the draft requires the acting agent, and its result was a free-form object where the draft fixes one. The SDKs I built it on enforced neither rule. That is a problem in the SDKs, so the fix belonged there.

So the week's job started, reconciling both SDKs to draft-03 rule by rule before anything else was allowed to depend on them. The first pass was a matrix of every rule against both implementations, and it caught one mistake of mine before any code moved. An earlier note said the Python action reference helper was the draft's section 4.2 form. It is a different digest. It hashes camelCase keys, a sorted scope array and a timestamp truncated to seconds, and the draft's form uses snake_case, one scope string and milliseconds. Following the note would have put the draft's label on the wrong digest, which the draft itself forbids. The helper stays, documented as pre-draft.

Two rules came out of the day. AuthorityDelegationV1 is the draft's authority record, and the older Delegation family in both SDKs is a compatibility surface that external verifiers should not target. And every change in the reconciliation declares why it exists: a repair toward the draft, hardening of an existing path, or evolution toward -04. A new rule is never presented as an old one.

The rest of the day was other people's work landing. The wallet_state re-date in the vocabulary merged once its keys and verifier were pinned next to its bytes, a clean install running 23 of 23, and the approval corrected an error in my own earlier review of vector 22. Cross-stack vectors from mcp-audit-gateway merged into the lab. And PriorSeal's author agreed to bind APS decision evidence through a sibling adapter on published surfaces rather than a fork.

Day 213: a pack pinned correctly went stale in four days, and a re-date pinned the bytes but not the keys

Two independent fixture packs for the conformance lab had been sitting in a pull request since the eleventh, conflicting with main. The rebase itself was the boring kind: one generated inventory table, two additive hunks, both sides rows that main had grown since the branch point. I kept all four rows and let the generator confirm it would have written the same thing. Both packs came through byte for byte. What was not boring was what one of them had become while it waited. The AAT pack is pinned to draft-sharif-agent-audit-trail-03 and carries a case it deliberately leaves unresolved, a genesis record signed by an independent recorder, because -03 says the recorder should use its own key and then verifies with the agent's. On September 15 the author published -04, which adds a signer key identifier and names the signing principal for each recording mode. The ambiguity the fixture records is no longer current. The pack is still correct, because it says what -03 says, but a reader arriving today would take a resolved question for an open one.

The fix was one paragraph, not a rewrite. The pack stays pinned to -03, the fixture stays unresolved against -03 only, and a dated note says which later revision closed it and how. No fixture, expected result or provenance line moved. Two things about that small edit are worth writing down. The pack's checksum manifest covers its own README, so a documentation change is a two-file change or the manifest fails. And my commit went up without a sign-off, which turned the DCO check red. The agent running the mechanics stopped there instead of amending in my name, because a sign-off is a certification, not housekeeping, and I gave the word explicitly. The amended commit has the same tree hash as the unsigned one. Merged by rebase, main carries the exact tree that was tested, six checks green on that commit.

The vocabulary got a harder review. A contributor re-dated a proposed signal at its sunset, as I had asked in the tracking issue, pinning his evidence by git revision, tree object and discovery-copy digest rather than by a URL that can move. All of those pins check out. I resolved the commit, matched the tree object, fetched the discovery copy and got the same digest byte for byte. Then the boundary problems. The note says each vector is an artifact exactly as the API issued it, and twelve of the twenty-three are deliberate derivations by their own descriptions. The runner that reports twenty-three of twenty-three imports the issuer's own verifier, so it is an author-produced result inside one trust domain, fine for the criterion, wrong to leave reading as independent. The two packages it installs are unversioned while the README documents that older versions accept two of the negative cases. And twelve vectors resolve keys from a live JWKS endpoint with no snapshot in the tree, so the signed artifacts are frozen and the keys that verify them are not. That is the same failure the re-date was supposed to close. Changes requested, status and issuer count left exactly as they were. One trap from the review itself: "eleven must be refused" is only true under a stated predicate, because every frozen vector now expects an expiry failure and two more carry a signed false verdict that is not a verification failure at all. Without the predicate the honest count is thirteen.

Two things closed without needing me to be right. The AuthZEN issue where I had offered a binding_hash vector set in July closed today because a maintainer's pull request landed an official directory with seven known-answer vectors and cited the larger set while doing it. I had read the closure as a loss for about an hour. It was a landing. My comment there is one sentence saying the larger set still exists if anything is uncovered. And a paper on arXiv characterizes APS delegation as a rooted tree with a single parent identifier, reporting that a tree-cascade baseline revokes every shared agent its edge-withdrawal method preserves. I checked that against the current draft. It is fair. Each delegation record has one parent, one chain authorizes an action, scopes and budgets are never unioned across chains, and revocation cascades through descendants. The limit is at the record level, not the agent level, an agent holding two unrelated chains keeps the second when the first is revoked, but a descendant created under the revoked one does not. No correction owed. A design question logged instead, and not the obvious fix of making the parent field an array.

One failure of my own to end on. To settle whether a client bug lived in the API or the client, a raw signed publish was sent to production from outside every tool we ship, and it landed: 201, readable, on search. The API is fine. The test card is not, because the throwaway key that signed it was deleted with the script that made it, and every removal verb on a card requires that key. A known test artifact now sits on the discovery surface until its expiry, or until I authorize a database write outside the protocol. The rule that comes out of that is small and I should not have needed it: a key used for a production write is kept until the cleanup is done.

Day 212: a rule sent as a question, and a requirement closed as too early

The evidence sufficiency rule went to CoSAI WS4 #189 for the reviewer who found the four defects to confirm or tighten. A negative claim over a scope passes only when field visibility and observation completeness are both established for that same scope, whether the field is absent or present but empty. If either is missing, the result is not_established with the unmet obligation named. Malformed input, unsupported verification, parser failure and internal errors stay outside the verdict. I posted it as a question. The merged cases stay candidates against proposed text.

The AEP run record's merge values went to the WasmAgent maintainer for his external evidence ledger. My first draft nearly sent git blob identifiers where content hashes belong, and nearly called a one-parent rebase result a merge commit. Both were fixed before posting. He recorded it with the two anchors and confirmed the assurance ceiling did not change.

The OWASP AISVS maintainer closed the binding requirement in #1152 as too aspirational for 1.01, to keep the backlog short, and left the door open. The test had already been tightened after a reviewer's correction, and that reviewer had accepted the requirement. I posted nothing after the close. It does not reject the finding, and it does not accept the requirement. A confusing line in the SDK's papers index was also fixed in #170.

Day 211: two records merged, and a manifest that cannot live inside the commit it names

The WasmAgent AEP run record had cited a moving branch. I repaired it to pin the component tuple, the publication commit and the lab revision as immutable references, and the four result files did not change. It merged in #94 as six commits. The claim is unchanged. It covers one independent layered run against the tuple published for one certified snapshot, with the JS and Rust native checks and the lab's semantic recomputation reported separately, and the semantic layer labeled author-produced.

Pinning it also explained an old manifest in the certified tuple. The protocol component SHA inside the certified tuple still carries the previous manifest, because a manifest that names a commit changes that commit. The maintainer confirmed in #92 that the model is the tuple plus a separate publication commit on the protected main branch. Both are now recorded side by side, so the next reader does not take the old manifest for an error.

The CoSAI WS4 #189 candidate cases merged in #91 after the reviewer reran the repaired head. He had reported four defects, and a second participant verified the repairs. The coverage premise is now explicit. Unsupported verification comes back as a non-verdict. A missing gating descriptor is a structural input error. And the harness is read-only, so reproducing a result cannot repair the artifact it checks.

Day 210: the fix was in the code and the shipped copy still taught the defect

An introduction you asked for was reported back to you as one someone else had asked of you, and the note you wrote yourself was attributed to the other person. On a network whose entire subject is who wants to meet whom, that is not a display bug. The server always knew the direction. The client relabelled it on the way out, and the labels it used were the ones published two days ago.

The code fix was done by morning. Three things would have let it ship with the old contract still in front of users. The skill that ships inside the package still named the fields the fix had replaced, so an agent following it reads a field that does not exist and reports an empty inbox. The test that guards that skill still carried the three dead names on its approved list, so the old contract could have been written back in without a single failure. And the suite exited 0 with all thirteen end-to-end tests skipped, because a default path did not resolve in that layout, which means the pass line was a claim about code that had not run.

All three are closed. The dead names are removed from the allow-list rather than renamed, so their return is a test failure. The pre-test guard now resolves the same path the test resolves and refuses to run without it, with a named flag for skipping on purpose.

It went out as a major with no compatibility alias. The old direction values were not merely renamed, they were inverted, so emitting them beside the new ones would keep teaching the defect to exactly the clients still reading it. The release note says the uncomfortable half of that out loud: an install that does not upgrade keeps working against the network and keeps showing every introduction backwards.

The publish tooling had its own version of the same fault. Its closing banner printed the procedure for initializing the thirty day compatibility clock for whatever version was being released, which for this release would have instructed the operator to restart a window that belongs to the previous one. The server refuses a second initialization, but a backstop is not a correct instruction. The banner now branches on the anchor release and prints do not touch it, with a verify-only check, for everything after it. The same release carries an errata correcting an earlier note that printed the thirty day cutoff where the publication instant belonged, a month apart.

On the conformance suite, a contributor's cross-stack family came back with five text edits done and three questions. Two were straightforward. The third was whether two exercised failure outcomes could land without an independent record, and I answered that they could not and that we still owed the record. That was wrong. The record already existed: the independent run we commissioned executed those checks and passed them, and the verifier derives each outcome from the record data before comparing it to the file's published code. What was narrow was our own earlier attribution, not the run behind it. The correction was an appended record assigning those layers to the run that made them, with the original claims left as published, and it does not extend to the third failure code in that class, which no vector raises.

Day 208: Mingle shipped, and a verifier that exits zero on seven mutations

Mingle shipped. Stage 2B closed with one withdraw action that removes unreleased contact, live fit and First Step together, and the race with a concurrent write resolves at the write in both orders. Stage 2C cut forty-six tools across three generations down to eight, with the old set behind a flag.

The run that mattered was two principals through real MCP processes against a real API. It found six call sites the mocks had passed. Review found seven severe items, two of them security, and all were fixed before publish. The server went live first. Then I published mingle-mcp 4.0.0 by hand under WebAuthn, and the server records that publish's registry timestamp as the start of a thirty day window for the legacy tools, so a later patch cannot move it. 4.0.1 followed the same day and changed only copy, six overclaims in the shipped skill and README.

A ScopeBlind maintainer merged our driver in #1 and asked what his three published files establish. I ran them locally without changing any input. The published verifier exits zero with top-level valid true on seven mutations. A log digest field is checked for shape and never compared to the file. The published command leaves out the provenance flag, so it never reads the Sigstore bundle. And the 28 of 28 agreement covers tool identity only, because the receipts carry no principal, context or input. I posted it as findings with questions back to him.

Day 207: nine fixes to Mingle, and the review of them found a reveal that released both values

A product review said Mingle had gone deep before it went wide. Nine correctness items shipped. Among them were an acceptance email that never fired, an opt-in with a hidden default, semantic search ordered by recency, no atomic way to replace a card, a purge the copy promised and the code did not do, a private key any local user could read, and bare sweep routes.

The review of that batch found the worst one. When one party revealed, the release handed out both parties' exact values. The fix tracks a releaser set for each dimension separately.

A binding audit went through forty-three mutating routes and recorded which request and commit signatures actually bind what the receipt claims. The fit and v2 containment gates are deployed. Two independent adversarial reviews approved the Stage 2A design. The MCP stayed unpublished, because the server it depends on was not live yet.

Day 206: the driver rewrote the policy it was being tested against, and the correction went upstream

A maintainer asked a simple question about our conformance driver: his fixture policy had not been valid Cedar until that morning, so why had our run reported a clean allow instead of an error. The answer was worse than an error. An earlier revision of the driver rewrote the policy text before evaluating it, turning an idiom Cedar rejects into one it accepts. That was disclosed in the pull request body at the time under a note claiming the semantics were unchanged. Measured against the bytes on disk, that claim was false.

Three configurations settle it. With the rewrite, the pre-fix policy gives allow, allow, deny, allow. Without the rewrite, the same policy gives allow, deny, deny, allow, because Cedar skips a policy that errors and returns a deny with the detail only in diagnostics. And the deny on the third case came back with an empty reason list: the forbid clause never evaluated at all, which means the only negative vector in the corpus had been passing for the wrong reason on any engine that skips erroring policies rather than rejecting them.

The shim also happened to produce exactly the clauses the maintainer later wrote himself, which is why it stayed invisible for three months. A rewrite that guesses right hides itself.

The driver now treats any Cedar diagnostic error as a failure rather than a decision, checked by putting the broken policy back and confirming it exits non-zero instead of emitting receipts. He merged it, reproduced the result on his own machine with cedarpy 4.8.7 and the published SDK, credited both findings, and wrote the rule into the upstream README: the policy under test is the on-disk bytes, a driver does not rewrite them for evaluation, and an evaluator that cannot run the policy must fail rather than deny. A separate finding went to its own issue, since the reference digest in the expected chain matches no version of the fixture policy that has ever existed in that repository.

Elsewhere, a boundary statement went to the CoSAI workstream that asked for it, on what a verifier may conclude when evidence is absent. Three layers stay apart: a runtime outcome, a receipt-chain outcome for a disclosed gap, and the property verdict. The verdict is returned after the applicable verification has run when the admissible evidence justifies neither pass nor fail, it must name the unresolved obligation, and it must not absorb verifier failures, because otherwise a broken verifier becomes conformant by returning the third value. For non-bypassability the outcomes are asymmetric: an observed alternate path can justify a fail, while incomplete coverage only blocks a pass. That asymmetry was wrong in my first draft and caught in review.

A gap in the governance vocabulary got its own issue. Six action classes, all digital-native, none of which names operating a sensor or an actuator, sitting beside a context dimension that already reads physical-world state from signed sensor readings. The physical world is on the input side of a decision with nothing matching on the action side. No new values are proposed, and one of the three legitimate outcomes is that the six are deliberately broad and nothing changes.

And the first reverify deadline in the crosswalk registry's history expired. A principal attestation on one cell reached its date with nothing to renew it: the upstream sample has been unmerged since May and no third party has attested. The tempting fix was to soften the control to a warning, which would have overridden a documented fail intent on the first day it ever fired. The claim was withdrawn instead, taking its deadline with it, with the attestation date and the reason kept in the notes. Absence is already a valid state for a registry claim, so nothing in the rules had to move.

Day 205: a pass before anything ships, and four errors that shared one shape

Claims, Context, Copy is now a required pass before any deliverable is shown. Claims means reverifying every fact, number and evidence chain against primary evidence rather than against the note that recorded it. Context means rechecking repository state, prior decisions, scope and public promises, and asking whether a true sentence is misleading where it is about to land. Copy is the writing pass. It has already caught more in two days than the habit it replaced caught in a month.

The second rule came from four errors in a single day that turned out to be the same error. The exit code of head read as the exit code of the program it was piped from. An empty regex match treated as proof that something was absent. A test selection matched by path reported as the repository's named contract suite. And a run another agent performed becoming I ran it in a draft headed for a standards thread. None of those was a lie and none was careless in the ordinary sense. Each one substituted something adjacent to the evidence for the evidence itself.

So the rule names the shape. A technical conclusion needs the exact command, its own exit code, and output that shows the property actually being claimed. A rerun on the same machine is a reproduction. An independent verification is someone else's machine.

Separately, AI attribution trailers are now stopped by a pre-push hook rather than remembered at review time. That moved the rule from a thing we intended to a thing the tooling enforces, which is a different kind of rule.

Day 204: a proof built to attack our own claim, and an audit workstream closed by reading disk

The admission-evidence proof closed on a real Gateway. Seven adversarial cases against a current runtime, constructed to break our own claim rather than support it. Two were established on a real host, three were legitimately local, and the stop condition held for the rest rather than being argued around.

The sentence it earns is narrower than the one we wanted. The gate proves what it saw and admitted, and a later hook can change the arguments the tool actually receives, so what can be produced at that boundary is bounded by where the boundary sits. No new behaviour shipped out of the proof, which was the point of building it. The upstream bug found along the way went to its own issue with its own reproduction.

The D1 to D9 security-audit workstream closed the same day, with each item verified against disk rather than against the index row describing it. The Ed25519 admissibility rule is the strict rule in all four SDKs. One item was ruled not to be a project at all: a dead helper with misleading labels, to be fixed the next time that surface is touched rather than carried as a security lane. The corpus-gap row stayed open and unblocked, because the vectors it needs finally have a shipped rule to encode.

Day 203: a proof that was wrong about itself, and a network that never existed

The second OpenClaw fix was merged this morning. It had been sitting in conflict since the maintainer put his own commit on the branch; he merged main into it himself and then merged the pull request, authorship kept. By the afternoon his runtime-floor change had landed too, 84 files raising the minimum Node to the first releases with lossless SQLite reads, and it closed the embedded-NUL issue as fixed with the report that started it named in the first line of the body. Two merges and one repo-wide change from a week of reading their source, which is a better return than any comment I could have written there.

The third fix nearly did not happen, and the reason is worth writing down. The completion path drops a requester agent id that spawn had already persisted, so on a multi-agent Gateway a finished child re-derives its owner, throws, and retries forever. The unit regression was clean. The review bot asked for a real Gateway trace, which is fair. The first real-Gateway run came back saying the fix changed nothing: same error, same count, with the line and without it. It was convincing enough that I started planning a withdrawal. The second run found the cause: the repository stamps its build with the git head, so an edit-run-revert cycle reuses the mutated build for the run that is supposed to be fixed. With the stamp deleted on both sides, the restored-row case fails without the line and delivers with it, three runs each way, and I reran both sides myself before believing either. That e2e is now the second commit on the pull request. The lesson is not about OpenClaw. A mutation pair without two rebuild lines in the log is not evidence, and the person most likely to be fooled by that is the one who wrote the fix.

Mingle got its first honest look at production. A read-only query on the live database showed that the 579 published cards the stats page had been reporting were seed scripts run in July: 504 principals with auto-generated ids, no links, no intros, arriving eighty in an hour. The twelve v3 cards were a test run of mine. The counter that read zero matches was a legacy field the newer code never touched. So there was never a matching failure to diagnose and never a retention failure. There was never a network. The server was rebuilt anyway, because it is the right server for the first ten people: expired is no longer spelled withdrawn, legacy 48-hour cards still disappear on schedule with their expiry logged after the delete and never inside it, every card transition goes into an event log, the stats come from the real tables, and a new match is pushed to both sides through an endpoint the client can poll. Nothing a user signs changed. It went live tonight.

Then the landing page was cut from 424 words to 185 and now says the product in one screen: find people through your agent. The MCP client polls for pending matches at session start, tells the difference between a card that lapsed and one you withdrew, and nudges before expiry rather than after. And I asked the OpenClaw community, in their own feature template, whether a persistent agent should hold a standing intent and only interrupt when both sides' agents agree there is a reason to talk, with the existing implementation offered for contribution and the promotion problem stated in the open. I had argued that morning for waiting until real introductions existed. I still think that was the safer order. It is up, and the answer will tell us more than another week of waiting would have.

Smaller: agentrust-trace 0.10.0 fixed the schema defect I reported and, by validating properly, exposed that the APS mapper's record fails the full verifier while passing the Level 0 grader, so the question of which one defines Level 0 is now an issue on their tracker; a Microsoft repository carried an invented expansion of my GitHub handle and a citation to an Internet-Draft that does not exist, corrected in one comment; and the proposed A2A identity-extension repo is not happening, because a repo nobody maintains is worse than a thread that works.

Day 202: the site got a second palette, and three more doors opened at OpenClaw

agent-passport.org now opens in blue, white and yellow. The night palette is still there, one click at the top right, and it did not move: every color on the landing page was converted from an inline value to a token, and substituting the night values back reproduces the old file byte for byte. The screenshot pass caught three things the byte proof could not, two shades of lime and one of black that the site had been using inconsistently, and each got its own token so neither theme lost it. Bright has one tier of subdued text where night has four, because the second tier fails the contrast floor on this blue; hierarchy is carried by size and weight instead.

The first OpenClaw fix, atomic writes on config recovery, was merged by the maintainer overnight with his own second commit on the branch. The second is in: a reply sent from a webchat session to a Google Chat space was going to a lowercased space id and failing with a 403; the canonical id was sitting in the session's stored delivery metadata the whole time. The first review asked for two things: proof through the transport, not through a URL builder, and a lookup that could not fall back to enumerating the session store. Both are answered with running tests, including a mutation that puts the old reader back and counts four store enumerations, and the review bot now rates it ready for a maintainer. What the change does not do is stated in the body: the exact-row accessor still runs the repository's canonical-key validation, which can scan the session table once on a fresh handle.

The third is a two-line renderer fix. gateway status could print a running service directly above a failed probe of a different Gateway without saying they were different targets; the data already carried the distinction and the renderer dropped it. The fixture that reproduces it was already in the suite. The review bot asked for real CLI output rather than the fixture, and the documented --port route produced it on this machine without installing anything. Both fixes are now rated ready for a maintainer.

The fourth door is an issue rather than a fix. Four session tests fail on Node 22.23.2, a runtime the project's own engines range and contributor guide endorse, because that Node's SQLite binding drops the tail of a text value at the first embedded NUL while Node 24 keeps it. The stored bytes are identical on both. The review bot traced it to the Node source lines and asked the maintainers whether to preserve Node 22 with lossless reads or narrow support. Within five hours the maintainer's own tooling opened a draft candidate that raises the runtime floor to the first Node releases carrying the upstream fix, with the compatibility alternative named and rejected in its body and the issue left open for his decision. The lossless-read alternative is scoped on our side and stays there unless it is asked for.

On the conformance side, attenu-guard released a nineteenth envelope vector and two people who had written neither the vectors nor the checkers scored it the same day with their own verifiers. The lab record for the new revision pins the file from four published sources, runs the two authors' own verifiers unchanged, and records the row-19 mutation: judging an envelope before the duplicate check makes the row report a bad signature and leaves a defective envelope witness-signed, which is the state the row exists to catch. The lab running those two pinned verifiers is the first pair of runs on this family that count as independent under the lab's rule.

Day 201: five conformance records in a day, and what the word independent is allowed to mean

The hosted MCP bridge at mcp.aeoess.com had been running inside the range of the advisory we published on Day 200. It now runs MCP 6.0.1 on SDK 6.0.1. Before the push, the tool manifest was diffed at both pins: 152 tools, no required input added anywhere, and exactly one behavioural change, which is the intended one, a timestamp without a zone is refused instead of accepted. Anonymous access still gets a 401. The authenticated smoke is a separate step and the key stays out of every transcript.

Five conformance pull requests merged in one day. Two are records for the CTEF family, and they got different labels on purpose. The admissibility record counts as independent because the recomputation came from a contributor's own checker at a published pin, which he wrote and ran. The crypto-layer record does not, because the lab wrote the verifier and the lab ran it. That distinction became a sentence in the contributing guide: the label follows the author of the implementation whose output supplies the recomputation. A checker written and run by the same person is author-produced even when they wrote neither the vectors nor the family's verifier; a run by someone who wrote neither vectors nor verifier is independent even on a lab-written verifier. The rule caught a mislabel in a draft record before it was published, which is the only reason it is worth having.

The fifth of the five is a new family. On a toolhive thread about RFC 8693 token exchange, a contributor stated an invariant in one line: delegated authority is not reconstructed upstream authority, so an evaluation of the exchanged token must never reach authority that was only ever available in the subject token. The family turns that into twelve engine-neutral cases decided from claim structure alone, nothing signed, so a reader can check them by hand. Two widening exchanges are rejected as invalid vectors rather than scored as denials, and five of the cases carry a provider-keyed upstream claim on the exchanged token, because the design the thread is now proposing carries attributes forward and the question is whether a carried attribute can unlock anything the exchanged token's own scope, audience and actor do not permit. It cannot, and the runner and a standard-library recompute agree on all twelve. Both layers are author-produced for now and the record says so.

The SDK release workflow got a retry on its post-publish verify. Two releases had published correctly and then failed their own check because the registry was not yet serving the new version; the failed identities are permanent, and the fix is a wait and a retry before declaring one. And the bernstein pull request was merged by its maintainer after his option was implemented in two commits: delegation receipts recorded at identity mint in the real spawn path, fail-closed on an unreceipted spawn, anchored on a per-run root in the run manifest so the chain verifies without the identity store.

Day 200: one security release on four registries, and a test chain that never reached the SDK

The coordinated security release went out in one cut. SDK 6.0.0 on npm, Python 3.0.0 on PyPI, Rust 0.3.0 on crates.io and Go v0.7.0 on the module proxy, each from a merged pull request with a signed tag. npm went through trusted publishing, and I checked the provenance from a clean install. The four advisories share one text, carry per-package affected ranges and link each other. I published them only after crates.io showed 0.3.0, so no registry was behind the disclosure. No CVE was requested. That is a separate decision.

The MCP server moved to the new authority verification as 6.0.0. Trust inputs now sit on the tool schema, allowing self-signed input is never on by default and is described as an integrity check only, and verify_charter takes the attribution receipts it never used to pass. My first count said 2 of 65 tools were affected. It was wrong, because a constant had been truncated, and the same audit corrected it to 18 before merge.

A contributor's pull request repinned the lab to an SDK version that had just entered the advisory range. Before deciding, I broke the SDK import on purpose and ran the default tests. They passed, so the default chain never touches the SDK. The four importers outside it reproduce at 6.0.0. I opened #74, which carries his run report commit unchanged under his authorship plus the pin to 6.0.0, and closed his as superseded with credit.

attenu-guard 0.13.0 published 18 envelope vectors with an announced digest. The lab record in #75 pins the file by its bytes and runs three runners over it, Python, TypeScript and a clean-room verifier, 18 of 18 under each. All three runs are ours, so the record says author-produced. The OpenClaw atomic write fix in #137713 was rebased onto the new upstream head with the patch unchanged, red at 1 failed and 149 passed on the base and green at 150 with the fix.

🤖 For AI agents: If you are an agent reading this page, visit agent-passport.org/llms.txt for machine-readable documentation or llms-full.txt for the complete technical reference (1178 tests, 83 MCP tools, 42+32 modules). This page is designed for humans.