The grafana deploy went through sorb's hands and never got its AAR. Writing it
now surfaced the actual mistake, which sits in gate 1 rather than in the rollout:
I wrote two acceptance criteria that need an authenticated grafana, and I have
none. That was as true when I wrote them as when I came to check them; I simply
never asked. Both are still unchecked today.
The decision to keep decided findings visible instead of feeding them to trivy's
ignore file is binding and lasting, and it existed only in a design document and
a README. It is ADR-0029 now, with the note that it was made and executed before
it was written down — the same fault ADR-0009 admits to about itself.
Both ledgers dated the 22nd and 23rd now carry one line about their commits
saying the 21st, since sorb chose to leave the dates rather than rewrite three
mirrored repositories over it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
A combined 3+4 gate row does not prove an L-sized run held gate 4, and every
ladder row owes its commit once the ledger closes. Both were findings on
yesterday's ledger too, and I wrote this one the same way anyway — which says
the lesson lives in the judge output rather than in how I fill the table.
The jump itself did exactly what gate 2 predicted: 278 high findings down to 89,
fifteen criticals to none, every component on the version the release ships. The
two flags I nearly forgot both earned their place — ingresses still publish the
public address, and the API certificate kept it.
What was expensive was not the version. A restart of the control plane tests
everything that leans on it, and two things did not come back on their own:
kube-state-metrics sat in crashloop, and alloy's log tailers died while logging
'will retry'. Twelve minutes without cluster logs, and nothing said so.
That same window emptied the metric the scan target set is derived from. The
query succeeded with zero rows, which counted as success, so the set fell from
54 to 12 and the next round deleted 42 reports — while coverage read 1.0. Three
rules I wrote the day before all stayed quiet, each for a defensible reason. The
gap sat exactly between them.
The stolperstein page for that failure class is marked harvested, and it is not:
the rule the framework took from it does not cover either case here. Both are
recorded there with what a future formulation would have to say instead.
Setting --node-ip alone would have been wrong in two ways I only found by
looking. servicelb publishes the node's address, so the traefik service and
every single ingress currently carry the public IP and would have started
carrying 10.0.0.2. And the API certificate's SAN list has no 10.0.0.2 in it, so
regenerating it while the kubeconfig points at the public address is how you
lock yourself out. The line now carries --node-external-ip and --tls-san
alongside.
Two risks I had assumed turned out not to exist: alloy scrapes pods only, never
the kubelet or the node role, and cert-manager solves HTTP01 over the public
name. Neither cares what the node advertises.
The install step runs with SKIP_START on purpose. The whole risk of this
undertaking is what the script writes into that unit, so it gets read before
anything starts rather than trusted.
ExecStart says 'k3s server server --disable=traefik --disable=servicelb
--node-ip=10.0.0.2'. The word server appears twice, so everything after the
second one is a positional argument and none of the flags is read. The node
annotation records all three; the node's internal IP is the public address
anyway; traefik.yaml sits deployed in the manifests directory. Four months of
a configuration that describes a cluster we do not have.
It matters here because the install script rewrites that unit. Whatever it
writes will not be the same broken line, and if the flags become live the next
start removes traefik and servicelb — the entire ingress — and moves the node's
advertised address onto the vSwitch.
Everything else in this gate came out smaller than gate 1 assumed. The traefik
chart major renders identically down to the container arguments, the only
difference being a docker.io prefix on the image. The restart costs about
seventy seconds, read off the container start times of the last one rather than
guessed.
And the yield is 29 high findings smaller than I claimed, because I had scanned
the newest tags instead of the ones this release ships. The acceptance criterion
is corrected down to −189 before the work, not after it.
I filed the fifteen criticals in k3s's bundled components as accepted, reasoning
that only a k3s upgrade could move them and that this was its own undertaking.
Then I repeated that in conversation as 'the biggest lever, a big project'
without ever checking which k3s release ships which components. It is v1.34.6 to
v1.34.10 — a patch on the minor we already run.
That patch carries traefik 3.7.8, coredns v1.14.6, metrics-server v0.9.0,
local-path v0.0.36 and helm-controller v0.16.26. Measured against what runs now:
278 high findings become 60, and all fifteen criticals go away, without touching
a single version line in our own manifests.
The risk is not k3s. It is that the same patch moves traefik's chart across a
major, and that a single node with a sqlite datastore means restarting k3s takes
the platform down with it. Both are named as decisions for sorb rather than
assumptions of mine.
Same rule as yesterday's ledger: an empty commit cell reads as pending, and
closing the run is what makes it due. All four are the research that produced
gate 1, so they carry its commit.
Fifty-four of fifty-four targets scanned, nothing missing, nothing orphaned.
Since the target set became derived rather than maintained, that number has
never been whole before.
Two of the six criteria I could not check. Grafana answers me with 401, so what
I have is that provisioning ran, read out of grafana's own log in loki —
provisioned is not rendered, and the two dashboards on ancient schema versions
only fail when somebody opens them. That is written down as a gap rather than
reported as eleven of eleven.
The counter I built the same evening turned out blind to its own case, and the
test I wrote for it could not see that, because it reimplements the arithmetic
and the faulty line never existed in the copy. Second time in two days for that
failure class, in the same project.
Also worth keeping: my filter for loki errors returned fifteen hits, thirteen of
which were loki logging my own query back at me.
Yesterday I rejected grafana 13.x on a number I had misread. Seventy against a
hundred and sixty-two high findings is not a comparison when the two images do
not contain the same things: 12.4.9 bundles no plugins at all, 13.2.0 bundles
thirteen, and every one of those findings sits in a plugin binary. The 13.2.0
core is clean where 12.4.9 carries the one critical still on our books.
Size M, so no design document — the schema reserves those for L, and raising the
class to reach a nicer template is the same trick as redefining a criterion to
pass it. Gate one lives in the ledger instead.
The only real danger here is grafana's database migration, which has no way back.
That is why a verified backup is an acceptance criterion rather than a step.
Closing a ledger changes what it owes. Six ladder rows had empty commit cells,
which reads as pending, and a gate row that recorded 3 and 4 together stopped an
L-sized run from proving it held gate 4 at all — sorb approved them in one go,
but one row per gate is what the size class is checked against.
The gate 5 cell named a hash that an amend had already invalidated. Writing a
commit into the file that records it is circular by nature; the honest form is
to point at the commit that carries the work and let this one carry the pointer.
Zero criticals without a decision across all fifty-four targets, every one of
them actually measured. What could not be fixed is decided instead: twenty-seven
entries, ninety-two CVE ids, each with a reason and a review date that a rule
watches.
Criterion 2 ends at one rather than zero, because the image that runs nowhere is
the rollback target for the one that does, and prefering it unscanned to keep a
number clean would be the wrong trade. Criterion 3 lags until the next scan
round, since the scanner runs daily and today is the day everything moved. Both
are written down as misses rather than redefined into hits.
The closing numbers are hand-computed against the live inventory, not read off
the dashboard. That is how the six criticals in the new synapse surfaced at all —
the pipeline still holds the report for the version we replaced this evening.
A runbook goes with it, because the expensive parts of this pass were never
trivy: a delivery path that does not report its own halt, a network rule a chart
jump outran, and three indentations I assumed instead of measuring.
Only two of the eight could prove anything new, and both did. Our call fork
against LiveKit v1.12.0 instead of v1.10.0 was undocumented anywhere; ten call
member events from two accounts settle it. And the ClamAV module under v1.158 did
not merely import — it rejected an EICAR file.
The real gain is not a CVE. Synapse v1.158 carries the fix for the defect that
caused this morning's incident, so the cause is gone rather than the symptom.
An AAR travels with it, because the login broke for twelve minutes and the cause
was mine. ESS 26.8.0 moves the auth service's Synapse endpoint from the direct
service to haproxy, and the ingress rule I wrote this morning listed three
sources, none of them MAS. Both renderings were in front of me; I had compared
them for our own adaptations and never asked which internal targets had moved.
Five version lines remove fourteen critical and 278 high findings from the
operating host. Three more were measured and dropped, which is the whole reason
for measuring: Grafana 13.2.0 clears every critical finding and nearly triples
high, so the minor jump inside 12.x wins by a distance; cadvisor moves five to
four; and python:3.13-slim does not move at all because the host already holds
the current build of that floating tag.
That last one is the second time today the assumption 'floating tag means stale'
was wrong. With Wiki.js it was wrong the other way round — the floating tag was
newer than the release tags that looked newer by number.
Checked against the new tools rather than the old: promtool v3.14.0 takes the
config and both rule files and passes the unit tests, amtool v0.34.0 takes the
alertmanager config.
The controllers that deploy everything else carried fourteen critical and 231
high findings. All four go to zero critical and high drops to 48, measured on
the target images first as usual.
The release manifest could not be used as it stands: it carries seven
deployments where we run four, so dropping it in would have added three
components nobody here operates. Generated for our set instead and checked
against what it replaces — same four deployments, four service accounts, eleven
CRDs.
For the second time today an intermediate state looked like damage and was none.
Right after the restart two HelmReleases read Ready=False with 'HelmChart does
not have an artifact' while their workloads kept running; the source controller
was rebuilding its artifacts. Rolling back on first sight would have aborted
something that was working.
Asked whether the work was documented per the framework, the honest answer was
no. The judge reports 26 findings across the three ledgers written in the later
part of today — and it had never been run on any of them. The four from the
earlier part are clean, so this is a regression in form during the session, not
a misunderstanding of it.
Two systematic errors. Gates were listed newest first, while the template says
one row per gate as it closes and the judge checks ascending order. And ladder
rows carrying a reused outcome were left with an empty commit column, which a
closed ledger counts as pending — the rung's commit is the one its result landed
in, and now says so.
The third was mine alone: a gate 5 row for a gate that has not closed. A row
without a commit is a claim without evidence, which is the one thing a ledger
exists to prevent. It is gone; what is still owed stands in the notes instead.
All seven ledgers now pass at zero findings.
The framework asks for a run ledger per undertaking, and this one was missing
while Authentik and cert-manager were already deployed. It exists now, and it
says so at the top: reconstructed from commits rather than kept as the work went,
which is exactly the gap a ladder trace is meant to close. Same day the AAR
obligation also only got met when asked.
The ladder rows carry what actually saved work today: scanning a target image
before upgrading, following two documented upgrade sequences instead of jumping,
and checking five potentially breaking changes against our own state — all five
turned out not to apply. Only one thing was built rather than reused, the
temporary disarming of Flux's automatic rollback.
The notes carry the failures plainly: an untested first commit of the narrowing
caught by its own sabotage check, three of my own numbers corrected, and a
boundary in Gate 3 that was never enforceable because Postgres hangs in the same
chart.
An AAR is not optional. AGENTS.md line 202 demands one after every deploy with a
handover and after every incident, and today held both: the account without a
profile row, and this deploy on the operating host. Neither existed until sorb
asked. I had reasoned the first one away using a comment in the template rather
than the rule itself.
That is FB-01 again, one day after it was partly fixed by writing the rule into
AGENTS.md — and in a sharper form. The knowledge gap was closed this time. The
rule stood where it belongs and was findable. It lost anyway, because a
secondary text offered a more convenient reading and nothing contradicted it.
The stolperstein records that distinction, because it changes what a fix would
have to do.
The incident AAR carries the more useful content: four causes claimed in a row,
each disproven by the next measurement, two of them stopped by sorb's objection
rather than by mine. The real cause was documented ten days earlier, and its
diagnostic shortcut had aged — a count of zero only catches accounts that never
placed a call, and this one had twenty-one, all old.
The previous commit dropped it into the slices table with a column missing.
Slice 1 is also corrected rather than quietly upgraded: its planned acceptance —
24 missing, 2 orphaned, coverage near 0.52 — was no longer observable once every
slice went live together and the first round ran straight through. What stands
in its place is the named blind spots, each now carrying a report.
Deployed by sorb, then measured rather than assumed: 56 targets, coverage at a
hundred percent, no orphaned report, and the round completing in about two
minutes. Every blind spot named in the issue now carries a report and both stale
entries are gone.
The news is not that the pipeline was broken. It answered reliably for the
twenty-nine images it knew about. The other twenty-seven had never been asked,
and they carry a hundred and twenty-seven additional critical findings —
including the public wiki, both Traefik layers, Prometheus, Grafana, Alertmanager
and the scanner's own image. That belongs to #0051, which is updated with the
new baseline; one of its candidates is not fixable but removable, since an
unused registry tag only entered the set because its repository has fewer tags
than the selection depth.
Three of my own numbers were wrong and are corrected in place: sixty-five
targets became fifty-six, up to sixty-five messages became thirty-nine, and the
normalisation changed every alert fingerprint so all known findings reported
once more. The last one was foreseeable from Gate 2, where normalisation is
already listed as mandatory; I did not follow it through to the alerting layer.
The shakiest call of Gate 3 measured better than feared. cAdvisor sees all eight
images the compose file declares and four more that appear in no compose file of
ours — the registry itself among them. Deriving from the running host beats
deriving from its description.
A new stolperstein carries the cheapest lesson of the day: a test that rebuilds
the unit under test proves the rebuild. Loading the real function exposed within
one run that the shipped file ignored its environment entirely.
The four rules, the derived target set and the deletion of orphaned reports are
built and checked as far as they can be without the host. What remains is a
deploy, and the first three items of its checklist are numbers already known by
hand: 24 missing, 2 orphaned, coverage near 0.52. If those disagree the
derivation is wrong.
Slice 4 records the more useful failure. Its test first rebuilt the loop instead
of loading it, which would have stayed green while the shipped file was broken.
Loading the real function exposed within a single run that the script set its
paths unconditionally and ignored the environment.
The slice changes no behaviour: the scanner still reads the old list. What it
adds is the ability to say how much of the estate is covered, and its acceptance
is a number already known by hand.
Both mistakes in this slice were caught by controls rather than by luck. The
counter-proof about tag ordering was itself wrong — name ordering loses v0.10.0,
not v0.8.0 — and the test failed until the reasoning was fixed. The first edit
to the compose file assumed the wrong indentation, and the assertion in front of
it stopped a half-applied change from being written.
Slice 1 changes no behaviour at all: the derivation runs, the metrics appear,
and the scanner keeps reading the old list. Its acceptance is a number already
known by hand — 24 missing, 2 orphaned, coverage near 0.52. If the derivation
disagrees, the derivation is wrong, not the hand count stale.
Rolling out is a handover, not a step of mine. Port 2248 on the operating host
answers from the matrix node and agent forwarding carries, but the key is not
authorised there, so the procedure written after this very stack applies:
whoever builds hands over, whoever deploys verifies and writes the AAR. The
quantities that procedure demands are in the doc: about 65 targets instead of
29, about 90 registry requests per derivation, and up to 65 possible messages a
round where yesterday brought 12.
Slice 2 sets the missing-targets rule above the current value on purpose and
only pulls it to zero in slice 3, so the room never learns to live with a red
alert. Slice 5 is where this can still fail quietly: if cAdvisor only shows what
runs, a stopped service is absent from the desired set and coverage still reads
100 percent.
Gitea's package API wants a token, so the tag timestamps come from the registry
itself: manifest, then config blob, then the created field, anonymously, about
ninety requests a round. The constraint of no new credentials survives.
The derivation is cached rather than run per scrape; at a fifteen second scrape
interval it would otherwise make some twenty-one thousand registry requests a
day.
Two assertions carry the design. A source that fails must leave its own share
empty while the others keep delivering, and when every source fails the target
file is not overwritten at all. Deleting reports follows the same rule: no
target file, no deletion, or a restart during a Prometheus outage would clear
the whole estate.
The shakiest call is written down as such: cAdvisor only sees containers that
run, so a service that happens to be down is missing from the desired set and
coverage still reads a hundred percent. That is the hole ADR-0026 warns about,
and Gate 4 has to check it against docker compose config before criterion 2
counts as met.
The set of targets that should be scanned is built where the set that was
scanned is already known: in the existing exporter. Anything else needs the
same derivation twice, and two derivations are two truths. That also settles
where the coverage metric comes from.
The AAR of 2026-08-01 decided one thing outright. Its third finding says a
stale-scan rule cannot report an image that never scanned, because no series
exists to hang the expression on. A coverage figure built from existing series
is therefore blind to exactly the gap it is meant to show, so it has to come
from the desired set instead.
Two prices are written down rather than discovered later: a target that leaves
the set must lose its report, or an image no one runs keeps reporting; and more
targets mean more messages, roughly 65 instead of 29 per round, which Gate 5
has to measure since the reporting path itself is out of scope.
ADR-0026 generalises it — derive targets, never maintain them — with the
condition that makes it safe: a derivation that fails looks like full coverage,
not like an outage, so its freshness is itself alerted.
Six countable criteria, the first of which is that images.txt stops existing and
is not replaced by another maintained file. Coverage of the running estate has
to reach 100 percent, measured as a set difference over normalised names, and
stay visible as a metric so the gap cannot return quietly.
The open question is how far "everything from the registry" reaches. Six repos
hold 36 tags; seven are CI artefacts and 25 of the remaining version tags run
nowhere. Scanning all of them triples the workload and puts findings about
v0.1.0 into the security room — the exact noise this work removes. Three
readings are laid out with a recommendation, not a decision.
The ledger records the rungs: nothing needed building for the sources. Both
image sets already live in the same Prometheus the scanner can reach, and the
registry answers an anonymous token. It also records two wrong turns of my own,
including reading a registry 401 as "needs credentials".
The gate rows can only name their commit once it exists, so they follow the
work rather than riding along with it.
The ladder rung about the DNS pointer said the Corefile imports
`custom/*.override`. Measured against the running Corefile it imports both —
`*.override` inside the main block and `*.server` at the end. The choice of
`.server` was not arbitrary: a second `hosts` block in the main block would
collide with `NodeHosts`. The rung now says so.
The last open line of the smoke list is no longer an argument. With all four
waves in force since 09:39, OpenID tokens were issued at 11:49:58, 11:57:18 and
11:57:56 — read from `open_id_tokens`, not from a self-report. Without one no
call starts, so the criterion is met by measurement.
ADR-0025 moves to accepted, the design doc to done/, and #0088 to done with its
closing section. The node-address gap is carried forward for the next harvest:
it bounds every claim these rules make.
A new stolperstein: while closing this piece of work, a user reported hanging
calls and a hanging identity reset. Four causes were claimed in a row and all
four were disproven by the next measurement, two of them after an objection
from sorb. The actual cause was an account without a row in the homeserver's
`profiles` table — documented ten days earlier and unrelated to egress. The
lesson is the order: on a report during a piece of work, measure the difference
between affected and unaffected users first, before auditing your own change.
Six deploy and helper workloads are restricted; the count of workloads
without a narrow rule went from fifteen to nine. The block is proven from
inside the pod - three truly external targets refused, CoreDNS answering as
the positive control - and the pod is distroless, so the probe ran in an
ephemeral container sharing its network namespace and labels. That is the
method the remaining waves need and it is written down rather than
rediscovered.
Rehearsing the rollback corrected a claim from Gate 2. "Flux makes a git
revert a rollback of minutes" holds only with an addition. The revert took
about forty seconds; restoring afterwards did nothing for four minutes,
although git.lab had the commit and the mirror reported finished. Only a
forced refetch of the GitRepository moved it.
So a rollback is a revert plus a nudge to the source. Someone who only
reverts and waits sees nothing for up to a poll interval and may conclude
the rollback itself is broken while it is merely slow. That belongs in the
record before a wave touches something users notice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
The owner pointed out that rohana is still reachable over the private
network and need not be called externally. Measured and correct: 10.0.0.3
answers as Gitea for the public host header and serves a valid Let's
Encrypt certificate for that very name, validated with full verification.
So the private path is not just reachable but TLS-clean, and no service
needs reconfiguring. What is missing is only the internal pointer - the
Corefile already imports the custom override directory, the ConfigMap
simply does not exist yet.
Two of the four class B workloads therefore stop leaving the cluster, and
the Gitea exception recorded in AGENTS.md becomes an internal rule. The
storage box stays external: no private path was named for it and I am not
assuming one.
The price is written into the design rather than skipped. An internal
pointer turns two paths into one - if 10.0.0.3 is down, rohana is
unreachable from the cluster although the public route would work, and the
failure would look like "Gitea is gone" instead of "the private path is
gone".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Two findings from the preparation change the design.
There is no split-horizon DNS: every public name of the platform resolves
inside the cluster to the node's external address. Services that reach each
other by public name therefore leave the cluster and come back through
Traefik.
And traffic to the node's own address is not subject to the policy.
Measured twice, from two workloads restricted since the 20th, with
different tools: the storage box, rohana and 1.1.1.1 are all blocked, the
node address is reachable, and CoreDNS answers as the positive control. The
restriction works - just not against the node itself. That is a property of
the k3s enforcement, not a manifest error, and no NetworkPolicy can close
it, because policies allow rather than forbid.
It cuts both ways. Class A grows and gets safer, because services talking
over public names keep working under restriction. And the ten workloads
restricted yesterday can still reach anything published through Traefik -
the gap between what the manifest promises and what holds belongs written
down rather than smoothed over.
I also walked into the trap my own acceptance criterion warns about. The
first counter-probe used /dev/tcp in a container whose sh does not have it,
so everything came back "blocked", including the reachable node. Same
mistake as on the 20th, in the undertaking whose criterion 4 names it.
Repeated with nc and a positive control; only then was the result worth
anything.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Two constraints measured rather than assumed. k3s has no separate CNI or
policy pod, so enforcement is standard NetworkPolicy: CIDRs only, no DNS
names. And the cluster is single-stack IPv4 - a v6 connection attempt fails
with OSError - which matters more than it looks: had it spoken IPv6, every
IPv4-only rule would have been a gate that looks tidy and holds nothing.
The fifteen workloads fall into three classes: internal only, exactly one
pinnable target, and genuinely broad. The manifests do not yield the
targets - external addresses live in SOPS secrets and in the images - so
class B comes from resolved names instead.
ClamAV is the hard case and is not talked away. Its signature source
resolves to Cloudflare with rotating addresses. Pinning them breaks
silently at the next rotation - signatures age, the service keeps running,
nobody notices - which is the failure class this project built half its
checks against. Allowing Cloudflare's ranges would look like a restriction
and barely be one. So its egress stays, with the reason in the manifest,
and mirroring signatures internally is named as the later option.
Rolled out in three waves by risk rather than one commit, so a failure says
which service stumbled. The counter-proof will use Python from inside a
pod against a target that is demonstrably reachable with broad egress -
on the 20th both a service IP and /dev/tcp in a dash container produced a
false "blocked".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Assigned by sorb. Size L: fifteen workloads, a real decision per service,
and mistakes land in production.
The acceptance criteria are counted in the running cluster rather than read
off the manifests - that distinction is what produced this undertaking in
the first place, since the issue's own remaining list was a week out of
date. Zero workloads without an egress rule, every exception naming where
and why with evidence that the service fails without it, a pre-agreed
smoke list, and a demonstration from inside a pod that a forbidden target
actually fails.
Non-goals name the temptations: no ingress changes, no new model or tool,
and no switching a feature off to save a rule. Synapse needs broad egress
for URL previews and already has the right control in
url_preview_ip_range_blacklist; disabling the feature would be the wrong
lever, as the issue has said since the 20th. flux-system is excluded too -
it pulls from the network and is the thing that rolls this change out, so
cutting its ground is a separate act with its own fallback.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Both issues described work that no longer exists. Without measuring, two
tasks would have entered planning half-done.
0083 is done. No single-file mount remains in the compose file; the running
exporter emits both metrics that only exist since the 19.08. change; and
the timing closes the chain - the mount commit carries 12:00Z while
Prometheus and Alertmanager have been up since 16:34Z and Loki since 17:08Z
the same day, so the redeploy came after the change. A "yes, it was rolled
out" would have claimed the same and proven nothing.
0088 stays, with a title that is no longer wrong. It said "13 ingress
rules, 1 egress", which was the state on 06.08.; measured today there are
ten workloads restricted to internal traffic and fifteen still open, not
the eleven the issue listed.
Measuring found more than it corrected. wikijs itself is unrestricted while
its database is, which nobody intended and nothing recorded. And the
namespaces authentik and monitoring carry only the metadata block, so their
egress is entirely open - that is not in the issue at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Size S owes no gates, but the ladder section is mandatory even when nothing
was built - an absent record and an ignored rule look identical.
"One ledger per session" does not quite hold here. This session's ledger
was closed at Gate 5 and then this task arrived. A second ledger
contradicts the wording; no ledger contradicts the sentence above it, that
every rule owes a trace or it is decoration. I chose the trace and wrote
the contradiction down rather than resolving it quietly: a session can hold
more than one finished run, and the concept currently knows only sessions.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Closing the run turned status to closed, and the judge - which checks a
closed ledger strictly and an open one loosely - reported fifteen findings
where it had reported none. All mine, in three classes: no ladder entry
carried the commit it became effective in, gate 4.4 carried no approval,
and the slices stood as gate rows although the template expects them in the
design document at size L ("size M records its slices").
The slice rows are removed, which touches existing rows that "appended,
never rewritten" protects. They should never have been there; the design
document has carried them with their commits all along.
Worth keeping, because it is the whole argument for the ledger: while the
run was open the record looked complete and was not. Only the act of
declaring it finished made the check strict enough to say otherwise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Five statements this undertaking contradicts, named rather than left
standing. Two of them are in this document's own earlier gates: the Gate-2
mapping had FB-06 as open when the framework had split the item and closed
that half in v0.1.3, and had FB-09 and FB-11 as open when v0.3.0 covers
them partly. Both only surfaced because the states were checked against the
rule text of the tags instead of the issue status.
The third is FB-12's "point 3 is fixed", whose reasoning - the error is
forgetting, not doing it wrong - was falsified by the same task. That page
is revised rather than appended to, which is the practice this project
harvested and the framework shipped in v0.1.3.
Criterion 5 stays partly met. The deterministic half ran against both of
the day's ledgers without findings; the verdict artifact must not be
written here, because its own template holds that a verdict produced in the
working session is void whatever it says.
Moving the document one level deeper broke three relative links, which
validate.py caught before the commit. Same class as everything else today:
the tool saw what the plan did not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
The row was written in the same commit as the work, so it could not name
that commit's sha - the second time in this run. The parallel run shows the
order that works: the work commit first, then a commit for the ledger row.
That is what "rows are appended as gates close" means.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Slice 4.4. judge.py caught gate 4.3 naming no commit - the row was written
before the commit existed and the sha never filled in. Classified as model
failure, correctly: the rule is clear, the template says so, I did not do
it. Filled in, then clean. A fabricated sha is refused and fails closed.
Both ledgers of the day pass the deterministic half, including the parallel
run's, which this session did not write.
The verdict artifact is missing on purpose. Its own template says a verdict
produced in the working session is void whatever it says, and an agent
judging its own run justifies rather than checks. Acceptance criterion 5 is
therefore partly met, recorded as partly rather than ticked.
First rule-coverage measurement: 14 of 23 rules observable, nine not -
among them "every line traces to the request", "a check owes proof it can
fail" and "incidents get an AAR". Decoration until they carry a duty to
leave a trace. That is the claim this whole series started from, now a
number instead of a suspicion.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Slice 4.3. Every finding is now a stolpersteine page carrying its state,
and the mapping is discharged mechanically: twelve findings in the
register's last hand-maintained revision, twelve pages naming their origin,
zero unassigned.
The states were verified against the rule text of the tags, not derived
from issue status - and that turned up two errors in my own Gate-2 mapping.
FB-06 was listed as open; the framework had actually split the combined
item and closed that half as issue 0028 in v0.1.3, whose WORKFLOW.md
carries "Name what this work made false" verbatim. FB-09 and FB-11 were
listed as open too; v0.3.0 covers them partly through the ledger and the
ladder's trace duty, so they are `partly` with the version named.
Two are `declined`: the framework considered them and will not cover them,
so they remain ours. That state exists because writing `open` for a decided
matter is the failure class this undertaking exists to clear.
The texts are not edited. They came out of git at the register's last
revision and moved unchanged; only repo-relative links were rewritten,
because inline links resolve against the file.
Also done: the duplicated half of the ADR rule is gone from the project
section - permanent exceptions have been upstream since v0.1.2 - and three
descriptions that still called the register an inbox now describe the
pages, the generated overview and the signpost.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Slice 4.2. The receipt duty from ADR-0024 was built yesterday on the
premise that "the error is forgetting, not doing it wrong". The same
upgrade falsified that on the same day: the receipt for schema.yaml was
written, it is truthful, it names correctly what was carried over - and two
fields were lost anyway. Nobody forgot anything.
pruefe_upstream_drift.py now holds our schema.yaml structurally against the
baseline: every type, field, enum value and required entry the baseline
carries must exist here. Adding is allowed, losing is not. Checked
counterfactually against the real loss, which it reports.
For the two other extended files there is no equivalent - they are Python,
and no structural comparison exists. That half stays open and is written
down rather than glossed over.
Building it taught the check two things about itself. It broke the existing
selftest because that fixture wrote schema.yaml as prose - the same
unrealistic fixture FB-12 already records as having made an assertion
worthless for this very script. And my counter-control reported "nothing"
for the removed crash guard, because a crash returns the same exit code as
a failure; only checking for a traceback showed the script never reaches
its summary without it. Both times the control was blunt, not the check.
Sixteen assertions, six deliberate breaks, none uncovered.
The occasion is recorded for the next harvest as two stolpersteine pages -
the first use of the mechanism ADR-0009 decided: a receipt certifies
attention rather than completeness, and a file cannot be authoritative and
immutable at once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Slice 1, the tracer bullet: schema, generator and one migrated page, so the
whole chain runs before eleven more depend on it.
Two fields were missing entirely. The v0.3.1 merge did not carry over
wiki-page's status and harvested_in, so the mechanism ADR-0009 decides was
not actually available here. A field-by-field comparison against the
vendored baseline found exactly those two and nothing else - the gap the
previous run predicted when it noted that reconciling the extended files is
a manual step with nothing to contradict it.
The generator gains a clustered section and writes the signpost; collect()
and apply_rules() already existed, so the change is a filter and a fifth
rule rather than a second reader or a new checker.
Running it corrected one of my own design errors immediately. The "without
state" group was written as a warning, and it flagged two perfectly correct
pages: status is optional per ADR-0009 and only meaningful on a page that
tracks a pattern. A permanent complaint with no subject trains people to
ignore the section, so the group is now neutral - with the downside written
into the code, since a pattern that lost its state now looks like an
ordinary page.
Eight controls, seven of them deliberate breaks: both generated files go
stale on a hand edit, an invented state value is refused, the new value is
accepted and clusters correctly, and harvested or declined without a named
version now fails.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
The owner objected to the plan carrying two findings as "open" when the
framework has considered and declined them, on the grounds that the enum
has no better value. He is right: replacing a state the field cannot
express with a state that is wrong is the exact class this undertaking
exists to clear.
Extending is established practice here rather than an exception - the
extension block in schema.yaml already declares two enum extensions, and
ADR-0024 covers the file. The portability cost is near zero because these
pages never travel upstream: ADR-0008 sends generalized failure classes,
which carry no frontmatter.
The value is `declined`, not `rejected`. Our issue enum already uses
`rejected`, and matching vocabulary would be cheaper, but the meanings
diverge: a rejected issue is finished, while a declined pattern persists
and stays ours to live with. Anyone reading `rejected` would think the page
is disposable.
Measured while deciding where to document it: HERKUNFT.md cannot be
maintained at all. It sits under docs/sources/ and the lock check refuses a
test commit against it, while ADR-0024 declares its table authoritative and
requires it to stay in step with the pair list. A file cannot be both the
maintained truth and immutable. Recorded for the next harvest; an accepted
record is not repaired in passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM