Restarting k3s took kube-state-metrics and alloy's log tailers down with it, so
kube_pod_container_info went empty in prometheus. The derivation asked, got a
clean response with zero rows, and counted it as success. Cluster targets went
from 39 to none, the total from 54 to 12, and the scan loop deleted every report
whose target had vanished. Coverage then read 1.0.
None of the four rules fired, and each for a defensible reason: the set was not
empty because the operating host still answered, and every timestamp was fresh
because an empty success updates it. The gap sat exactly between them.
A source that has delivered before and now delivers nothing is treated as a
failure: its previous targets are kept, its timestamp ages, and the stale rule
takes over. A fifth rule watches the total for a drop of more than 40 percent,
and it deliberately also fires on a deliberate shrink — losing 40 percent of the
checked estate is worth a line either way.
The six rancher decisions are gone too. They described versions that no longer
run: the k3s patch took all fifteen of their criticals with it.
The alloy chart's 1.x line pins its config-reloader sidecar by digest. Kubernetes
then reports that container's image as a bare sha256 and puts the usable
reference in image_spec alone, which the derivation was not reading. An hour
after the chart bump the sidecar would have gone unscanned — the exact hole
#0106 exists to close, reopened by an upgrade rather than by neglect.
Worse than missing: the bare digest passed normalisation as repository 'sha256'
with the hex as its tag, so it would have entered targets.txt, failed every pull,
and shown up as a permanent coverage gap pointing at nothing.
Both ends are closed and both are asserted, including that the query still asks
for image_spec — an assertion on the parsing alone would stay green while the
data never arrives.
Ten more decisions cover what is fixed in the repo but not yet rolled out on the
operating host, dated a week out so the alert speaks up if the deploy does not
happen.
Disabling the filter left all twenty-two assertions green, which means nothing
tested it. The counter-case I had written exercises the skip in the derivation,
not the filter inside the registry selection, so the feature went out with no
coverage at all — the project's own favourite failure, in the change that was
meant to remove noise.
Three assertions now cover it: three repositories of which one runs, the
counter-case showing the same call returns all three without estate knowledge,
and one that pins what the rule must not do — it excludes repositories, not old
tags, so a rollback target of a running service still counts.
Same sabotage as before now turns one of them red.
Decision by sorb: narrow the target set to what is operated rather than tidying
the registry. A repository now counts only while at least one of its tags is
running, which drops element-desktop-build and windows-vm and 61 of their 62
critical findings with them. Of those 62, exactly two had a fix available; they
are build artefacts nobody runs, so remediation was never the right answer.
The narrowing stays derived rather than maintained: the running set is the one
the derivation already builds, so there is no second list to keep in step
(ADR-0026).
Order now matters. The registry selection needs the running estate to tell a
rollback target from a build artefact, so it runs after the other two sources
and is skipped when they yield nothing.
That last part changes behaviour deliberately. A derivation where both estate
sources answer successfully but empty used to count as complete and merely
unusable; it now reports the registry as failed. If nothing is running at all,
that is an outage rather than a normal state, and it should say so instead of
hanging on a single boolean. The test carries the new contract with that
reasoning written next to it, plus a counter-case proving the registry does run
when the estate is known.
The scan loop is untouched and still reads images.txt, so nothing about this
deployment behaves differently. What changes is that the exporter now knows what
*should* be scanned and can say how much of it is: three sources, none of them
new infrastructure. The cluster and the operating host both already sit in the
same Prometheus, and the registry answers an anonymous token — the same token
dance Trivy performs to pull.
The coverage numbers are the point of this slice. Once deployed they have to
read 24 missing, 2 orphaned and about 0.52, because that is what was counted by
hand on 2026-08-21. A different answer means the derivation is wrong, not the
hand count stale.
Failure handling is the substance rather than an afterthought. A source that
fails costs only its own share; its freshness timestamp keeps ageing instead of
disappearing, because a series that vanishes can never fire a rule — the third
finding of the 2026-08-01 AAR. When every source fails the target file is left
untouched, so a restart during a Prometheus outage cannot clear the estate.
Twenty-one assertions cover it, each paired with its counter-proof: names that
must collapse and names that must not, a source filter that is shown to matter
by removing it, and time-ordered tag selection against the name ordering that
would silently drop v0.10.0. Sabotaging the normalisation turns eleven of them
red, so the suite demonstrably can fail.