52 Commits
Author SHA1 Message Date
Thore Cimbal 937a17ecf8 cve: seventeen decisions pointing at nothing, found the first time the counter could see
The unblinded counter went live with this pull and immediately reported
seventeen. Eleven belong to images the operating-stack deploy replaced — those
decisions did their job and are gone.

The other six are one CVE. Trivy's database moved CVE-2026-57433 in perl-base
from critical to high, and with that a single reclassification made six
decisions across six images obsolete at once. A severity is not a property of
our estate; it belongs to somebody else's database and changes underneath us.
That is now written next to the format description, because the next person will
otherwise wonder why an entry rots without anything here changing.

Fifteen entries and 58 identifiers remain.
2026-08-21 12:00:00 +00:00
Thore Cimbal 9a3b1d1dd5 cve: a successful query can still be worthless, and it cost 42 reports
Restarting k3s took kube-state-metrics and alloy's log tailers down with it, so
kube_pod_container_info went empty in prometheus. The derivation asked, got a
clean response with zero rows, and counted it as success. Cluster targets went
from 39 to none, the total from 54 to 12, and the scan loop deleted every report
whose target had vanished. Coverage then read 1.0.

None of the four rules fired, and each for a defensible reason: the set was not
empty because the operating host still answered, and every timestamp was fresh
because an empty success updates it. The gap sat exactly between them.

A source that has delivered before and now delivers nothing is treated as a
failure: its previous targets are kept, its timestamp ages, and the stale rule
takes over. A fifth rule watches the total for a drop of more than 40 percent,
and it deliberately also fires on a deliberate shrink — losing 40 percent of the
checked estate is worth a line either way.

The six rancher decisions are gone too. They described versions that no longer
run: the k3s patch took all fifteen of their criticals with it.
2026-08-21 12:00:00 +00:00
Thore Cimbal c7aaf388f5 monitoring: the same situation has three different numbers
Grafana 13.2.0 shows 162 findings, 63 series, or 28 vulnerabilities depending on
what is being counted, because the same go stdlib repeats across thirteen plugin
binaries. Prometheus counts series — one per CVE, package and version per target
— so on the dashboard the jump reads 70 to 63, which is fewer, not more.

That makes my original rejection wrong twice over: I compared images with
different contents, and I compared a series count against an occurrence count.
The comment above the version line now carries all three numbers so the next
person does not repeat it, and the README says what a series actually is.
2026-08-21 12:00:00 +00:00
Thore Cimbal d68991629e cve: the stale-decision counter was blind to the case it exists for
I wrote 'and ziel in soll' into that condition this evening, and it means a
decision only counts as stale while its image is still deployed. The moment an
image leaves the inventory — a version bump, which is exactly when decisions go
stale — the entry stops being checked and sits in the file forever. Grafana
12.0.0 and 12.4.9 proved it hours later: both entries survived the jump to
13.2.0 and the counter reported zero.

The cost of removing the guard is that a decision made before the scanner reaches
its image shows up here until the next round. The alert waits 24 hours, which
covers it.

The existing arithmetic test could not have caught this, because it reimplements
the calculation and the faulty condition was never in the copy. New tests drive
the real collect() instead, including one that the old behaviour fails.

Both dead grafana entries removed: 28 decisions down to 26.
2026-08-21 12:00:00 +00:00
Thore Cimbal 3aee13800d monitoring: grafana 13.2.0, loki 3.7.6, alloy v1.18.1
Grafana's core binary is clean at 13.2.0 where 12.4.9 carried the one critical
still on our books. The image also reports far more high findings, and the
comment above the line explains why that is not a reason to go back: 13.x ships
thirteen datasource plugins inside the image, every one of those findings lives
in a plugin binary, and none of the affected datasources is one we use.

Loki drops from 41 high to 8, alloy from 44 to 14. Neither carries state that
has to migrate.

Grafana's does, and it migrates without a way back. sorb accepted that: the
volume holds preferences and history, while both datasource UIDs are pinned in
provisioning and all eleven dashboards come from files — so even losing it
entirely restores to the same UIDs the 657 dashboard references expect.
2026-08-21 12:00:00 +00:00
Thore Cimbal c942ce9328 cve: decide the one critical the new grafana brings with it
12.4.9 is what runs now, and its single remaining critical sits in a bundled go
dependency that grafana has to update, not us. Recording it before the scanner
reaches that image, so the count does not go to one and stay there.

Worth noting why it was not in the earlier list: the old scanner did not know
CVE-2025-41115 in grafana 12.0.0 at all. Trivy 0.74.0 found it on the first
round, on an image we were already replacing — an outdated scanner is outdated
detection, which was the argument for bumping it, now with an example.
2026-08-21 12:00:00 +00:00
Thore Cimbal 2c0caacfb0 cve: the cadvisor decision cited a commit that never touched it
I wrote that v0.55.1 was waiting in 516641b. It is not — cadvisor is not in this
compose file at all; it runs in the portainer stack next door, so nothing here
sets its version. The finding stands and the measurement stands, but the reason
attached to it was false, and a decision is only worth its reason.

Same shape as the portainer agent: measured, outside our deployment path, and
therefore dated with the others rather than with this week's rollout.
2026-08-21 12:00:00 +00:00
Thore Cimbal 1ce4b55b75 monitoring: the hardening lived on the host and not in here
Prometheus and Loki accept writes without authentication and were bound to
0.0.0.0; node-exporter likewise. That was fixed on CFGMON during the firewall
work today and never came back to the repository, so this file still described
three open ports and a firewall as the only thing in front of them.

The consequence is worse than a stale comment. A pull would have reverted the
binding and reopened all three, and nothing here or there would have said so.
It did not happen only because git refused the pull over the local edit — the
accident that saved it is not a control.

Ports now match what actually runs: the private vSwitch address for the hosts
that push, localhost for the host itself, and node-exporter on localhost alone.
2026-08-21 12:00:00 +00:00
Thore Cimbal 422b651f43 cve: synapse carries six of its own, and they were nearly missed
The pipeline still holds the report for v1.151, so the six criticals in v1.158
never showed up in any query against it. They surfaced only from re-deriving the
target set against the live cluster and folding in the pre-deploy scans by hand.
Five sit without a fix in the debian base, one in the bundled tool's go runtime;
the chart sets the version, so there is nothing to take.

With that entry the count over all 54 desired targets is zero open.
2026-08-21 12:00:00 +00:00
Thore Cimbal d5e2995314 cve: a digest-pinned container was about to fall out of the target set
The alloy chart's 1.x line pins its config-reloader sidecar by digest. Kubernetes
then reports that container's image as a bare sha256 and puts the usable
reference in image_spec alone, which the derivation was not reading. An hour
after the chart bump the sidecar would have gone unscanned — the exact hole
#0106 exists to close, reopened by an upgrade rather than by neglect.

Worse than missing: the bare digest passed normalisation as repository 'sha256'
with the hex as its tag, so it would have entered targets.txt, failed every pull,
and shown up as a permanent coverage gap pointing at nothing.

Both ends are closed and both are asserted, including that the query still asks
for image_spec — an assertion on the parsing alone would stay green while the
data never arrives.

Ten more decisions cover what is fixed in the repo but not yet rolled out on the
operating host, dated a week out so the alert speaks up if the deploy does not
happen.
2026-08-21 12:00:00 +00:00
Thore Cimbal 44cc030f0b cve: a critical with nothing left to do is allowed, but only on the record
Sixty-three criticals on running images have no fix to take. Writing them into
a trivy ignore file would have been the obvious move and the wrong one: trivy
drops ignored findings from its output, so afterwards 'zero because fixed' and
'zero because we looked away' render identically. Every finding stays in
trivy_vuln_info. The decisions sit beside them in entscheidungen.json and are
counted, not subtracted.

Each entry names its CVEs one by one. A blanket entry per image would also
swallow the next finding that shows up there, which is the finding you would
most want to see. The loader rejects an entry without ids, and rejects a review
date it cannot parse — rejecting the whole file, because a half-read decision
list is worse than none.

An unreadable file leaves everything counted as open. Getting that direction
backwards would mean a typo reads as 'all decided', and nobody would notice.
The end-to-end run caught the same mistake in the other half: when the derived
target set is empty the set is unknown, not empty, so the open count now falls
back to every report rather than to zero.

Three python services move off the debian base while we are here — 3.13-slim
carried four criticals with no fix, 3.13-alpine none. All three run on the
stdlib alone and TLS was checked inside the image before the switch.
2026-08-21 12:00:00 +00:00
Thore Cimbal 516641bbba monitoring: five image versions, measured before choosing (#0051)
The operating stack carried 30 critical and 676 high findings. These five lines
remove fourteen and 278 of them. Every target was scanned before it was written
into the file:

  trivy          0.58.2  -> 0.74.0    critical 4 -> 0, high 98 -> 0
  grafana        12.0.0  -> 12.4.9    critical 7 -> 1, high 70 -> 3
  prometheus     v3.3.1  -> v3.14.0   critical 2 -> 0, high 46 -> 2
  alertmanager   v0.28.1 -> v0.34.0   critical 1 -> 0, high 41 -> 2
  node-exporter  v1.9.1  -> v1.12.1   critical 1 -> 0, high 38 -> 8

Three candidates were measured and rejected, which is the point of measuring.
Grafana 13.2.0 clears all seven critical findings but takes high from 70 to 162,
so the minor jump inside 12.x beats the major one by a wide margin. cadvisor only
goes from five critical to four. And python:3.13-slim is unchanged — the host
already holds the current build, so a repull buys nothing.

Configuration was validated against the new tools rather than the old ones:
promtool v3.14.0 accepts prometheus.yml with both rule files, all 20 plus 5
rules, and the rule unit tests; amtool v0.34.0 accepts alertmanager.yml.

⚠️ Needs a deploy on the operating host. Recreate rather than up: the images
change, and the CVE targets themselves move with trivy and the exporter.
2026-08-21 12:00:00 +00:00
Thore Cimbal 59c75c8a01 cve: actually test the narrowing — the previous commit shipped it uncovered
Disabling the filter left all twenty-two assertions green, which means nothing
tested it. The counter-case I had written exercises the skip in the derivation,
not the filter inside the registry selection, so the feature went out with no
coverage at all — the project's own favourite failure, in the change that was
meant to remove noise.

Three assertions now cover it: three repositories of which one runs, the
counter-case showing the same call returns all three without estate knowledge,
and one that pins what the rule must not do — it excludes repositories, not old
tags, so a rollback target of a running service still counts.

Same sabotage as before now turns one of them red.
2026-08-21 12:00:00 +00:00
Thore Cimbal 4758e35150 cve: only scan registry repositories that are actually running (#0051)
Decision by sorb: narrow the target set to what is operated rather than tidying
the registry. A repository now counts only while at least one of its tags is
running, which drops element-desktop-build and windows-vm and 61 of their 62
critical findings with them. Of those 62, exactly two had a fix available; they
are build artefacts nobody runs, so remediation was never the right answer.

The narrowing stays derived rather than maintained: the running set is the one
the derivation already builds, so there is no second list to keep in step
(ADR-0026).

Order now matters. The registry selection needs the running estate to tell a
rollback target from a build artefact, so it runs after the other two sources
and is skipped when they yield nothing.

That last part changes behaviour deliberately. A derivation where both estate
sources answer successfully but empty used to count as complete and merely
unusable; it now reports the registry as failed. If nothing is running at all,
that is an outage rather than a normal state, and it should say so instead of
hanging on a single boolean. The test carries the new contract with that
reasoning written next to it, plus a counter-case proving the registry does run
when the estate is known.
2026-08-21 12:00:00 +00:00
Thore Cimbal c5b20a1e47 docs: the README told the reader to maintain a file that no longer exists
Three passages still described cve/images.txt, and one of them instructed the
next person to keep it in step with the stack on every change — the exact habit
ADR-0026 abolished, and the one that had left coverage at 52 percent.

The section now says there is nothing to keep in step, names the three sources
the set is derived from, and carries the warning that matters: a failed
derivation shrinks the desired set, which makes coverage look better rather than
worse. Whoever edits this must not trim the freshness stamps away.

The compose comment said the same thing and is corrected with it.
2026-08-21 12:00:00 +00:00
Thore Cimbal 07875ba6bb cve: a target that leaves the set loses its report (#0106)
Without this, an image dropped from the desired set keeps reporting: the
exporter reads every json in the results directory and takes the target from
Trivy's own ArtifactName. That is precisely what the security room showed on
2026-08-20, when it carried HIGH findings for threadnet-web:v0.3.0, an image
that runs nowhere.

Deletion only ever runs against a list that was successfully read. The guard at
the top of the round already skips everything when the list is missing or
empty, so a restart during a Prometheus outage cannot clear the estate.

The round became a function so the test can load the real one. The first version
of that test rebuilt the loop instead, and a rebuilt test proves the rebuild —
it stayed green while the shipped file set its paths unconditionally and ignored
the environment entirely. Loading it exposed that within one run.

Two of my own errors are fixed here as well. The driver read
`runde || sleep A && sleep B`, which groups left to right, so a missing list
would have slept the wait AND the full day — exactly what the short wait exists
to prevent. And the paths were hardcoded where the exporter already took them
from the environment.

Removing the guard as a deliberate sabotage turns the dangerous case red:
untouched reports drop from two to zero.
2026-08-21 12:00:00 +00:00
Thore Cimbal 386b65a9e0 cve: the scanner reads the derived set, images.txt is gone (#0106)
Criterion one of this piece of work was that the hand-maintained list stops
existing and is not replaced by another maintained file. It is deleted here; the
scanner reads what the exporter derives.

Without a usable list the loop does nothing at all and retries in a minute
rather than sleeping out the full day, because after a deploy the exporter only
writes the file on its next scrape. The old file staying put during a failed
derivation is deliberate: the loop then keeps working from the last known state
instead of running into nothing.

Probed rather than assumed, with trivy stubbed: no list skips the round, an
empty list skips the round, and a list with a comment header scans exactly its
two entries.
2026-08-21 12:00:00 +00:00
Thore Cimbal b989987d69 cve: four rules and three panels guard the derivation (#0106)
Gate 3 planned three rules; there are four. The fourth covers a case the others
miss entirely: every source answers cleanly but empty. Then nothing is missing,
because the desired set is empty, the freshness stamps are current, and nothing
is scanned at all. The Python suite already carries that case as "an empty set
is not the same as success", so the rule belongs with it.

These are the first rule unit tests in this stack. Each rule has a case where it
must fire and one where it must stay silent, because a rule that always fires
cannot be told from a correct one otherwise. Two sabotages confirm the tests
bite: an unreachable threshold on the source-freshness rule makes the expected
alert vanish, and removing the six hour grace period makes the missing-targets
rule fire at five hours where the test demands silence.

The grace period is not padding. A full round over roughly 65 images takes time,
so right after a deploy the gap is real rather than wrong.

The dashboard gains coverage and unscanned-image counters in the two free slots
of the top row, and a source-freshness bar at the bottom, so no existing panel
moves. That bar is the only place where a failed derivation can be told apart
from success.
2026-08-21 12:00:00 +00:00
Thore Cimbal 27770b6ec7 cve: derive the target set and measure coverage — observing only (#0106)
The scan loop is untouched and still reads images.txt, so nothing about this
deployment behaves differently. What changes is that the exporter now knows what
*should* be scanned and can say how much of it is: three sources, none of them
new infrastructure. The cluster and the operating host both already sit in the
same Prometheus, and the registry answers an anonymous token — the same token
dance Trivy performs to pull.

The coverage numbers are the point of this slice. Once deployed they have to
read 24 missing, 2 orphaned and about 0.52, because that is what was counted by
hand on 2026-08-21. A different answer means the derivation is wrong, not the
hand count stale.

Failure handling is the substance rather than an afterthought. A source that
fails costs only its own share; its freshness timestamp keeps ageing instead of
disappearing, because a series that vanishes can never fire a rule — the third
finding of the 2026-08-01 AAR. When every source fails the target file is left
untouched, so a restart during a Prometheus outage cannot clear the estate.

Twenty-one assertions cover it, each paired with its counter-proof: names that
must collapse and names that must not, a source filter that is shown to matter
by removing it, and time-ordered tag selection against the name ordering that
would silently drop v0.10.0. Sabotaging the normalisation turns eleven of them
red, so the suite demonstrably can fail.
2026-08-21 12:00:00 +00:00
Thore CimbalandClaude Opus 5 2eec485bef alerts: the silence alarms now cover the media restore probe too
A second monthly probe exists (restore-drill-media). Its failure was
already covered - BackupJobFailed matches the prefix - but its silence was
not, and silence is the failure this rule set was written against.

RestoreDrillStale now matches the prefix instead of one exact name, and
names the affected probe in the message: with two drills, "the restore
probe has not run" no longer says which. BackupCronJobMissing gains the new
cronjob, because a series that disappears is not an alert in Prometheus,
it is quiet.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
2026-08-21 12:00:00 +00:00
Thore Cimbal 7b25b6e5ae monitoring: die letzten fuenf Einzeldatei-Mounts auf Verzeichnisse (#0083)
Docker haengt Einzeldatei-Mounts am Inode auf; git pull ersetzt die Datei per
Rename, der Container zeigt weiter auf die alte. Fuer prometheus, alertmanager,
cve-scan und grafana war das laengst behoben - fuer loki, alloy und die drei
Python-Skripte nicht.

Bei den Skripten war die Falle sogar schlimmer: Python liest beim Start, ein
Neustart ist ohnehin noetig und wird deshalb fuer ausreichend gehalten. Der
Prozess startet neu, der Mount zeigt weiter auf den alten Inode, der Container
meldet Erfolg und faehrt alten Code.

Geaendert:
  loki           ./loki:/etc/loki           (-config.file zieht mit)
  alloy          ./alloy:/etc/alloy         (Pfad bleibt gleich)
  matrix-alerts  ./alertmanager:/app
  release-watch  ./alertmanager:/app
  cve-exporter   ./cve:/app

Belegt statt angenommen: docker compose config gueltig; alle vier Pfade in
Wegwerf-Containern nachgesehen; und die Gegenprobe - Datei im laufenden
Container geaendert, Pruefsumme wechselt. Genau das ging vorher nicht.

⚠️ Einmalig noetig: docker compose up -d --force-recreate loki alloy
matrix-alerts release-watch cve-exporter. Danach nie wieder.
2026-08-20 12:00:00 +00:00
Thore Cimbal fa26393649 monitoring: Gameserver-Dashboard neu bauen (#0002)
Die bisherige Fassung fragte elf Metriken des loens2-Exporters ab
(pterodactyl_server_cpu_absolute, _memory_mebibyte, _uptime_milliseconds ...).
Keine davon existiert - der Exporter hat nie funktioniert, das Dashboard war
seit jeher leer, und es ist niemandem aufgefallen.

Neu auf dem, was tatsaechlich anliegt: fuenf pterodactyl_server_*-Metriken des
Exporters, der die Panel-Datenbank liest, plus die fuenf gameserver:*-Recording-
Rules fuer CPU, Speicher und Netz.

Zwei Fallen in den Daten, beide beruecksichtigt:

- status, players und max_players tragen KEIN uuid und KEIN game, nur server_id
  und server_name. Die Zusatzangaben in der Tabelle kommen deshalb ueber einen
  Join auf pterodactyl_server_info.
- players liefert NaN, wenn der Server steht. Ohne den Vergleich > -1 zoege ein
  einziger gestoppter Server die Gesamtsumme auf NaN.

Alle elf Queries vor dem Commit gegen den laufenden Prometheus geprueft, nicht
nur geschrieben - genau der Fehler, an dem die alte Fassung starb.
2026-08-20 12:00:00 +00:00
Thore Cimbal a07c6eda09 monitoring: Game-Host-Anbindung nach der Gegen-Uebergabe fertigstellen
Die Host-Seite ist deployt und verifiziert; der Betreiber hat drei Abweichungen
zu ba59518 gemeldet, alle uebernommen:

1. pterodactyl_exporter zeigt auf 9810 statt 9531. Der alte Exporter (loens2)
   hat NIE funktioniert - er rief die Client-API mit einem Application-Key ueber
   http auf und lieferte null pterodactyl_*-Metriken. Ersetzt durch einen, der
   die Panel-MariaDB read-only liest und keinen API-Key braucht.

2. metric_relabel_configs sind hinzugekommen. Sie lebten bisher im lokalen
   Prometheus des Game-Hosts und gehoeren immer dem SCRAPENDEN Prometheus - mit
   dessen Rueckbau also hierher. Ohne sie waeren Gameserver nur unter ihrer UUID
   sichtbar und je Serie rund 20 container_label_* uebrig geblieben; erwartete
   Serienzahl faellt von 3615 auf ~1700.

3. instance = gameserver statt pterodactyl, damit Metrik und Log denselben
   Schluessel tragen (promtail setzt host=gameserver). Kostet keine Historie -
   die Targets sind am 2026-08-20 zum ersten Mal ueberhaupt up gegangen.

Neu: gameserver-rules.yml nimmt den Join Container-UUID -> Klarname vorweg.
cAdvisor kennt Gameserver nur unter ihrer UUID, und Relabeling kann keine
zweite Metrik nachschlagen.

Mit promtool gegen die hier tatsaechlich laufende Version 3.3.1 geprueft (die
Uebergabe nutzte 3.0.0): Konfiguration gueltig, 2 Regeldateien, 5 Regeln.

⚠️ Die Uebergabe nennt --force-recreate wegen der Inode-Falle (#0083). Fuer
Prometheus gilt das hier NICHT mehr: Das Verzeichnis wird gemountet
(./prometheus:/etc/prometheus), genau um diese Falle zu beseitigen. Ein
restart bzw. SIGHUP genuegt.
2026-08-20 12:00:00 +00:00
Thore Cimbal ba59518ef0 monitoring: Game-Host privat scrapen statt ueber die oeffentliche IP (#0002)
Die Targets standen seit dem ersten Scrape auf der oeffentlichen IP und waren
nie up. Ursache jetzt belegt: Die Exporter-Container auf dem Game-Host
veroeffentlichen ihre Ports gar nicht - docker ps zeigt "9100/tcp" ohne
0.0.0.0-Mapping, sie sind nur im Docker-Netz erreichbar. Weder Firewall noch
Bind-Adresse, wie bisher vermutet.

Targets auf 10.0.0.4 umgestellt (Host liegt im vSwitch, 22/80/443 antworten
dort). cadvisor bewusst auf Host-Port 8081: 8080 ist von coolify-proxy belegt.

Neu angebunden: pterodactyl_exporter auf 9531 - laeuft dort seit drei Monaten
und war nie im Monitoring.

README nachgezogen; sie nannte weiterhin die oeffentliche IP und beschrieb die
Targets als "antworten aktuell nicht", ohne den Grund.

⚠️ Wirksam erst, wenn die Ports auf dem Game-Host auf 10.0.0.4 veroeffentlicht
sind. Bis dahin bleiben die Targets down - unveraendert zum Zustand davor, nur
mit richtiger Adresse und dokumentierter Ursache.
2026-08-20 12:00:00 +00:00
Thore Cimbal 7ee42f90e6 monitoring: Dashboard fuer blockierte ClamAV-Inhalte (management #0077)
Sichtbar machen, wie oft ClamAV tatsaechlich etwas blockiert - bisher nur in
Logzeilen zu finden.

Zwei Abweichungen vom Issue-Text, beide gemessen statt angenommen:

1. Das Issue schlaegt {app="synapse-main"} vor. Das Label app existiert in
   dieser Loki nicht - Cluster-Logs liegen unter
   job="loki.source.kubernetes.k8s_logs" mit instance="<ns>/<pod>:<container>".
   Die vorgeschlagene Query haette dauerhaft ein leeres Panel ergeben.

2. Das Issue kennt nur das Synapse-Modul. Seit dem Client-seitigen Scannen gibt
   es eine zweite Quelle (clamav-http-scanner), und die ist die wichtigere: Sie
   greift beim Senden UND beim Empfangen und deckt damit auch verschluesselte
   Raeume ab. Das Synapse-Modul sieht nur unverschluesselte Uploads - deshalb
   steht es dort auf null, waehrend der Client sieben Treffer zeigt.

Drittes Panel zeigt Fail-open-Faelle: Ist clamd nicht erreichbar, laesst der
Scanner Dateien bewusst durch. Ohne dieses Panel bliebe genau das unsichtbar -
rot ab dem ersten Fall.

Alle acht Queries gegen die laufende Loki geprueft, nicht nur geschrieben.
2026-08-20 12:00:00 +00:00
Thore CimbalandClaude Opus 5 9a106151f1 monitoring: a read error must not erase the CVE timeline (management #0082)
The exporter pruned first_seen on every scrape, keeping only what it had just
seen. A report that failed to parse - a file being written, a brief I/O error -
was skipped by a silent `continue`, so its findings never entered seen_keys and
their first-seen timestamps were deleted for good. Nothing reported it, and
"first seen" simply restarted at now.

Pruning is now limited to targets whose report was actually read this round.
Proven both ways against a throwaway results directory rather than by reasoning:
make one report unreadable and its entry survives while read_errors counts 1; fix
the other report but drop its finding and that entry is pruned as before. The
distinction is the point - the old behaviour was not too aggressive, it was
indiscriminate.

Two numbers now leave the exporter: trivy_reports_total and
trivy_report_read_errors, with alerts on both. They cover what TrivyScanStale
cannot reach by construction - a target that never produced a report has no series
for time() to compare against, so it stays quiet no matter how long it has been
broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 12:00:00 +00:00
Thore CimbalandClaude Opus 4.8 be6b5f0657 docs: add group-rules pointer (CLAUDE.md) and repo-specific AGENTS.md
Field test F-011 found four of five components carried no pointer file, so a
session landing here had no path to the group rules at all. CLAUDE.md is the
one-line pointer the check looks for; AGENTS.md links the canonical rules in the
management repo (with the Gitea mirror URL for readers outside the lab) and
otherwise carries only what is specific and easy to get wrong here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-15 12:00:00 +00:00
sorbandClaude Opus 4.8 e9c13dcf22 feat(monitoring): scrape alertmanager, alert on failed delivery; fix stale README
Alertmanager was configured as an alerting target but never scraped, so its own
metrics were absent: a silently breaking alert chain could not report itself —
the same blind spot as a missing series, now at the end of the chain. Adds the
operating_alertmanager scrape job and AlertDeliveryFailing on
alertmanager_notifications_failed_total.

The README still claimed alert delivery was deliberately muted via a
room=security null receiver. That route is gone; alertmanager.yml routes
everything to the matrix receiver, so the backup alerts added yesterday do get
delivered. Documentation asserting the opposite is dangerous in both directions,
so it now states the current wiring and keeps the alert-storm history as
background.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-14 12:00:00 +00:00
sorbandClaude Opus 4.8 4cb9bfb2bc fix(monitoring): mount prometheus/alertmanager config dirs, not single files
Deploying e0808ba surfaced the trap in practice: docker pins a single-file bind
mount to the inode, git pull replaces files by rename, so the container kept
serving the old alerts.yml while SIGHUP reported a successful reload — host and
container md5 differed and BackupCronJobMissing simply was not there.

The trap was already documented in this README, in detail, with the correct
command and a verification snippet, and it still bit. A footgun you avoid only by
reading gets stepped on eventually, so remove it structurally: directory mounts
resolve through the path on every access. Config paths are unchanged, so
--config.file keeps working. Remaining single-file mounts (loki, alloy, the
scripts) are named in the README as still needing --force-recreate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-14 12:00:00 +00:00
sorbandClaude Opus 4.8 e0808bad90 fix(alerts): a missing series must not silence the backup alerts
Verification on CFGMON showed RestoreDrillStale could never fire: restore-drill
has no last_schedule_time series until its first scheduled run (a manually
triggered job does not set it), and an expression over a missing series yields
nothing. BackupNotRunning shares the flaw — deleting a CronJob removes the very
series the alert reads, so it goes quiet instead of firing.

Fall back to kube_cronjob_created, but aggregate with max by(namespace, cronjob):
'or' matches including __name__, so a bare fallback would return BOTH series and
the never-updating created timestamp would fire permanently once the window
elapsed. Add BackupCronJobMissing so a vanished CronJob is itself the alert.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-14 12:00:00 +00:00
sorbandClaude Opus 4.8 1bbff5e749 feat(alerts): alert on backup failure, stalled backups and stale restore drill
There was no rule covering backups at all: a failed nightly job would have gone
unnoticed, which is precisely the silent failure management #0030 is about.
BackupJobFailed catches a failed run, BackupNotRunning catches a CronJob that
stopped scheduling, and RestoreDrillStale fires when the monthly drill stops —
an unverified backup is an assumption again, so the absence of the check is
itself worth alerting on.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-14 12:00:00 +00:00
Thore CimbalandClaude Fable 5 32f89afebf BACKLOG.md durch ein README ersetzt
Das Repo hatte auf oberster Ebene kein README - das einzige Frontdoor-Dokument
war ein BACKLOG.md, das seit 2026-07-30 nur noch auf 'sorb/Backlogs' verwies.
Dieses Repo gibt es unter dem Namen nicht mehr (umbenannt zu 'management',
kanonisch auf git.lab statt Gitea), und ein dateibasiertes Backlog widerspricht
ADR-0005: alles Offene ist ein Issue. Der alte Gitea-Link funktioniert nur noch
ueber einen 301-Redirect.

Ein Wegweiser, der auf ein Modell zeigt, das abgeloest wurde, ist schlechter als
keiner - deshalb ersetzt statt geflickt. Das neue README beschreibt, was das
Repo ist, verweist auf monitoring/README.md als Betriebsanleitung und nennt die
einschlaegigen Issues mit ihren aktuellen Nummern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-02 14:46:04 +02:00
Thore CimbalandClaude Fable 5 ff87cb25fb monitoring: CVE-Alarme aggregiert pro Image + Receiver-Robustheit (gitops#51)
Entscheidung sorb 2026-08-01 (Option 1 aus #51): Alarme als
count by (target, severity) statt pro CVE (~58 Serien statt ~1200),
CVE-Details bleiben im Dashboard (trivy_vuln_info unveraendert).
Receiver: inkrementelles save_state nach jedem Alarm, 1s-Sende-Drossel
(Synapse rc_message), recent_resolved-Dedup gegen doppelte Fallback-Haken
bei Batch-Retries, Teilfehler -> 502 liefert nur den Rest nach.
Stumm-Route + Null-Receiver entfernt - Zustellung wieder scharf.
promtool/amtool/py_compile gruen. UNGETESTET bis Deploy (Uebergabe-Issue).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 16:11:25 +02:00
0bd77e28d8 monitoring: CVE-Alarme vorerst stumm, Deploy-Fallstricke dokumentiert (gitops#47)
Die CVE-Pipeline aus #47 ist deployt und sammelt Daten, die Alarm-Zustellung
ist aber bewusst abgeklemmt: room="security" routet auf einen Null-Receiver.

Grund: die Regeln erzeugen eine Alarm-Instanz pro CVE pro Image (bei 14 von 29
gescannten Images bereits 59 CRITICAL / 445 HIGH). group_by legt alle in eine
Gruppe, matrix-alerts.py schickt eine Nachricht pro Alarm -> Schwall. Dazu
steht save_state() hinter der Sende-Schleife: bricht ein Send ab (Synapse
rate-limitet nach ~10 mit 429), wird kein State gespeichert, der Receiver
antwortet 502 und Alertmanager wiederholt die ganze Gruppe -- mit leerer
Deduplizierung. Details und Weg zum Scharfschalten im README.

Ausserdem dokumentiert: Config-Aenderungen an prometheus.yml/alerts.yml/
alertmanager.yml werden von "docker compose up -d" NICHT aktiv. Die Dateien
sind einzeln gemountet, Docker haengt den Mount am Inode, git pull erzeugt
beim Umbenennen einen neuen -- der Container sieht weiter die alte Datei.
Es braucht --force-recreate; ein SIGHUP laedt nur den alten Inhalt erneut.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 16:09:37 +02:00
Thore CimbalandClaude Fable 5 cdfadc0f26 monitoring: Security-Dashboard-Ordner provisioniert (gitops#47)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 13:42:32 +02:00
Thore CimbalandClaude Fable 5 b6007c50fd monitoring: CVE-Pipeline v1 (gitops#47) - Scanner, Exporter, Regeln, Routing, Dashboard
- cve-scan: Trivy-Loop ueber die 29 real deployten Images (Cluster-Inventur
  2026-08-01 + Prod-Web-Image); 24h-Intervall, Fehler einzelner Images
  blockieren nicht
- cve-exporter: Stdlib-Exporter mit first_seen-State (Zeitstrahl), Schema
  trivy_vuln_info/_count/_first_seen/_last_scan gemaess Pflichtfeldern
- 3 Alertregeln (CRITICAL sofort, HIGH mit 24h-Daempfung, Scan-Frische) -
  promtool SUCCESS 9 rules; alle mit room=security
- matrix-alerts: Label-basiertes Raum-Routing (MATRIX_ROOM_<NAME>), Edits
  landen im richtigen Raum via State
- Grafana-Dashboard cve-overview: Severity-Stats, CVE-Tabelle mit
  NVD-Link/Fix-Version/first-seen, Zeitstrahl, Verlauf
UNGETESTET bis zum Deploy auf CFGMON (compose up -d).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 13:42:07 +02:00
Thore CimbalandClaude Fable 5 f45e01219b monitoring: release-watch in eigenen CVE-/Release-Raum (gitops#47), SBOM-Abgrenzung dokumentiert
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 10:10:16 +02:00
Thore CimbalandClaude Fable 5 8c06329e33 monitoring: Release-/Advisory-Watch fuer Element-Upstreams (gitops#22)
Stdlib-Daemon im matrix-alerts-Muster: pollt die GitHub-Release-Atom-Feeds
von synapse/ess-helm/element-web/mas/element-call alle 6h und meldet neue
Eintraege als Notiz in den Alerts-Raum (Security-Verdacht mit 🚨 markiert).
Erstlauf setzt nur den State. UNGETESTET bis zum Deploy auf CFGMON.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 04:51:53 +02:00
6ffab68583 monitoring: State von matrix-alerts persistent machen
Der Receiver (ea33c3b) merkt sich die Firing-Nachricht pro Alarm, um sie beim
Resolved per Edit abzuhaken. Die State-Datei lag aber unter /tmp im Writable
Layer des Containers. Das ueberlebt ein "compose restart", nicht aber ein
"up -d", das den Container neu baut -- also genau jeden Deploy. Danach haetten
alle offenen Alarme ihre Event-Zuordnung verloren und sich ueber den Fallback
als separate Nachricht aufgeloest, statt die urspruengliche abzuhaken.

Jetzt: named volume matrix_alerts_data auf /state, Pfad per MATRIX_STATE_FILE.

Verifiziert: Testalarm eingekippt, State geschrieben, Container per
--force-recreate neu gebaut (Container-ID 7e013a99d892 -> b1f95380dede),
State unveraendert vorhanden.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-01 04:51:08 +02:00
Thore CimbalandClaude Fable 5 ea33c3b7b0 monitoring: Resolved hakt die Firing-Nachricht per Edit ab statt neu zu posten
Ein Alarm = eine Nachricht: Firing-Event-IDs werden je
Alertmanager-Fingerprint gemerkt (State-Datei, ueberlebt Restarts),
Resolved ersetzt die Originalnachricht per m.replace mit
durchgestrichenem Text + Haken. Ohne bekannte Zuordnung Fallback auf
eigenstaendige Resolved-Nachricht. Re-Notifies desselben Alarms
erzeugen keine Doppelnachrichten mehr.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-08-01 03:24:52 +02:00
Thore CimbalandClaude Fable 5 5b09a12287 monitoring: echte Alerts-Raum-ID eingetragen, Bot eingerichtet (gitops#32)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-07-31 23:43:04 +02:00
Thore CimbalandClaude Fable 5 682adbdd54 monitoring: Alert-Zustellung auf eigenen @alerts-Bot + dedizierten Raum umgestellt
Entscheidung 2026-07-31 (gitops#32): eigener Bot statt maintenance-notify
(Token-Trennung MATRIX-Host vs. CFGMON), eigener Alerts-Raum im
Operating-Space statt Wartungsraum. Bot-Account existiert bereits.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PKhFj1S3UdD6xL2fbWPeYj
2026-07-31 23:26:32 +02:00
Thore CimbalandClaude Fable 5 9594ec80ec monitoring: Alerting vorbereitet - Alertmanager, 6 Alert-Regeln, Matrix-Receiver (gitops#32)
Regeln decken die real erlebten Fehlerklassen ab: TargetDown (coturn-Klasse),
KubePodRestartLoop (36k-Restarts-Klasse), OOM-Kills, RAM-/Swap-/Disk-Druck.
Zustellung in den wartung-Raum ueber einen minimalen Stdlib-Webhook-Receiver
(gleiche Machart wie maintenance-notify, gitops Issue #24); Bot-Token kommt
beim Deploy per .env (Vorlage in .env.example).

UNGETESTET/DEPLOY-PENDING: promtool/amtool-Lint auf dem Mac an haengendem
Docker-Hub-Pull gescheitert - vor dem Deploy auf CFGMON ausfuehren (Images
liegen dort bereits):
  docker run --rm -v $PWD/monitoring/prometheus:/cfg:ro --entrypoint promtool prom/prometheus:v3.3.1 check rules /cfg/alerts.yml
  docker run --rm -v $PWD/monitoring/alertmanager:/cfg:ro --entrypoint amtool prom/alertmanager:v0.28.1 check-config /cfg/alertmanager.yml

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 18:30:40 +02:00
Claude b067c93677 BACKLOG.md: auf zentrales Backlogs-Repo verweisen
Die offenen Punkte betreffen inzwischen mehrere Hosts, dieses Repo
beschreibt aber einen Stack auf einem Host. Inhalte sind nach
sorb/Backlogs umgezogen und dort pro Host strukturiert; hier bleibt nur
der Verweis, damit die Liste nicht an zwei Stellen auseinanderlaeuft.

Die Historie der urspruenglichen Eintraege bleibt bis 25bb5dd erhalten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:37:38 +02:00
Claude 25bb5dd44e BACKLOG.md: DNS-Bereinigung bei IONOS aufnehmen
IONOS legt pro Subdomain automatisch einen Mail-Satz (MX, SPF,
DKIM-CNAMEs, autodiscover) und ein www.-Paar an, auch fuer Hosts ohne
Mail. Betrifft selendis und matrix vollstaendig, rohana und game nur
beim www.-Paar.

Ersatzloses Loeschen waere schlechter: ohne SPF gibt es keine Aussage
mehr, und ohne MX weichen Absender per RFC 5321 auf A/AAAA aus -- Mail
an @rohana.axion1337.de landete dann auf Port 25 des Hosts. Richtig ist
Null-MX (RFC 7505) plus SPF -all plus DMARC p=reject.

Konkrete Record-Listen fuer rohana und selendis dokumentiert; der User
setzt diese beiden direkt um. game, matrix, ftp und die Apex-DMARC-
Policy bleiben offen -- bei matrix erst klaeren, ob der Server Mail mit
Absender @matrix.axion1337.de verschickt.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:25:13 +02:00
Claude adb8cf3a11 BACKLOG.md: IPv6 und www.rohana ergaenzen
rohana, www.rohana und selendis haben jetzt AAAA-Records auf
2a01:4f8:c17:93eb::1. Let's Encrypt validiert damit bevorzugt ueber
IPv6, die Hetzner-Firewall braucht fuer 443 also eine Regel mit Quelle
::/0 zusaetzlich zu 0.0.0.0/0 -- sonst scheitert die Erneuerung Ende
September trotz offenem IPv4.

Ausliefern ueber IPv6 funktioniert bereits (rohana und selendis
antworten mit 200, Traefik lauscht auf [::]:443).

www.rohana matcht keinen Traefik-Router und liefert das Traefik-Default-
Cert aus. Bei einer Subdomain ist das www.-Praefix ueberfluessig;
Empfehlung ist Loeschen der beiden Records statt Router + Cert-SAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:17:17 +02:00
Claude 8c7a57dbf1 BACKLOG.md: offene Punkte festhalten
Zwei Punkte, die bewusst nicht sofort erledigt wurden:

1. Cert-Erneuerung ab Ende September 2026 braucht Port 443 aus dem
   offenen Internet (TLS-ALPN-01). Betrifft auch Gitea. Alternative:
   Umstellung auf DNS-01, dann ohne offenen Port.
2. Pterodactyl-Host 157.90.155.206 ist nicht erreichbar (kein Ping,
   Ports 8080/9100 dicht), daher 2 Prometheus-Targets down. Bestand
   schon vor dem Rework.

Zusaetzlich notiert: oeffentlich erreichbarer Remote-Write-Receiver auf
9090 ohne Auth, und dass die Grafana-Admin-Credentials aus .env nicht
fuer die HTTP-API gelten.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:11:38 +02:00
Claude edac97e931 monitoring: Alloy-Storage persistieren
--storage.path=/var/lib/alloy/data war gesetzt, aber ohne Volume: die
Positions-Datei lag im Container-Layer und war bei jedem Recreate weg.
Alloy las danach alle Docker-Logdateien von vorn, worauf Loki alle
Eintraege aelter als 7 Tage mit HTTP 400 abwies
("timestamp too old", reject_old_samples). Sichtbar als Fehler-Burst bei
jedem Deploy; betroffen waren nur Alt-Logzeilen bis zurueck zu 2025,
keine aktuellen Daten.

Verifiziert: Positions-Datei liegt jetzt in monitoring_alloy_data und
ueberlebt --force-recreate (28 -> 31 Zeilen), zweiter Recreate erzeugt
0 Fehler. 7 Container liefern weiterhin Logs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:10:04 +02:00
Claude a400f8a4ba monitoring: Grafana-Certresolver auf letsencrypt korrigieren
Das Label nannte den Resolver "le", Traefik kennt ihn aber als
"letsencrypt" (--certificatesresolvers.letsencrypt.acme.*). Traefik
protokollierte daher "Router uses a nonexistent certificate resolver"
und lieferte fuer selendis.axion1337.de sein Default-Self-Signed-Cert
aus. Der Fehler stammt aus dem Altbestand in /opt/monitoring und wurde
bei der Bestandsaufnahme unveraendert uebernommen.

Verifiziert: Let's Encrypt-Cert (YR2) ausgestellt, gueltig bis
2026-10-28, https://selendis.axion1337.de/login antwortet mit 200.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 14:00:58 +02:00
sorbandClaude Opus 5 9deb205dce Merge branch 'rework/monitoring'
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-30 12:13:01 +02:00