Files
management/docs
Thore CimbalandClaude Opus 5 f466e87f7a docs: the media probe runs monthly, and its failure is provably visible
restore-drill-media, on the 4th at 05:20, an hour after the database
probe. The name is the alerting: BackupJobFailed already matches
restore-drill.*, so a failure needs no new rule and reaches the
maintenance room through the single alertmanager route.

Every link is measured rather than assumed. A control run with a
deliberately damaged file reported one mismatch and failed the job. The
job name was checked against the rule's regex. The route was read from the
config. And delivery works today: seven notifications sent, none failed,
alertmanager up.

The probe re-proves its own comparison on every run: after passing, it
alters one shared file by a byte and fails with "this probe proves
nothing" if the comparison stays quiet. That a comparison has only ever
said "equal" is a guess, and running monthly does not change it.

Silence is covered too - the stale and missing alarms now match the prefix
and name which probe is affected - but those two rule changes sit in
threadnet-operating and are not live: the operating stack does not pull by
itself, and Prometheus still serves the old expressions, measured through
its rules API. The failure alarm is unaffected and already armed. Written
down as open rather than reported as done.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
2026-08-21 12:00:00 +00:00
..
2026-08-21 12:00:00 +00:00