Eight commits from the past two days carry the wrong author identity, made in an
agent session that set user.email by hand in fresh clones — an hour after that
same session wrote the canonical identity into AGENTS.md.
sorb's call is to fix them with the next history pass rather than force-pushing
two repos over eight commits. The issue exists anyway because gruppenpruefung
reports them on every run: without a recorded reason the next session starts
'repairing' them, or worse gets used to red findings, which is exactly what
happened with the TargetDown noise in #0002 the same morning.
Notes the structural prevention too — an includeIf block setting the identity for
group clones — since writing the rule down demonstrably did not prevent breaking it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Declaring threadnet-wiki as a component took the check from 23 to 72 findings,
46 of them Wiki.js git-storage commits. Those are the same class ADR-0009 already
exempted for the rotation bot — written without a human present, so attributing
them to a person would be wrong — but the exemption held exactly one name.
Matching is on the address rather than the display name on purpose: the Wiki.js
account shows up as 'Administrator', which is far too generic to silence findings
with. Down to 25, and the checks that should still fire do: eight non-canonical
commits remain, all of them mine from the past two days.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The cluster was running coturn 4.10.0 while :latest pointed at 4.17.2, which is
the concrete harm the issue describes: nobody knew what ran, a reschedule would
have jumped seven minor versions unannounced, and the CVE scan was measuring a
moving target. Now pinned to 4.17.2 and verified beyond 'the pod is up' — a STUN
binding request from the public internet succeeds and the server reports the
caller's external address, so the relay path itself is proven.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The group check flagged four undeclared projects. Two are real components and are
now declared, both deliberately without a mirror: threadnet-wiki because its
content flows the other way (Wiki.js to Gitea, canonized to git.lab — a mirror
back would close the loop and overwrite edits), and notfallhandbuch per ADR-0016.
The other two are cleanup rather than declaration: project 42 'wiki' looks like a
superseded first attempt, dead since 2026-08-12, and 43 is already marked for
deletion. Declaring either would misrepresent them — the phase enum has no state
for 'abandoned'.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
All five components now carry the pointer the check looks for, verified by
gruppenpruefung.py dropping from 27 to 23 findings — exactly the four pointer
findings. Each AGENTS.md carries only what is specific and easy to get wrong
there: for the forks, that the README is upstream material describing something
else entirely; for thread-net-git, that its small compose file hosts the Flux
source; for threadnet-operating, the two lessons this session paid for.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
W4 point 4 closed with gitops 60aaf0e: the lab WireGuard config now has a repo
home. The root CA turned out to already have one (ci/lab-ca-chain.crt is exactly
the aXionLabs chain), so that half of the point was quietly already met.
All eight now carry a named resolution with a reference, two of them as their own
ADRs, honouring this issue's own rule of documenting rather than silently fixing.
The only thing left is the rotation, which is dated follow-up work in #0015 rather
than an open contradiction.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
sorb's call: rotate everything once at acceptance rather than piecemeal now.
That is the lower-risk order — the mirror credential is still unidentifiable and
the mirrors feed the Flux source, so four separate revocations would mean four
separate ways to break it silently. The inventory and the ordering stay valid, so
the later rotation is execution rather than analysis.
Recorded what the deferral accepts rather than leaving it implicit: the exposed
WireGuard key and PATs stay valid, five never-used tokens remain (one with
manage_runner and k8s), and 'acceptance' is not a dated milestone — which is
exactly how security work rots. The existing due date stays as a review anchor,
not a rotation deadline.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Inventoried the 24 git.lab PATs by metadata only — last_used_at separates
'needed' from 'lying around': four are in active use, five are active but never
used at all (one with manage_runner and k8s scope), and several names exist twice
because a replacement was created without revoking the old one. All six push
mirrors are healthy, but GitLab masks both parts of the mirror URL, so the
credential remains unidentifiable — and it is a Gitea token, which the PAT list
cannot answer for. Hence the ordering: set a dedicated mirror credential first,
revoke second. The revocations themselves are sorb's; from here a never-used
token is indistinguishable from a staged one.
W5 resolved: the secrets rule now has a bootstrap exception, since on a headless
host it was only satisfiable by violating it. W4 splits — point 5 is #0015 (plus
the WG key, which no token inventory covers), point 4 is demonstrably undone.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Corrects only the split-DNS line of ADR-0004 (frozen once accepted, hence a
separate ADR). Rather than documenting 'four zones exist', it measures what each
one does: ~lab and ~axionlabs.de resolve names that exist only internally or
differently (git.lab, and ca.axionlabs.de as real split-horizon to the step-ca),
~axion1337.de carries the internal-only git.axion1337.de, and ~lab.de carries
nothing at all while routing a foreign public domain through the lab resolver —
so it goes.
This also answers the audit's rollback option: reverting to ~lab alone would have
broken internal CA and git resolution. The purpose was never written down, which
is why rolling back would have been blind.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Queried the lab resolver directly from inside the lab VLAN and compared every
record type against the public view: A, MX, TXT, CNAME, subdomains that exist
only publicly, plus records created and deleted yesterday. Not a single
divergence — the UDM holds no zone of its own and forwards live; the aa flag it
sets is a UniFi quirk and was what made the hypothesis look plausible.
That disproves the risk I asserted earlier in this issue, where I called the
pinned ACME resolvers 'load-bearing'. They are good practice, not a safety net
against W1, and the claim stood as fact for an hour. Corrected in place.
W1 is therefore documentation-only. What remains is that the purpose of the three
extra zones is recorded nowhere, which is why a blind rollback is the worse option.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
W8 is moot: sorb confirms UDM SSH was disabled long ago.
W1 turns out to interact with #0007, which did not exist when the audit was
written. CFGMON routes ~axion1337.de to the lab resolver, and Traefik's DNS-01
renewal verifies TXT propagation — had it used the system resolver, that check
would ask the UDM and might never see the challenge record, failing renewal
silently until the certificates expire. It does not, because the config pins
public resolvers explicitly; that line is load-bearing rather than cosmetic and
is now documented as such. What the UDM actually answers for the zone remains
unverified, with the commands to check it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
W2 was already covered — the ping trap lives in the textbloecke, and the firewall
exception is conditional. W3 fixed: cfgmon.md listed the runner as running though
it was dismantled on 2026-08-01; row removed and, rather than leaving the open
question, the page now states that the service table is current state while the
sections below are history. W6: addendum practice had proven itself twice but was
undefined, so the AAR template now makes it a rule — append-only and dated, so the
original mistake stays readable. W7: ADR-0009 unified the identities but AGENTS.md
only said 'canonical author identity' without naming it; now spelled out.
W1, W4, W5 and W8 need sorb's decision and are written up with what each one
costs if left alone — W8 (root SSH on the gateway) being the sharpest.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The original silences expired on 2026-08-04, so TargetDown had been firing every
4h for eleven days — noise that dulls the very alert path the backup work in
#0030 depends on. New silence is scoped to the two GAME jobs rather than the
alertname alone, so future TargetDowns for anything else still get through, and
it carries an expiry that forces a re-decision if the vSwitch move has not
happened by then. The date is now this issue's de facto deadline, so it also
went into the wartegrund where STATUS surfaces it.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Went through the seven imported waiting issues and replaced the generic
'reason is in the GitLab history' placeholder with the real blocker, which
completes #0041. Three of the seven were not merely imprecise but wrong:
- #0025: the deploy had long landed; screenshots confirm 24 aggregated messages
in the security room (limit 29), summing to the known 126 CRITICALs.
- #0014: the A/B/C decision exists as ADR-0008 (option A). Half its open question
is now answered — MATRIX has no docker group at all, so the root-equivalence
does not apply there.
- #0027: the blocking Struktur-Workshop happened on 2026-08-06 and produced three
ADRs, but W1 and W3 were spot-checked and are still unresolved.
The remaining four wait on a named action by sorb. Measured from here: the GAME
exporters are still filtered (and their silences expired on 2026-08-04, so
TargetDown has been firing every 4h since), while CFGMON's 9090/3100 are already
closed from the internet — so #0008 is about making that state deliberate rather
than an acute exposure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The handover issue still sat in waiting while the deploy had long landed:
ff87cb2 is an ancestor of HEAD (CFGMON now runs e9c13dc), the rules aggregate
per image so the per-CVE flood is structurally impossible, matrix-alerts.py
saves state incrementally inside the send loop, and notifications_failed_total
is 0 across 80 series. The null-receiver kill switch is gone.
Recorded honestly what was not observed: whether aggregated messages actually
arrived in the security room once. Delivery is now permanently monitored via
AlertDeliveryFailing, so a future failure reports itself instead of relying on
someone looking.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Records the deployed state: fallback expressions, directory mounts (which proved
themselves on the very next rollout, where a SIGHUP reload was genuinely enough),
and alertmanager now scraped so delivery failures are visible. Also notes the
README correction — delivery had not been muted since gitops#51, and docs saying
otherwise would have made a missing alert look expected.
Left open deliberately: TrivyScanStale has the same missing-series gap but no
natural equivalent to kube_cronjob_created, so it needs a decision rather than a
reflex; and phase A/B of the restore drill still needs a throwaway environment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Alert rules deployed and verified on CFGMON. Verification surfaced two variants
of the same failure class: a missing time series silences an alert instead of
firing it (fixed with a created-time fallback aggregated via max by, plus an
absent() alert for a vanished CronJob), and a SIGHUP reload that reported success
while serving the old file from a stale inode.
The second one matters most: that trap was already documented in detail, with the
right command and a check, and it still bit — so it was removed structurally
(directory mounts) rather than documented harder.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Monthly restore-drill CronJob (verified before commit) plus the finding that
mattered more: there was no backup alerting at all, so a failed nightly job
would have gone unnoticed. Added BackupJobFailed/BackupNotRunning/
RestoreDrillStale; the last one alerts on the absence of the check itself.
Alert rules still need deploying on CFGMON.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Standing exception to ADR-0001 (everything is push-mirrored to Gitea). The
handbook necessarily maps the infrastructure, the backup locations and where the
keys are kept; mirroring it onto the internet-facing host that is itself one of
the covered failure cases would hand a post-compromise attacker their next step.
Confidentiality over availability, with a local clone closing the availability
gap. Records the rejected alternatives so the question does not reopen.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Not mirroring it is deliberate: the handbook maps the infrastructure, backup
locations and where the keys live, so putting it on the internet-facing Gitea
hands an attacker the roadmap once the stack is compromised. My earlier
recommendation only weighed availability and was wrong. Local clone covers the
availability gap. Flagged that this is a standing exception to ADR-0001 and
would warrant its own ADR.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The databases are no longer an assumption: notfall.sh stage 3 restored all
three Borg repos into a throwaway postgres inside the pod (synapse 31908 rows,
MAS 16085, authentik 325149, wiki 251), isolated from production and repeatable.
Procedure and tool now live in git.lab/axion1337.chat/notfallhandbuch so an
emergency needs one clone; this repo keeps a pointer. Still open: phase A/B on
an empty host, Synapse media, a repeat cadence, and mirroring that new repo off
git.lab.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Derived from the running system: prerequisites (age key from the vault first),
the three Borg repos with their archive layout, bootstrap order (host/K3s, the
two manual secrets, Flux, kustomization dependencies), data restore and
verification. Lives in git rather than Wiki.js on purpose — the wiki runs on the
cluster being restored.
Two previously undocumented findings: Flux pulls from Gitea rather than git.lab,
so a simultaneous loss of rohana requires repointing gotk-sync first; and
consumers must be scaled down before pg_restore or their startup schema collides.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
sorb's call: Gitea on rohana is a push mirror of the canonical git.lab, so a
nightly gitea dump would back up a copy — effort not justified, cron stays off.
Documented the one non-mirror asset for the record: the container registry holds
four images the cluster pulls (incl. threadnet-web and the backup image itself),
which is rebuild time rather than data loss and is covered by #0022/#0033.
For #0030, sorb confirms the age key is also in the password vault, dissolving
the circular dependency found earlier. Noted that the vault is now part of the
restore path and must lead the procedure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Cluster-side backups are healthier than assumed: three nightly Borg jobs to a
Hetzner Storage Box, all completing with plausible volumes and working prune
(synapse 199MB/247 files, authentik ~150MB, wikijs 223kB DB-only). No silent
failures.
Critical finding for #0030: the Borg passphrase and SSH key needed to READ those
backups are SOPS-encrypted under a single age key that exists only in the cluster
being backed up and on one laptop — no documented cold copy. Losing both makes all
three repos permanently unreadable. Cold escrow must precede any restore drill.
For #0010 this shrinks the work: the Storage Box + Borg pattern already exists and
is proven, so Gitea needs only its own repo there. The disabled cron (no backups
since 2026-07-30) remains separately urgent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Verified via DoH: _dmarc.axion1337.de is p=reject (sorb changed it), subdomains
inherit reject with sp= absent per RFC 7489, and the noted DKIM gap was a false
alarm — IONOS uses s1-ionos/s2-ionos/s42582890 selectors, all present with valid
keys. Apex SPF left at ~all deliberately: real mail flows over the apex and DMARC
already enforces reject, so -all adds little while risking silent send breakage.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ftp.axion1337.de deleted and verified (NXDOMAIN). All four DNS hygiene issues
from today's batch (#0001, #0003, #0005, plus #0007 earlier) are now closed:
rohana and selendis hardened with Null-MX/SPF -all/DMARC reject, matrix and
www.game removed, ftp removed. Production A/AAAA records untouched throughout.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
matrix.axion1337.de was already removed by sorb independently (platform runs
under .chat), so www.matrix went with it (confirmed NXDOMAIN via two
independent DoH resolvers). www.game deleted through IONOS and verified.
#0005 down to a single remaining item: ftp.axion1337.de.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Both done step-by-step with sorb through the IONOS panel, externally confirmed
via DoH: rohana (www removed, Null-MX, SPF -all, DMARC reject — was previously
unprotected) and selendis (IONOS Mail service deactivated to unlock MX
deletion, 3 DKIM CNAMEs + www removed, SPF edited in place, Null-MX, DMARC
reject). Service-record lesson noted for the remaining matrix cleanup.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Previous commit set three issues in-progress, tripping the WIP<=2 rule. The
verification is done; the DNS mutations are sorb's to run in IONOS, so #0001
and #0003 move to waiting (with wartegrund) while #0005 drives the batch.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Verified current axion1337.de zone state via DoH: www.matrix/www.game are
IONOS-default records nothing serves (cluster routes .chat, no matching cert);
rohana is unhardened (no Null-MX/-all/reject) with www.rohana still present;
selendis mail-set untouched; matrix mail-set + autodiscover present; ftp is
IONOS-hosting ballast. Attached a consolidated per-name IONOS action list; the
mutations are sorb's to run in IONOS (no API access from here). Batch in-progress.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Cert renewal switched to DNS-01 (IONOS): test issuance validated end-to-end
(LE YR1, valid to 2026-11-12), the shared letsencrypt resolver now renews
rohana/selendis via DNS-01, so the September renewal needs no open port 443.
Syncs the canonical file with the already-closed git.lab tracker issue 7.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Traefik stack is thread-net-git (manual compose deploy on CFGMON). DNS-01 diff
is ready (tlschallenge -> dnschallenge/ionos + IONOS_API_KEY via host .env).
Two human dependencies remain: create the IONOS API key and deploy+verify on
CFGMON (no SSH from here).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Consolidate the 13 findings + learnings from the Wiki.js AAR into
docs/wiki/stolpersteine/wikijs.md (config/deploy, theming, navigation,
locale/timezone incl. the standing fork patch, access control, git-storage),
link it from the wiki index, and set the AAR status to harvested.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Capture the follow-up work in the Wiki.js AAR: locale migration mechanics
(migrateToLocale only patches pages; rebuild tree+index; nav must move to the
new locale), the new-user timezone source (DB column default), API-only wiki
editing, and the standing upstream deviation (startup sed on users.js) with its
upgrade-check anchor. Stumbles 11-13 + a dated Nachtrag section.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
#0023 (Docusaurus navbar logo) is moot since Wiki.js replaced Docusaurus
(ADR-0014) -> rejected. #0007 (cert renewal, due 2026-09-28) gets a concrete
plan: pursue DNS-01 (approach B, already recommended) before mid-September,
with the port-opening fallback A as a dated calendar checkpoint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sidebar links render href=target verbatim, so page targets need a leading
slash; without it they resolve relatively and 404 from any sub-path. Captured
as stumble #10 for forkers; fixed in gitops set_navigation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Wiki.js→Gitea (sorb/ThreadNetWiki) is wired and verified end-to-end (status
operational, page create/delete propagates). Move ADR-0015 to accepted with an
implementation note; mark the git-storage part of #0048 done. Remaining: the
canonize job Gitea→git.lab (needs the target repo decision).
The cluster cannot reach git.lab (deliberate lab-independence), so #0048's
git.lab-repo-as-storage is infeasible. ADR-0015 routes Wiki.js content
cluster→Gitea→canonize to git.lab, reusing the accepted TURN-rotation pattern;
status proposed, pending sorb's ratification and two prerequisites (Gitea repo +
deploy PAT). Note the flagged contradiction on #0048 and the live theming
progress (logo, shared login background, dark default) on #0050. STATUS regen.
wiki.axion1337.chat (public, Let's-Encrypt TLS, native Authentik-OIDC login, no
forward-auth) added to #0048 (ingress/cert pattern) and #0049 (redirect URI, same
URL for user and admin, role decides). Kept out of ADR-0014 deliberately: accepted
ADRs are not edited, and the hostname is a deployment detail, not a new decision.
ADR-0014 records the decision: Wiki.js replaces Docusaurus for the platform wiki
— the only option meeting both hard requirements (per-group abschotten AND
docs-as-code in git). BookStack ruled out (DB-only, no git). Scope excludes
homelab/docs; neckbeard docs stay in management; content in a dedicated wiki repo
(not a branch, not a monorepo). ADR-0007 set to superseded. #0047 resolved
(decided: Wiki.js). Build issues 0048 (deploy + git storage), 0049 (OIDC + roles/
abschottung: admins write, users read-only), 0050 (theming, colours+logo extracted
from homelab/wiki). #0046 becomes the umbrella. STATUS regenerated; all gates green.
#0024 decided: axionwiki.lab (development-time). #0020 decided: stay on
Docusaurus with Authentik forward-auth, not the last word on the surface. Both
closed with the decision recorded. New follow-ups the user asked to keep:
0046 (move the wiki into the ThreadNet Server Suite, M4) and 0047 (re-examine
surface alternatives beyond BookStack afterwards, M2). STATUS regenerated;
validate, gen_status --check, upstream_drift and pruefe_prosa green.
The job got past the git fix but then failed the urllib call to https://git.lab
with CERTIFICATE_VERIFY_FAILED: gruppenpruefung.py uses urllib's default trust,
which in python:3.12-alpine does not include the private aXionLabs CA. Point
SSL_CERT_FILE at the repo's ci/lab-ca-chain.crt (the same chain curl --cacert
uses); Python honours it in the default SSL context. Verified locally: the
context loads the 2 lab CA certs.
gruppenpruefung.py runs 'git log --all' over the sibling clones (line 108) but
its job used python:3.12-alpine with no before_script, so it would hit the same
FileNotFoundError: 'git' as validate did. Dormant only because the job runs on
schedule/web, not push. Add 'apk add git'. The job's API/CA and GITLAB_TOKEN
prerequisites remain tracked separately (management#31).
pruefe_prosa.py shells out to 'git cat-file' to verify cited commit SHAs, but
the python:3.12-alpine image has no git and before_script only installed pyyaml.
The job crashed with FileNotFoundError on every push since the migration
(pipelines #255, #257). Add 'apk add git' to before_script. Verified green in
the same image locally: validate, gen_status --check, upstream_drift and
pruefe_prosa all pass.