Scheduled rotation (Issue #38). New secret generated, re-encrypted with
the scoped rotation age key, checksum/rotated-at annotations bumped so
Flux restarts coturn + synapse-main on merge. Please review and merge.
Every other release in this repo pins an exact version. Exactly two used a
range, and exactly those two stopped receiving updates: 0.12.6 is the last 0.x
alloy chart and 5.37.0 the last 5.x kube-state-metrics, so both ranges had been
sitting on a ceiling that reads as 'stays current'. A pinned version at least
looks stale.
Rendered both versions of each chart against our real values before changing
anything, because a chart jump moved an internal endpoint out from under a
network rule earlier today. Same objects, same selectors, same service names;
kube-state-metrics names its targetPort instead of numbering it and swaps
endpoints for endpointslices, neither of which our policies or scrapes touch.
nginx for the docs server goes along: 1.26-alpine carried two criticals,
1.31.4-alpine measures clean.
Until ESS 26.4.0 the auth service pointed straight at matrix-stack-synapse-main,
which allow-ingress-synapse permits, so nothing was needed here. ESS 26.8.0 moves
matrix.endpoint to the haproxy service — and MAS walked into the default-deny.
Login broke with 500 'failed to provision device' on /oauth2/token.
The rule was mine, from #0088 this morning. It named who may reach haproxy, and
the list was complete for the topology of that hour. A chart decided otherwise
six hours later.
The CVE gain is the smallest of today's steps: 17 critical findings become 6 and
217 high become 120 across Synapse, MAS, element-admin and lk-jwt-service.
The real reason is in the Synapse changelog rather than the CVE column. v1.157.0
fixes a bug introduced in v1.150.0 where reactivating a deactivated and erased
user did not restore their profile, breaking login, name changes and invitations.
We run v1.151.0, inside that range, and that is precisely what happened to an
account this morning.
Everything breaking in the span was checked against our state and none of it
applies: no Synapse workers, the stable auth integration already in use rather
than MSC3861, and no livekitAuth values set. Our elementWeb override to the fork
survives — both chart versions were rendered with our real values and compared.
Not 26.8.1: it appeared today. A fresh backup of the Synapse and MAS databases
plus media was taken by hand beforehand, and sorb holds a host snapshot.
Seven minors in sequence, each with the latest patch, exactly as the
documentation requires. Every step was verified before the next one started:
three deployments on the new version and fifteen certificates Ready.
The measured point of the exercise: v1.14.0 carried 8 critical and 128 high
findings across controller, webhook and cainjector; v1.21.1 carries 0 and 24.
The next scan round will show it in the dashboard rather than in this message.
This is the step with the RBAC narrowing: the cert-manager-edit aggregate
ClusterRole no longer grants create on challenges or create, patch and update on
orders. Those resources belong to cert-manager's own ACME workflow. Nothing here
is bound to that ClusterRole — zero bindings across all namespaces — so no
tooling loses a permission it was using.
Previous step verified: three deployments on v1.19.6, fifteen certificates
Ready.
This is the step with the ACME metric label change: the high-cardinality path
label on certmanager_acme_client_request_count and _duration_seconds is replaced
by a bounded action label. Nothing here uses those metrics — neither the
operating stack's rules nor any dashboard references them — so no dashboard or
alert has to follow.
Previous step verified: three deployments on v1.18.6, fifteen certificates
Ready.
One minor at a time with the latest patch, as the documentation requires. The
previous step is verified: all three deployments on v1.17.4 and fifteen
certificates Ready.
One minor at a time with the latest patch, as the documentation requires. The
previous step is verified: all three deployments on v1.16.5 and fifteen
certificates Ready.
One minor at a time with the latest patch, as the documentation requires. The
previous step is verified: all three deployments on v1.15.5 and fifteen
certificates Ready.
Seven minors behind, and the documentation allows only one minor at a time with
the latest patch of each; skipping is offered solely as uninstall and reinstall.
Measured on the images rather than assumed: v1.14.0 carries 8 critical and 128
high findings across controller, webhook and cainjector; v1.21.1 carries 0 and
24. Both potentially breaking changes on the way were checked against our state
and do not apply — no dashboard or alert uses the ACME metrics whose label
changes, and nothing is bound to the cert-manager-edit ClusterRole whose
permissions narrow.
Fifteen certificates are Ready before this starts; that is the check after every
step.
Both steps are verified: migrations applied without inconsistency, blueprints
unchanged, and a real login recorded server-side on each version. Outside a
migration window an automatic rollback is the right behaviour, so remediation
goes back to three retries.
The intermediate step is verified: migrations applied without inconsistency,
both pods ready, blueprints unchanged at seventeen flows and two providers, and
a real login through Element, MAS and Authentik recorded server-side at 14:51.
This step adds the remaining high findings. Measured on the images: 2026.5.6
carries 5 critical and 92 high, 2026.8.0 carries 5 and 21. What stays is
perl-base and libxml2, neither of which has a fix available.
A fresh backup was taken first. The one from before the intermediate step sits
on the old schema and would be the wrong way back now.
Authentik's documentation forbids skipping major releases and requires the
latest minor of each before moving on, so 2026.5.6 comes before 2026.8.0 rather
than being an optional stop.
It is also where the benefit is. Measured on the images rather than assumed:
2026.2.3 carries 27 critical and 477 high findings, 2026.5.6 carries 5 and 92.
This one step removes every critical finding that can be removed here; what
remains is perl-base and libxml2, neither of which has a fix. The second step
adds high findings only.
Both breaking changes of this release were checked against our state and do not
apply: the deprecated Postgres connection options are not set anywhere, and
there are no outposts whose version would have to match.
A manual backup was taken immediately before this, because the nightly one is
hours old and the way back is a database restore rather than a version rollback.
The release carried upgrade remediation with three retries and no strategy. The
Flux CRD is explicit: the strategy defaults to rollback, remediation runs between
each attempt, and the last failure is remediated as well whenever retries exceed
zero. A failing upgrade would therefore have rolled Helm back to the old version
up to four times, against a database Django had already migrated forward — the
migration inconsistency Authentik's own documentation warns about, triggered by
this line.
Here the way back is a database restore, not a version rollback, so an automatic
rollback cannot help and can only deepen the damage. It goes back to three once
the jump is done; outside a migration window the remediation is right.
Nothing else changes: the chart version, the values and the install remediation
are untouched.
Its configuration was read from the running process yesterday: the
homeserver by internal service name, Authentik by its public name. Public
names resolve to the node address, and traffic there is not subject to the
policy at all - so the upstream path survives the restriction while
everything else goes.
Held back until now on purpose, because a mistake here hits sign-in and the
analysis alone is not an acceptance. The owner is testing a login against
this change; if it fails, the revert is one commit and about a minute,
rehearsed in wave 1.
Congruence rechecked: twenty-one excluded, the same twenty-one narrowly
ruled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
An open egress that nobody explains reads as neglect. These five are
decisions, so the reasons sit at the rule, where the next reader looks -
not in an issue and not in an ADR alone.
Synapse federates to arbitrary servers and previews addresses users pick;
there is no list to write, and the right control sits a level higher in
url_preview_ip_range_blacklist. Coturn and the SFU relay media to arbitrary
clients - that is the service. ClamAV pulls signatures from a CDN, and
pinning would break the update silently, which is the failure class this
project keeps finding.
MAS is now only there out of caution, and the comment says so. Read from
the running process today: its upstream is Authentik at its public name and
the homeserver at the internal service name. The public name resolves to
the node address, which the policy does not cover anyway - so MAS is
restrictable and only the sign-off is missing, because a mistake there hits
sign-in.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
authentik and monitoring carried only the metadata block, so every pod
there could reach anything. Both now follow the pattern from matrix: the
broad policy keeps an exclusion list and stays as the catch-all, and each
workload gets a rule of its own.
The destinations were read, not guessed. Alloy ships to 10.0.0.3 on 3100
and 9090 - taken from its running configuration. Authentik sends mail
through smtp.ionos.de:587, and that is load-bearing rather than optional:
the blueprints use password recovery and invitations by mail. The database
and kube-state-metrics speak to nobody outside.
Mail gets a /27 rather than two /32. The name resolves to .97 and .113
today, both in the provider's own block; a third address would break mail
with nobody watching a rule, and thirty-two addresses of one provider's
mail infrastructure is the smaller price. That is the opposite call to
ClamAV on purpose - this is not a CDN in front of half the internet.
Deliberately not allowed: authentik's version check. It sits behind
Cloudflare with rotating addresses and nothing depends on it, so the call
fails and gets logged. The manifest says so, because whoever finds that
error later should know it is intended.
Congruence checked in both namespaces before pushing: excluded and
narrowly ruled are the same sets.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Both backups reach the Hetzner storage box over SSH; the TURN rotation and
Wiki.js reach rohana. Everything else is now refused for them.
Two separate policies rather than entries in egress-nur-intern, because a
policy's egress rules apply to every pod it selects: hanging the storage
box off the shared policy would hand it to fourteen workloads that have no
business there.
rohana is allowed as 10.0.0.3/32, not as its public address. The DNS
pointer from the previous commit sends the public name down the private
path, and the certificate is valid there - both measured after the change,
not assumed. So the group's one deliberate Gitea exception no longer leaves
the cluster at all.
Congruence checked before pushing, and this is the check worth keeping:
twenty names are excluded from the broad rule and exactly the same twenty
are covered by a narrow one. An entry on one side only would either cut a
workload off completely or leave the restriction inert, and neither is a
syntax error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
rohana.axion1337.de resolves publicly to a Hetzner address, while the host
is reachable over the private network at 10.0.0.3 - and serves a valid
Let's Encrypt certificate for that very name there, measured with full
verification. Without an internal pointer every access from the cluster
leaves it for no reason and needs an outbound exception.
A dedicated zone rather than a second hosts block: the Corefile already
runs hosts /etc/coredns/NodeHosts in the main block, so a second one there
would collide. The .server import at the end of the Corefile takes a zone
of its own, and the reload plugin picks the change up without a restart.
The price is in the file, not in a commit message nobody rereads: two paths
become one. If 10.0.0.3 is down, rohana is unreachable from the cluster
although the public route would work, and the failure looks like "Gitea is
gone" rather than "the private path is gone". The comment says where to
look first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
matrix-rtc-authorisation-service, wikijs-config, init-secrets,
synapse-check-config and both deployment markers have no demonstrable need
to leave the cluster. They now get DNS plus the pod and service networks
and nothing else.
The name goes into BOTH lists, and that is the whole point: NetworkPolicies
are additive, so as long as the broad policy still selects a pod and allows
0.0.0.0/0, a second and stricter rule for the same pod changes nothing. The
file already carried that warning as a comment; this change obeys it rather
than rediscovering it.
Checked before pushing: both selector lists are congruent - no workload is
excluded from the broad rule without receiving the narrow one, and none the
other way round. A one-sided entry would either open a pod completely or
cut it off entirely, and neither shows up as a syntax error.
Class A is deliberately the first wave: these workloads run at deploy time,
so a mistake surfaces at the next rollout rather than in a user's face.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
Counterpart to restore-drill.yaml: that one covers the databases, this the
Synapse media store. Separate job on purpose - the database probe needs a
throwaway Postgres, this one the production PVC read-only, and folding two
different permission and failure pictures into one job makes an emergency
harder to diagnose, not easier.
The name is deliberate. BackupJobFailed already matches restore-drill.*,
so a failure is covered without a new rule and reaches the maintenance
room through the single alertmanager route.
It re-proves its own comparison every run. After the check passes, one
shared file is altered by a byte and the comparison must report it -
otherwise the job fails with "this probe proves nothing". A comparison
that has only ever said "equal" is a guess, and that stays true when it
runs monthly rather than once.
The intersection carries the proof, not the totals: the backup is a
snapshot while production keeps running, and Synapse prunes its own
preview caches. One-sided files are therefore tolerated in url_cache and
url_cache_thumbnails and are an error anywhere else.
The production PVC is mounted read-only and only emptyDir is written. On
today's single node a second read-only mount of the ReadWriteOnce volume
is unproblematic; the comment says what would have to change if a second
node ever appeared, because that failure would alert without anything
being wrong with the data.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2Q4Ri8NGwyTZzScvKnWFM
169.254.0.0/16 fehlte - das Netz des Cloud-Metadaten-Dienstes. Die Vorschau holt
Adressen, die Nutzer bestimmen; die Sperre gehoert deshalb auch an diese Stelle
und nicht nur in die NetworkPolicy eine Ebene tiefer. Dazu 100.64.0.0/10.
Aufgefallen bei sorbs Frage, warum ich vorgeschlagen hatte, die URL-Vorschau
abzuschalten. Der Vorschlag war falsch: url_preview_ip_range_blacklist ist die
richtige Kontrolle an der richtigen Stelle und war gesetzt - ich haette pruefen
muessen statt eine Funktion in Frage zu stellen. Beim Pruefen kam diese Luecke
heraus, die mehr wert ist als der Vorschlag.
Sie sprechen ausschliesslich cluster-intern: die drei nginx-Dienste, beide
Postgres, der ClamAV-Vorschalter, haproxy, die Wiki-Gruppenpruefung und die
beiden Matrix-Bots. Fuer sie faellt der pauschale 0.0.0.0/0-Satz weg; DNS und
Cluster-Netz bleiben.
⚠️ Der Ausschluss steht in egress-block-metadata selbst, nicht in einer neuen
strengen Policy. NetworkPolicies sind additiv - solange irgendeine Policy
0.0.0.0/0 fuer einen Pod erlaubt, aendert eine zweite, strengere nichts. Genau
dieser Fehler waere die naheliegende Umsetzung gewesen.
Zwei Dinge vorher gemessen statt angenommen:
- waehlt auch Pods, die das Label gar nicht tragen (29 Pods, 1 Treffer,
notin liefert 28). Ohne diese Gewissheit haetten drei Pods ohne
app.kubernetes.io/name ihren Metadaten-Schutz verloren.
- Die Selektoren gegen den laufenden Cluster geprueft: 20 behalten breiten
Ausgang, 9 laufende Pods werden eng - Summe 29, also genau die Gesamtzahl.
Zwei Policies, weil ein podSelector nicht ueber zwei Label-Schluessel odern kann;
drei Pods im Namespace fuehren nur .
Rollback: beide neuen Policies loeschen, Ausschluss zuruecknehmen.
Der erste echte Lauf scheiterte mit "Connection refused", obwohl Service,
Endpunkt, Portname, Pod-Marken und Policy-Regel alle korrekt waren.
Ursache gemessen, nicht geraten: Die NetworkPolicy-Regeln fuer einen NEU
erzeugten Pod sind beim Start des Containers noch nicht programmiert. Im selben
Pod nacheinander: psql sofort abgewiesen, roher TCP-Test zu, nach 20 Sekunden
psql erfolgreich.
Ein frueherer nc-Test aus einem separaten Pod lief zufaellig spaet genug und sah
sauber aus - er hat mich zunaechst in die falsche Richtung geschickt.
Der Job wartet jetzt bis zu 60 s auf Erreichbarkeit und meldet erst danach einen
Fehler. Ohne das scheiterte er jede Nacht an einem Wettlauf, und man gewoehnt
sich an den roten Job - genau die Abstumpfung, die #0104 abgestellt hat.
⚠️ Dieselbe Klasse kann andere kurzlebige Jobs im Namespace treffen
(wikijs-backup, synapse-backup, restore-drill). Nicht geprueft, eigener Befund.
Letztes offenes Abnahmekriterium des Issues: Der Fall 'angemeldet, aber ohne
Gruppe' muss erkennbar sein. Bei @apo fiel er erst auf, weil er sich beschwerte -
Wiki.js meldet diesen Zustand nirgends.
Taegliche Stichprobe auf der Datenbank statt Instrumentierung des Logins: Ein
Patch am Anmeldeweg waere dafuer unverhaeltnismaessig, und ein Tag Verzug ist
hinnehmbar - vorher lag der Fall tagelang unbemerkt. Ausgabe geht ueber Alloy
nach Loki, Suchbegriff 'Wiki-Nutzer ohne Gruppe'.
Der Pod bekommt eine eigene NetworkPolicy-Regel. Ohne sie kaeme er nicht an die
Datenbank, und zwar lautlos - die Datenbank laesst nur wikijs und wikijs-backup
zu (AGENTS.md: jeder neue Pod braucht seine eigene Regel).
Das Skript prueft zuerst, ob die Datenbank ueberhaupt antwortet. Ohne diese
Gegenprobe saehe 'keine Nutzer ohne Gruppe' bei unerreichbarer DB genauso aus wie
ein sauberes Ergebnis - stiller Erfolg ist hier das groessere Risiko als ein Fund.
Beide Zweige gegen die laufende Datenbank getestet, lesend: der gute Fall meldet
'OK (4 geprueft)', der Warn-Zweig nennt die Betroffenen und den Behebungsweg.
Mit dem Chart-Default Cluster verteilt kube-proxy NodePort-Pakete selbst weiter
und ersetzt dabei die Absenderadresse. Der SFU sieht den Client deshalb nie unter
seiner echten Adresse - am 2026-08-19 war das gewaehlte ICE-Paar prflx 10.42.0.1,
eine peer-reflexive Adresse aus dem Cluster-Netz. Das macht die Kandidatenwahl
unnoetig instabil; Local ist die dokumentierte LiveKit-Empfehlung fuer Kubernetes.
Ohne Nachteil, weil der Cluster aus einem Knoten besteht - der uebliche Preis
(nur Knoten mit laufendem Pod nehmen Verkehr an) kann nicht greifen. Bei einem
zweiten Knoten neu zu bewerten.
Lokal gegen den Chart gerendert, nicht angenommen: beide NodePort-Dienste
bekommen Local, der ClusterIP-Dienst bleibt unberuehrt.
Behebt NICHT das schubweise ICE-Flappen (16.08. = 119 Wechsel, 18.08. = 0) - das
ist aelter und liegt an der Zahl der Netzwerkpfade auf Client-Seite.
Release des Upstream-Anschlusses (ADR-0022). Gleicher Quellstand wie rc.3
(8ca03fe), das die Abnahme bestanden hat - nur unter Release-Nummer. Die Images
sind nicht bitgleich, weil der Build die Versionszeichenkette aus git describe
ins Artefakt backt.
Produktion laeuft damit erstmals auf einem Fork, der wieder an der
Upstream-Historie haengt.
Rueckhebel: Tag zurueck auf v0.5.4.
Zweiter Anlauf des Upstream-Anschlusses (ADR-0022). Entfernt den Merge-Rest,
der rc.2 die Raumliste brach, und bringt einen typecheck-Job mit, den
docker_web als needs fuehrt - dieses Image ist das erste, das ohne bestandene
Typpruefung gar nicht haette entstehen koennen.
Abnahme steht aus, in dieser Reihenfolge: Raumliste laedt, ClamAV per Zip,
ClamAV per .png (der umgezogene Bild-Pfad), Call-Teilnehmerliste.
Rueckhebel: Tag zurueck auf v0.5.4.
react-soft-crash bei sorb (Rageshake 2026-08-19 15:17, Safari):
"Setting 'feature_room_list_sections' does not appear to be a setting."
aus SettingsStore.getValue in RoomListItemViewModel.generateItemSync - also
bei jedem Raumlisteneintrag.
Fehler in der Merge-Aufloesung von ADR-0022: Upstream hat den Labs-Schalter
feature_room_list_sections entfernt (Sektionen laufen jetzt ueber
RoomList.showSections). Settings.tsx hat Upstreams Fassung uebernommen, in
RoomListItemViewModel.ts blieb die alte getValue-Zeile daneben stehen.
Der Build konnte das nicht fangen: getValue nimmt einen String, der Fehler
entsteht erst zur Laufzeit.
Kandidat kommt nach dem Fix als rc.3 zurueck.
Kandidat, kein Release. Bringt den Merge aus ADR-0022 in Produktion, damit die
Abnahme an einem echten Client stattfinden kann.
Zu pruefen sind die zwei Patches, die der Merge verschieben musste:
ClamAV-Fehlermeldung im Bild-Pfad und die Call-Teilnehmerliste in der Raumliste.
Der Datei-Pfad (Zip) ist unberuehrt und diente heute als Ausgangswert - der
Scanner meldete die EICAR-Datei erwartungsgemaess zweimal, beim Senden und beim
Empfangen.
Rueckhebel: Tag zurueck auf v0.5.4.
Decision sorb. Hides the edit button beside the server name, so the homeserver can no
longer be switched through the UI, and the 401/403 login error now names the server
rather than staying generic.
Honest about its reach, in the comment as well as here: it is a surface restriction.
MatrixChat still takes hs_url from the query string in two registration flows without
consulting this setting, so a crafted link is unaffected. Against
GHSA-wrcp-5v3v-3j6v - open since 2026-07-20, affecting everything below 1.12.22 while
we run 1.12.17 - it narrows the way in without closing it. The update in management
#0099 remains the actual fix.
The matching line went into the desktop client separately, since that one carries its
own config.json.
Decision sorb. Measured basis rather than preference: four months of operation with
zero destinations, zero remote users and zero rooms with outside participation, while
the federation API answered publicly - the delegation routes it over 443, so 8448
being shut never mattered.
An empty list federates with nobody and one entry opens it for exactly that domain,
so the capability stays one line away rather than gone.
The comment records what must not be done instead, because it is not obvious and it
would look correct: blocking /_matrix/federation at the edge. lk-jwt-service verifies
OpenID tokens through /_matrix/federation/v1/openid/userinfo and reaches it over the
public name - no hostAliases, ClusterFirst DNS - so a path-level block kills group
calls. Synapse serves that endpoint without an X-Matrix signature, so the whitelist
does not touch it.
Caught while validating: the first version of this edit split the auto_join block,
moving auto_join_rooms_for_guests under federation. Functionally identical after the
fragments merge, wrong to read, and fixed before pushing - the diff is now 20 added
lines and nothing moved.
First egress rule in matrix, authentik and monitoring. It allows DNS, the cluster
ranges and the whole internet, and denies only 169.254.0.0/16 - link-local, where
Hetzner serves instance metadata unauthenticated to any pod.
Deliberately narrow. The textbook cut, 0.0.0.0/0 except RFC1918, would have severed
two things here, both over 10.0.0.3 on the private Hetzner network: Alloy writes
metrics and logs there, and the TURN rotation reaches Gitea through a hostAlias to
that address. Private ranges therefore stay open.
The payoff is modest and should be stated as such: measured from a pod, the service
answers with instance-id, hostname, region, MAC and network config, while userdata
and public-keys are empty. No credentials are exposed here, unlike the AWS case this
hardening usually targets. It costs nothing though, and it closes the class.
Two preconditions checked rather than assumed, because both are the usual way this
breaks: kube-system carries kubernetes.io/metadata.name so the DNS rule actually
matches, and the cluster is IPv4-only so 0.0.0.0/0 really does cover everything.
Rollback is deleting the one policy per namespace.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Decision sorb, and it was already made in #0049: normal Authentik users read the
user documentation, admins are admins. The role model was implemented; the way in
was not. selfRegistration created an account on first login and autoEnrollGroups
was empty, so the account landed in no group at all - and since Guests is stripped
of every permission, the user saw nothing and was told nothing about why. That is
#0103, and it happened to a real person.
Admins stay manual: membership in "authentik Admins" arrives through the groups
claim and is not affected by this baseline. betrieb/* keeps its default deny, so
the separation #0049 verified end to end still holds - it only stops applying to
people who were never let in at all.
The lookup aborts if wiki-anwender is missing rather than silently enrolling into
nothing, which would reproduce the exact failure this fixes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On Safari LiveKit silently skipped the sender track swap, so the raw microphone
stayed on the wire regardless of the suppression level. The fork now verifies
and enforces the swap and states the outcome in the console.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The gate opens after the passed two-person acceptance: checkbox and slider are
back in the in-call audio settings. Rollback lever for any regression is the
gate in threadnet-call, not a deployment revert.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ships threadnet-call df4e5ee: the AI filter attaches to the microphone track
after publication with its own AudioContext on just that track. The feature
gate stays closed, so this behaves identically to v0.5.1 for every user; a
single test client opts in via two localStorage keys. The gate opens only
after the filter passes a two-person call - standing rule from #0054.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ships threadnet-call dcc8643: the AI filter's off-path is byte-identical to
upstream again (no processor key, noiseSuppression untouched) and the feature is
hard-gated off until the webAudioMix decision. The gate also covers clients that
still have the setting enabled in localStorage. Regression tests pin both cases
and were demonstrably red on the broken code.
Acceptance is a real two-person call after the rollout; v0.4.3 remains one
tag-revert away.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The console pinned it: with the filter on, LiveKit refuses the processor because
Element Call constructs the room without webAudioMix, so no local audio track
ever carries an AudioContext. With the filter off the same build does publish
its track, so the opt-out path itself is intact.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
sorb asked for it: three hypotheses about the broken unmute were disproven from
the outside, so the browser console is the only remaining source. Calls stay
broken while this runs. Goes back to v0.4.3 as soon as the console is captured.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Calls connect but no participant can unmute, and the SFU log shows not a single
published track. v0.5.0 is the first production image carrying the AI noise
suppression code in the audio capture path, and v0.4.3 is the last image calls
demonstrably worked on. Restoring service first; the cause is still open.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ships ADR-0018: DeepFilterNet3 as an opt-in filter in the call widget, default
off, checkbox plus slider, 35 % by default. The image now carries 23 MB of model
assets under /widgets/element-call/assets/dfn3/; they load when the user turns
the filter on, not on page load, so anyone leaving it off pays nothing.
The .7 package would have shipped a filter that was dead inside the widget and
nowhere else. Verified through the chain instead of trusting the green build:
npm package, node_modules, webpack output, and the CI artifact all carry the
assets at the path the widget requests. The last link — the running pod — gets
checked after this syncs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Reporting worked but ended in silence: report_event.admin_message_md was unset,
so a user who reported content saw no indication of whether it reached anyone or
whom to follow up with. For a moderated community that is an open edge.
sorb's decision is route B — reports stay in the server's event_reports store and
are reviewed through Element Admin; Draupnir deliberately does not get server
admin rights, which polling would have required. The message therefore names a
person rather than promising an automatism, and @sorb is the only admin who can
actually see the reports.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Authentik's own uniqueness is case-sensitive, so 'Boje' and 'boje' pass as
distinct while Matrix treats them as the same localpart. ADR-0011 closed the
takeover vector with on_conflict:fail, but that only bites at login: the user
registers happily and fails later with no explanation. This policy answers where
the mistake is made.
Deliberately reads only prompt_data and never request.user — the stage runs in an
anonymous enrollment context, which is exactly what the previously attached system
policies crashed on.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The running container was 4.10.0 while :latest had moved on to 4.17.2 — with
imagePullPolicy IfNotPresent the node keeps whatever it pulled once, so nobody
knew what was actually running and the next reschedule onto a fresh node would
have jumped seven minor versions silently. That is the concrete case #0052 is
about, and it also explains why the CVE scanner reported against a moving target.
Pinned to 4.17.2, which is both current and what :latest resolves to today, so the
scan results finally describe the thing that runs. The config uses only long-lived
core options (realm, use-auth-secret, relay-ip, cert/pkey), none of them removed
in that range. busybox in the init container goes 1.28 to 1.36, the version this
repo already uses elsewhere.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Restores the Borg archives into a throwaway postgres inside the pod and passes
only when rows actually land — the pg_restore exit code is not proof, counted
rows are. Production is never touched; the repos are only read.
Automated rather than a documented cadence: a check nobody performs is the same
mistake as an untested backup, one level up. Runs on the 4th at 04:20, after the
nightly jobs. Verified manually before commit (synapse 31908 rows, MAS 16085,
wiki 251).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>