chore: upgrade the neckbeard baseline from v0.1.1 to v0.3.1

Three minor releases at once. Byte-identical surface: only WORKFLOW.md
actually changed (+128/-2); CLAUDE.md and the four templates are
untouched. AGENTS.md takes the new upstream prefix (+22/-2) and keeps our
project section unchanged.

The two declared-extended files were reconciled by hand, because nothing
compares them: validate.py gains upstream's check_vendored_portable, and
schema.yaml gains vendored, project_section_marker, the judged link field
and the ledger and verdict types. gen_status.py needed nothing — upstream
did not touch it. Recorded in the new HERKUNFT.md, including that this
reconciliation has no contradictor and will be forgotten next time.

Newly adopted: check_harvest.py and judge.py, plus docs/ledger/ and
docs/verdict/ with their templates. Deliberately not adopted:
check_locked.py — pruefe_sperrliste.py has done that job here since
2026-08-20, and two tools for one rule is maintenance without gain.

The v0.1.1 baseline stays where it is; docs/sources is immutable.
This commit is contained in:
Thore Cimbal
2026-08-21 12:00:00 +00:00
parent 8191d03137
commit 33a0ee7aec
25 changed files with 4070 additions and 10 deletions
+461
View File
@@ -0,0 +1,461 @@
#!/usr/bin/env python3
"""check_harvest.py — refuse a harvest that carries adopter specifics.
A harvest travels from an adopting project into the framework repository.
Only generalized statements may cross: the failure class, its effect, its
cause, and the framework change they argue for. Names of the adopter, its
products, hosts, stack components, people, paths and identifiers must not.
Three surfaces are checked, because a real leak used all of them:
* file contents
* file names
* commit messages in the range being handed over
Fails closed (exit 1, never "clean"):
* term list missing, unreadable, or empty after stripping comments
* a revision range git cannot resolve
* a file that cannot be read as text — a harvest is prose, and a blob
nobody can read is precisely what nobody can review
Deliberately NOT a fourth surface: the author and committer identity of
those commits. It is a repository-wide property rather than something a
harvest carries in — the framework's own history and LICENSE hold the same
name — so a run over any branch would report it every time. A check that
is permanently red reports nothing, and teaches people to skip it. Commit
identity belongs to the repository's publication decision, not to this
check; verify it there, once, and not in every harvest.
⚠️ This check supplements human review, it does not replace it. A denylist
finds only the nouns somebody thought of, and the paragraph above names a
surface it does not look at by design. Treating a green run as proof of
absence is the same mistake that made the leak it exists to prevent.
The term list belongs to the adopter and is never shipped with the
framework — the framework must not store the names it exists to keep out.
Two rules for writing one:
* one term per line, `#` starts a comment. Matching is case-insensitive
and respects word boundaries, so `zephyr` finds "Zephyr-Web" and
"zephyr.example" but not the middle of an unrelated English word.
Wrap a term in asterisks — `*zephyr*` — to match inside words too,
for the rare name that hides in a compound with no separator.
* a term the framework itself uses is shared vocabulary, not a secret.
Listing it only produces noise that teaches people to ignore the check.
Scope. With `--range`, the lines that range **adds** are read, plus the
names of the files it touches and the messages of its commits. A harvest
answers for what it writes, not for what the repository already carried:
regenerating a shared index would otherwise drag every pre-existing line
of that index into the result. Without `--range`, every tracked file is
read in full — an audit of the repository, which is a different question.
Findings print the term and a masked excerpt: enough to locate the leak,
without repeating the secret in full wherever the output ends up.
Usage:
python scripts/check_harvest.py --terms <file> [--range <a>..<b>] [path ...]
python scripts/check_harvest.py --selftest
"""
from __future__ import annotations
import io
import os
import re
import subprocess
import sys
import tempfile
from contextlib import redirect_stdout
from functools import lru_cache
from pathlib import Path
from typing import Iterator, NamedTuple
MASK = "***"
COMMIT_SEP = "\x1e"
FIELD_SEP = "\x1f"
class HarvestError(RuntimeError):
"""The check could not answer the question. That is never a pass."""
class Finding(NamedTuple):
surface: str # "content" | "filename" | "commit" | "unreadable"
where: str # path, or commit sha
line: int | None # None for filenames, commit subjects, unreadable
term: str
excerpt: str
def git(root: Path, *args: str) -> str:
"""Run git. A failure raises — it never becomes an empty result.
⚠️ This is the core of failing closed. An unresolvable range makes git
exit non-zero; turning that into "" would turn it into "no changes"
and therefore into a silent pass.
"""
done = subprocess.run(["git", "-C", str(root), *args],
capture_output=True, text=True)
if done.returncode != 0:
raise HarvestError(f"git {' '.join(args)}: {done.stderr.strip()[:160]}")
return done.stdout
def load_terms(path: Path) -> list[str]:
"""Read the adopter's term list. Missing, unreadable or empty is an error."""
try:
raw = path.read_text(encoding="utf-8")
except OSError as err:
raise HarvestError(f"term list unreadable: {path} ({err})") from err
terms = [z.strip() for z in raw.splitlines()]
terms = [z for z in terms if z and not z.startswith("#")]
if not terms:
raise HarvestError(f"term list is empty: {path}")
return terms
@lru_cache(maxsize=None)
def term_pattern(term: str) -> re.Pattern[str]:
"""`*x*` matches inside words; a bare term respects word boundaries.
⚠️ Word boundaries are the default because a substring match on a
proper noun hits ordinary language: short personal names sit inside
perfectly ordinary English words — "rene" inside "serene", and the
name that forced this inside the word "authored". Measured, not
hypothesised: a substring run flagged every commit trailer here.
"""
if len(term) > 2 and term.startswith("*") and term.endswith("*"):
return re.compile(re.escape(term[1:-1]), re.I)
return re.compile(rf"(?<!\w){re.escape(term)}(?!\w)", re.I)
def mask(text: str, term: str) -> str:
"""Replace the matched term, keeping the surrounding context readable."""
return term_pattern(term).sub(MASK, text).strip()[:120]
def scan_text(text: str, terms: list[str], *, surface: str,
where: str, numbered: bool = True) -> list[Finding]:
found: list[Finding] = []
for nr, zeile in enumerate(text.splitlines() or [text], start=1):
for term in terms:
if term_pattern(term).search(zeile):
found.append(Finding(surface, where, nr if numbered else None,
term, mask(zeile, term)))
return found
def iter_files(root: Path, paths: list[str],
rev_range: str | None = None) -> Iterator[Path]:
"""Scope: the range's own files, or every tracked file when none given.
A harvest is checked against what it adds. Scanning the whole repository
instead surfaces pre-existing content that the harvest never touched —
noise that buries the findings that matter.
"""
if not paths:
if rev_range:
roh = git(root, "diff", "--name-only", "--diff-filter=d", rev_range)
else:
roh = git(root, "ls-files").replace("\0", "\n")
for name in roh.splitlines():
if name and (root / name).is_file():
yield root / name
return
for roh in paths:
p = Path(roh)
if p.is_dir():
yield from (q for q in sorted(p.rglob("*")) if q.is_file())
elif p.is_file():
yield p
else:
raise HarvestError(f"path does not exist: {p}")
def _kurzname(root: Path, datei: Path) -> str:
"""Path for the report — never a crash.
⚠️ A file handed in explicitly may sit outside the repository, and
`Path.relative_to` raises for those. A check that dies on a path it was
asked to read answers nothing; the reason it is not fatal is that the
only thing wanted here is a label.
"""
try:
return str(datei.relative_to(root))
except ValueError:
return str(datei)
def scan_files(root: Path, paths: list[str], terms: list[str],
rev_range: str | None = None) -> list[Finding]:
found: list[Finding] = []
for datei in iter_files(root, paths, rev_range):
rel = _kurzname(root, datei)
found += scan_text(rel, terms, surface="filename", where=rel,
numbered=False)
try:
inhalt = datei.read_text(encoding="utf-8")
except (UnicodeDecodeError, OSError):
found.append(Finding("unreadable", rel, None, "-",
"not readable as text — cannot be reviewed"))
continue
found += scan_text(inhalt, terms, surface="content", where=rel)
return found
def scan_diff(root: Path, rev_range: str, terms: list[str]) -> list[Finding]:
"""Read the lines a range *adds*, not the files it happens to touch.
⚠️ Reading whole files reports content the harvest never wrote — a
generated index regenerated by the harvest carries every pre-existing
line of the repository into the result. Those belong to the repository
and its own publication decision; a harvest answers for what it adds.
"""
roh = git(root, "diff", "--unified=0", "--no-color", rev_range)
found: list[Finding] = []
datei, nr = "", 0
for z in roh.splitlines():
if z.startswith("+++ "):
ziel = z[4:].strip()
datei = ziel[2:] if ziel.startswith("b/") else ziel
elif z.startswith("@@"):
treffer = re.search(r"\+(\d+)", z)
nr = int(treffer.group(1)) if treffer else 0
elif z.startswith("Binary files") and datei and datei != "/dev/null":
found.append(Finding("unreadable", datei, None, "-",
"binary — cannot be reviewed as text"))
elif z.startswith("+") and not z.startswith("+++"):
if datei and datei != "/dev/null":
inhalt = z[1:]
for term in terms:
if term_pattern(term).search(inhalt):
found.append(Finding("content", datei, nr, term,
mask(inhalt, term)))
nr += 1
for name in git(root, "diff", "--name-only", "--diff-filter=d",
rev_range).splitlines():
if name:
found += scan_text(name, terms, surface="filename", where=name,
numbered=False)
return found
def scan_commits(root: Path, rev_range: str, terms: list[str]) -> list[Finding]:
roh = git(root, "log", f"--format=%H{FIELD_SEP}%B{COMMIT_SEP}", rev_range)
found: list[Finding] = []
for block in roh.split(COMMIT_SEP):
block = block.strip("\n")
if FIELD_SEP not in block:
continue
sha, nachricht = block.split(FIELD_SEP, 1)
found += scan_text(nachricht, terms, surface="commit",
where=sha[:12], numbered=False)
return found
def report(findings: list[Finding]) -> None:
for surface in ("content", "filename", "commit", "unreadable"):
teil = [f for f in findings if f.surface == surface]
if not teil:
continue
print(f"\n {surface} ({len(teil)}):")
for f in teil:
ort = f"{f.where}:{f.line}" if f.line else f.where
print(f" {ort} — term {f.term!r}")
print(f" {f.excerpt}")
def selftest() -> int:
"""Positive and negative controls. A check only ever seen green is a guess."""
fehler: list[str] = []
def pruefe(name: str, bedingung: bool) -> None:
print(f" {'ok ' if bedingung else 'FAIL'} {name}")
if not bedingung:
fehler.append(name)
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
git(root, "init", "-q", "-b", "main")
git(root, "config", "user.email", "selftest@invalid")
git(root, "config", "user.name", "selftest")
liste = root / "terms.txt"
liste.write_text("# comment\nzephyr\nrene\n*brand*\n\n", encoding="utf-8")
terms = load_terms(liste)
(root / "clean.md").write_text("a generalized statement\n", encoding="utf-8")
(root / "body.md").write_text("one\ntwo Zephyr three\n", encoding="utf-8")
(root / "zephyr-notes.md").write_text("clean body\n", encoding="utf-8")
# "serene" contains the name "rene" — the false-positive class that
# forced word boundaries. "xbrandy" is the opposite case, opted
# into with *…*.
(root / "english.md").write_text("a serene afternoon\n", encoding="utf-8")
(root / "touched.md").write_text("old line says Zephyr\n", encoding="utf-8")
(root / "compound.md").write_text("xbrandy\n", encoding="utf-8")
git(root, "add", "-A")
git(root, "commit", "-q", "-m", "initial, no secret here")
(root / "clean.md").write_text("still generalized\n", encoding="utf-8")
(root / "late.md").write_text("added later, says Zephyr\n", encoding="utf-8")
# Touched by the range, but the term sits in a line the range did
# not write — the generated-index case, in miniature.
# Line 1 is pre-existing and must stay silent; line 2 is added by
# the range and must be reported *at line 2* — which also proves the
# hunk header is read instead of assumed.
(root / "touched.md").write_text(
"old line says Zephyr\nadded line also says Zephyr\n", encoding="utf-8")
git(root, "add", "-A")
git(root, "commit", "-q", "-m", "mentions Zephyr in the message")
treffer = scan_files(root, [], terms)
inhalt = [f for f in treffer if f.surface == "content"]
namen = [f for f in treffer if f.surface == "filename"]
commits = scan_commits(root, "HEAD~1..HEAD", terms)
pruefe("1 term in a file body is found, with its line number",
any(f.where == "body.md" and f.line == 2 for f in inhalt))
pruefe("2 term in a filename is found while the body is clean",
any(f.where == "zephyr-notes.md" for f in namen)
and not any(f.where == "zephyr-notes.md" for f in inhalt))
pruefe("3 term in a commit message is found while the tree is clean",
len(commits) == 1)
# The listed term is lower case; the file says "Zephyr". The match
# is only proof of case-insensitivity if the source really differs.
pruefe("4 a lower-case term matches a capitalized occurrence",
"Zephyr" in (root / "body.md").read_text(encoding="utf-8")
and any(f.term == "zephyr" and f.where == "body.md"
for f in inhalt))
pruefe("5 the excerpt masks the term instead of repeating it",
all(MASK in f.excerpt for f in inhalt))
leer = root / "empty.txt"
leer.write_text("# only comments\n", encoding="utf-8")
pruefe("6 an empty term list raises instead of passing",
_raises(lambda: load_terms(leer)))
pruefe("7 a missing term list raises instead of passing",
_raises(lambda: load_terms(root / "nope.txt")))
pruefe("8 an unresolvable range raises instead of reporting nothing",
_raises(lambda: scan_commits(root, "nosuchref..HEAD", terms)))
pruefe("9 clean input against a non-empty list finds nothing",
scan_files(root, [str(root / "clean.md")], terms) == [])
pruefe("10 a bare term does not match inside an unrelated word",
not any(f.where == "english.md" for f in treffer))
pruefe("11 a *term* does match inside a word, when opted into",
any(f.where == "compound.md" and f.term == "*brand*"
for f in treffer))
# Sharp in both directions: the range's own file must be found, and
# the older ones must not. An assertion that only checks for "nothing"
# would also pass if the scoping read no files at all.
bereich = scan_diff(root, "HEAD~1..HEAD", terms)
pruefe("12 with a range, exactly that range's own additions are read",
{f.where for f in bereich} == {"late.md", "touched.md"}
and any(f.where == "body.md" for f in treffer))
beruehrt = [f for f in bereich if f.where == "touched.md"]
pruefe("13 a pre-existing line in a touched file is not reported",
len(beruehrt) == 1)
pruefe("14 an added line is reported at its real line number",
bool(beruehrt) and beruehrt[0].line == 2)
# Everything above tests the scanners directly, which leaves the
# dispatch in main() unproven: it could call the whole-file reader,
# or skip commit messages entirely, and every assertion above would
# still pass. Two discriminators that only the wiring can satisfy:
# whole-file mode reports touched.md at line 1 (the pre-existing
# line) where the diff reader reports line 2, and a missing
# scan_commits call removes the commit section from the output.
zuvor = Path.cwd()
puffer = io.StringIO()
try:
os.chdir(root)
with redirect_stdout(puffer):
main(["--terms", str(liste), "--range", "HEAD~1..HEAD"])
finally:
os.chdir(zuvor)
ausgabe = puffer.getvalue()
pruefe("15 main() routes a range to the diff reader, not whole files",
"touched.md:2" in ausgabe and "touched.md:1" not in ausgabe)
pruefe("16 main() actually reads the commit messages of the range",
"commit (" in ausgabe)
# A path handed in explicitly may sit outside the repository. That
# crashed with a ValueError until a positive control tried it.
aussen = Path(tempfile.gettempdir()) / "check_harvest_outside.md"
aussen.write_text("mentions Zephyr\n", encoding="utf-8")
try:
draussen = scan_files(root, [str(aussen)], terms)
except ValueError:
draussen = []
finally:
aussen.unlink(missing_ok=True)
pruefe("17 a path outside the repository is reported, not fatal",
len(draussen) == 1 and draussen[0].term == "zephyr")
print()
if fehler:
print(f"selftest: {len(fehler)} assertion(s) failed")
return 1
print("selftest: 17 assertions passed")
return 0
def _raises(fn) -> bool:
try:
fn()
except HarvestError:
return True
return False
def main(argv: list[str]) -> int:
if "--selftest" in argv:
return selftest()
terms_pfad: Path | None = None
rev_range: str | None = None
paths: list[str] = []
i = 0
while i < len(argv):
if argv[i] == "--terms" and i + 1 < len(argv):
terms_pfad, i = Path(argv[i + 1]), i + 2
elif argv[i] == "--range" and i + 1 < len(argv):
rev_range, i = argv[i + 1], i + 2
elif argv[i].startswith("--"):
print(f"unknown option: {argv[i]}")
return 2
else:
paths.append(argv[i])
i += 1
if terms_pfad is None:
print(__doc__.strip().splitlines()[-3].strip())
print("check_harvest: --terms is required. Unchecked is not clean.")
return 2
root = Path.cwd()
try:
terms = load_terms(terms_pfad)
if rev_range and not paths:
findings = scan_diff(root, rev_range, terms)
findings += scan_commits(root, rev_range, terms)
else:
findings = scan_files(root, paths, terms)
if rev_range:
findings += scan_commits(root, rev_range, terms)
except HarvestError as err:
print(f"check_harvest: {err}")
print("The question could not be answered — that is not a pass.")
return 1
umfang = f"{len(terms)} term(s)"
if rev_range:
umfang += f", commits {rev_range}"
if findings:
print(f"check_harvest: {len(findings)} finding(s) — {umfang}")
report(findings)
print("\nA harvest carries the failure class, its effect and its cause —")
print("never the adopter's names. Generalize, then run this again.")
return 1
print(f"check_harvest: no findings — {umfang}")
print("⚠️ A denylist finds only what someone listed. Human review still applies.")
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1:]))
+564
View File
@@ -0,0 +1,564 @@
#!/usr/bin/env python3
"""judge.py — read a run's trace against the workflow. Build form A.
A judge cannot examine behaviour, only traces. This one reads a session
ledger (`docs/ledger/`) against `WORKFLOW.md`'s rules and against git,
and reports where the two disagree. It is deterministic, stdlib-only and
model-agnostic; the inferential half of judging is build form B, a
separate ritual in fresh context (see WORKFLOW.md, "Judging a run").
What it checks
* gate rows are ordered, without duplicates, and complete for the
declared size class — S owes no gates, and is not judged as if it did
* every status comes from the four-value vocabulary
* a gate that has a successor carries an approval
* every named commit resolves, is an ancestor of HEAD, and the commits
run in the same order as the gates they belong to
* the ladder section exists, and each entry names what was searched,
what was found, and an outcome beginning `reused:` or `built:`
The commit checks matter more than the rest: they are the only part that
compares the ledger against evidence the ledger's author did not write.
What it deliberately does not do
* judge a session that has no ledger. History cannot be instrumented
after the fact, and a check that reports every past session forever is
a check nobody reads.
* judge quality. Whether a design document is real or filler is form B's
question, and no script can answer it.
* resist tampering. The ledger is written by the agent it describes.
The threat model is drift, not sabotage — ADR-0010 records this as an
explicit non-goal, because a judge that suggests otherwise is worse
than none.
Coverage
`--coverage` reports which rules of the framework can be observed at all
and by what — a script, the ledger, or nothing. Rules at "nothing" are
either unobservable or inert, and both are worth knowing. This is a
report, never a gate: it must not turn red, or it becomes the standing
finding nobody can fix.
Fails closed (exit 1, never "clean"): a ledger that cannot be read or
parsed, and a commit git cannot resolve.
Usage:
python scripts/judge.py --ledger docs/ledger/<file>.md [--root .]
python scripts/judge.py --coverage
python scripts/judge.py --selftest
"""
from __future__ import annotations
import re
import subprocess
import sys
import tempfile
from pathlib import Path
from typing import NamedTuple
STATUSES = ("DONE", "DONE_WITH_CONCERNS", "NEEDS_CONTEXT", "BLOCKED")
GATES_FOR_SIZE = {"S": set(), "M": set(), "L": {"1", "2", "3", "4", "5"}}
OUTCOME_PREFIX = ("reused:", "built:")
PENDING = "pending"
PLACEHOLDERS = {"", "-", "—", "n/a", "none", "tbd", "todo", "?"}
class JudgeError(RuntimeError):
"""The check could not answer the question. Never a pass."""
class Row(NamedTuple):
cells: list[str]
def get(self, i: int) -> str:
return self.cells[i].strip() if i < len(self.cells) else ""
class Finding(NamedTuple):
bucket: str # "model-failure" | "framework-gap"
rule: str
detail: str
def git(root: Path, *args: str) -> str:
"""Run git, or raise. A failure is never an empty result."""
proc = subprocess.run(["git", "-C", str(root), *args],
capture_output=True, text=True)
if proc.returncode != 0:
raise JudgeError(
f"git {' '.join(args)} failed: {proc.stderr.strip() or 'no output'}")
return proc.stdout
# --------------------------------------------------------------------------
# Reading the ledger. Frontmatter is read with re rather than PyYAML: the
# envelope is already validated by validate.py against schema.yaml, so this
# only needs the two or three scalars it acts on — and adopters copying one
# file should not inherit a dependency.
# --------------------------------------------------------------------------
def parse_ledger(path: Path) -> tuple[dict, list[Row], list[Row]]:
try:
text = path.read_text(encoding="utf-8")
except OSError as exc:
raise JudgeError(f"cannot read ledger {path}: {exc}") from exc
match = re.match(r"^---\n(.*?)\n---\n", text, re.S)
if not match:
raise JudgeError(f"{path}: no frontmatter block")
meta: dict[str, object] = {}
current: str | None = None
for line in match.group(1).splitlines():
item = re.match(r"^\s+-\s*(.+)$", line)
if item and current:
meta.setdefault(current, [])
if isinstance(meta[current], list):
meta[current].append(item.group(1).strip().strip('"'))
continue
kv = re.match(r"^([a-z_]+):\s*(.*)$", line)
if kv:
value = kv.group(2).strip().strip('"')
current = kv.group(1)
meta[current] = value if value not in ("", "[]") else []
for field in ("type", "size", "status"):
if not meta.get(field):
raise JudgeError(f"{path}: frontmatter is missing '{field}'")
if meta["type"] != "ledger":
raise JudgeError(f"{path}: type is '{meta['type']}', not 'ledger'")
body = text[match.end():]
return meta, _table(body, "Gates"), _table(body, "Ladder")
def _table(body: str, heading: str) -> list[Row]:
"""Rows of the Markdown table under '## <heading>', header excluded."""
section = re.search(rf"^##\s+{heading}\s*$(.*?)(?=^##\s|\Z)",
body, re.S | re.M)
if not section:
return []
rows: list[Row] = []
for line in section.group(1).splitlines():
line = line.strip()
if not line.startswith("|"):
continue
cells = [c.strip() for c in line.strip("|").split("|")]
if all(re.fullmatch(r":?-{2,}:?", c) for c in cells):
continue # the ---|--- separator
rows.append(Row(cells))
return rows[1:] if rows else rows # drop the header row
# --------------------------------------------------------------------------
# The checks
# --------------------------------------------------------------------------
def check_gate_order(rows: list[Row], size: str, closed: bool) -> list[Finding]:
out: list[Finding] = []
seen = [r.get(0) for r in rows]
numbers = [s for s in seen if s.isdigit()]
if len(set(numbers)) != len(numbers):
out.append(Finding("model-failure", "workflow/gate-order",
f"a gate is recorded twice: {numbers}"))
if numbers != sorted(numbers, key=int):
out.append(Finding("model-failure", "workflow/gate-order",
f"gates are out of order: {numbers}"))
owed = GATES_FOR_SIZE.get(size.upper(), set())
if closed:
missing = sorted(owed - set(numbers))
if missing:
out.append(Finding("model-failure", "workflow/gates-for-size",
f"size {size} owes gates {missing}, not recorded"))
return out
def check_status_vocabulary(rows: list[Row]) -> list[Finding]:
out: list[Finding] = []
for row in rows:
status = row.get(3)
if status and status not in STATUSES:
out.append(Finding("model-failure", "agents/status-vocabulary",
f"gate {row.get(0)}: '{status}' is not one of "
f"{', '.join(STATUSES)}"))
return out
def check_approvals(rows: list[Row]) -> list[Finding]:
"""A gate with a successor must carry an approval."""
out: list[Finding] = []
for i, row in enumerate(rows[:-1]):
if row.get(2).lower() in PLACEHOLDERS:
out.append(Finding("model-failure", "workflow/stop-before-next-gate",
f"gate {row.get(0)} has no approval, but gate "
f"{rows[i + 1].get(0)} was started"))
return out
def check_commits(root: Path, rows: list[Row]) -> list[Finding]:
"""The only part measured against evidence the author did not write."""
out: list[Finding] = []
order: list[tuple[str, int]] = []
for row in rows:
sha = row.get(1)
if not sha or sha.lower() in PLACEHOLDERS:
out.append(Finding("model-failure", "ledger/commit-required",
f"gate {row.get(0)} names no commit"))
continue
try:
full = git(root, "rev-parse", "--verify", f"{sha}^{{commit}}").strip()
except JudgeError as exc:
raise JudgeError(
f"gate {row.get(0)} names commit {sha}, which does not "
f"resolve — unchecked is not passed ({exc})") from exc
ancestor = subprocess.run(
["git", "-C", str(root), "merge-base", "--is-ancestor", full, "HEAD"],
capture_output=True, text=True)
if ancestor.returncode != 0:
out.append(Finding("model-failure", "ledger/commit-reachable",
f"gate {row.get(0)}: commit {sha} is not an "
f"ancestor of HEAD"))
continue
depth = int(git(root, "rev-list", "--count", f"{full}..HEAD").strip())
order.append((row.get(0), depth))
ranked = [gate for gate, _ in sorted(order, key=lambda p: -p[1])]
stated = [gate for gate, _ in order]
if ranked != stated:
out.append(Finding("framework-gap", "ledger/commit-order",
f"the commits run in the order {ranked}, the gates "
f"claim {stated}"))
return out
def check_design_doc(meta: dict, root: Path, closed: bool) -> list[Finding]:
"""Size L owes a design document. The coverage table claimed this
before anything checked it — the inventory drift its own entry warns
about, caught on the first read."""
if str(meta.get("size", "")).upper() != "L" or not closed:
return []
related = meta.get("related") or []
if isinstance(related, str):
related = [related]
if any(r.startswith("docs/design/") for r in related):
return []
return [Finding("model-failure", "workflow/design-doc-for-L",
"size L, but the ledger names no design document in "
"'related'")]
def check_ladder(rows: list[Row], closed: bool) -> list[Finding]:
out: list[Finding] = []
if not rows:
out.append(Finding("framework-gap", "agents/ponytail-ladder",
"the Ladder section is empty — the gate was not "
"passed. A session that built nothing says so."))
return out
for i, row in enumerate(rows, start=1):
searched, found, outcome = row.get(0), row.get(1), row.get(2)
if searched.lower() in PLACEHOLDERS or found.lower() in PLACEHOLDERS:
out.append(Finding("model-failure", "agents/ponytail-ladder",
f"ladder entry {i} names no concrete candidate"))
if not outcome.lower().startswith(OUTCOME_PREFIX):
out.append(Finding("model-failure", "agents/ponytail-ladder",
f"ladder entry {i}: outcome must begin "
f"'reused:' or 'built:', got '{outcome[:30]}'"))
if closed and row.get(3).lower() in PLACEHOLDERS | {PENDING}:
out.append(Finding("model-failure", "ledger/commit-required",
f"ladder entry {i} is still 'pending' in a "
f"closed ledger"))
return out
def judge(root: Path, ledger: Path) -> list[Finding]:
meta, gates, ladder = parse_ledger(ledger)
closed = meta.get("status") == "closed"
size = meta.get("size", "L")
findings = check_gate_order(gates, size, closed)
findings += check_status_vocabulary(gates)
findings += check_approvals(gates)
findings += check_commits(root, gates)
findings += check_design_doc(meta, root, closed)
findings += check_ladder(ladder, closed)
return findings
# --------------------------------------------------------------------------
# Rule coverage. A report, never a gate.
#
# The inventory is maintained by hand because the rules live in prose and
# nothing derives them mechanically. That is a known weakness: it will drift
# from AGENTS.md and WORKFLOW.md unless someone updates it, and nothing
# contradicts that drift today.
# --------------------------------------------------------------------------
COVERAGE = [
# (rule, where it is written, what observes it)
("artifact frontmatter and enums", "AGENTS.md §5", "validate.py"),
("link targets are files, never directories", "AGENTS.md §5", "validate.py"),
("STATUS.md is generated, never hand-edited", "AGENTS.md §4", "gen_status.py --check"),
("accepted ADRs are never edited", "AGENTS.md §4", "check_locked.py"),
("docs/sources is immutable", "AGENTS.md §5", "check_locked.py"),
("a harvest carries no adopter specifics", "WORKFLOW.md", "check_harvest.py"),
("PROJECT.md answers Gate 0", "AGENTS.md §2", "validate.py"),
("a design doc is done iff it sits in done/", "WORKFLOW.md", "validate.py"),
("the ponytail ladder was walked", "AGENTS.md §1", "ledger"),
("a size class is declared for the run", "AGENTS.md §3", "ledger"),
("gates run in order", "WORKFLOW.md", "ledger"),
("a gate is approved before the next begins", "WORKFLOW.md", "ledger"),
("every slice reports one of four statuses", "AGENTS.md §1", "ledger"),
("size L owes a design document", "WORKFLOW.md", "ledger"),
("simplicity: the minimum that solves it", "AGENTS.md §1", None),
("surgical changes: every line traces to the request", "AGENTS.md §1", None),
("slice 1 is a tracer bullet", "WORKFLOW.md", None),
("a check owes proof it can fail", "WORKFLOW.md", None),
("a delivering artifact owes one real result", "WORKFLOW.md", None),
("a closeout names what it made false", "WORKFLOW.md", None),
("reproduce before fixing", "WORKFLOW.md", None),
("incidents get an AAR", "WORKFLOW.md", None),
("a permanent exception owes an ADR", "AGENTS.md §1", None),
]
def coverage_report() -> int:
by_script = [r for r in COVERAGE if r[2] and r[2] != "ledger"]
by_ledger = [r for r in COVERAGE if r[2] == "ledger"]
unobserved = [r for r in COVERAGE if r[2] is None]
total = len(COVERAGE)
print(f"rule coverage: {total} rules inventoried\n")
for title, group in (("observed by a script", by_script),
("observed by the ledger", by_ledger),
("not observable today", unobserved)):
print(f" {title} — {len(group)}")
for rule, where, how in group:
suffix = f" [{how}]" if how else ""
print(f" {rule} ({where}){suffix}")
print()
observed = len(by_script) + len(by_ledger)
print(f" {observed}/{total} observable, {len(unobserved)} decoration "
f"until they gain a duty to leave a trace.")
print("\nThis is a report, not a gate. A rule at zero coverage is either")
print("unobservable or inert; neither is a violation of anything.")
return 0
# --------------------------------------------------------------------------
# Positive control. A gate that has only ever been seen green is a
# hypothesis — WORKFLOW.md, Gate 4.
# --------------------------------------------------------------------------
LEDGER_HEAD = """---
type: ledger
date: 2026-08-21
size: {size}
status: {status}
related:
{related}---
## Gates
| gate | commit | approval | status | note |
|---|---|---|---|---|
{gates}
## Ladder
| searched | found | outcome | commit |
|---|---|---|---|
{ladder}
"""
def _write(path: Path, *, size="L", status="closed", gates="", ladder="",
related=' - "docs/design/d.md"\n'):
path.write_text(LEDGER_HEAD.format(size=size, status=status,
gates=gates, ladder=ladder,
related=related),
encoding="utf-8")
def selftest() -> int:
passed, failed = 0, []
def check_that(name: str, condition: bool) -> None:
nonlocal passed
if condition:
passed += 1
print(f" ok {passed + len(failed)} {name}")
else:
failed.append(name)
print(f" FAIL {passed + len(failed)} {name}")
with tempfile.TemporaryDirectory() as tmp:
root = Path(tmp)
git(root, "init", "-q", ".")
git(root, "config", "user.email", "selftest@invalid")
git(root, "config", "user.name", "selftest")
shas = []
for i in range(5):
(root / f"f{i}.txt").write_text(f"{i}\n", encoding="utf-8")
git(root, "add", "-A")
git(root, "commit", "-qm", f"c{i}")
shas.append(git(root, "rev-parse", "--short", "HEAD").strip())
led = root / "ledger.md"
good_ladder = ("| validate.py | a generic engine | reused: schema entry "
f"only | {shas[0]} |")
rows = lambda gates: "\n".join(
f"| {g} | {shas[i]} | owner | DONE | n |" for i, g in enumerate(gates))
_write(led, gates=rows("12345"), ladder=good_ladder)
check_that("a clean ledger produces no findings",
judge(root, led) == [])
_write(led, gates=rows("13245"), ladder=good_ladder)
check_that("a gate out of order is reported",
any(f.rule == "workflow/gate-order" for f in judge(root, led)))
_write(led, gates=rows("11234"), ladder=good_ladder)
check_that("a duplicated gate is reported",
any("twice" in f.detail for f in judge(root, led)))
_write(led, gates=rows("123"), ladder=good_ladder)
check_that("a closed size-L ledger missing gates 4 and 5 is reported",
any(f.rule == "workflow/gates-for-size"
for f in judge(root, led)))
_write(led, size="S", gates="", ladder=good_ladder)
check_that("size S is not judged against gates it does not owe",
not any(f.rule == "workflow/gates-for-size"
for f in judge(root, led)))
_write(led, gates=f"| 1 | {shas[0]} | owner | FINISHED | n |",
ladder=good_ladder)
check_that("a status outside the vocabulary is reported",
any(f.rule == "agents/status-vocabulary"
for f in judge(root, led)))
_write(led, gates=(f"| 1 | {shas[0]} | | DONE | n |\n"
f"| 2 | {shas[1]} | owner | DONE | n |"),
ladder=good_ladder)
check_that("a gate whose successor began without approval is reported",
any(f.rule == "workflow/stop-before-next-gate"
for f in judge(root, led)))
_write(led, gates="| 1 | deadbee | owner | DONE | n |",
ladder=good_ladder)
check_that("a commit that does not resolve raises instead of passing",
_raises(lambda: judge(root, led)))
git(root, "checkout", "-q", "-b", "side", shas[0])
(root / "side.txt").write_text("x\n", encoding="utf-8")
git(root, "add", "-A"); git(root, "commit", "-qm", "side")
off = git(root, "rev-parse", "--short", "HEAD").strip()
git(root, "checkout", "-q", "main" if _has(root, "main") else "master")
_write(led, gates=f"| 1 | {off} | owner | DONE | n |", ladder=good_ladder)
check_that("a commit that is not an ancestor of HEAD is reported",
any(f.rule == "ledger/commit-reachable"
for f in judge(root, led)))
_write(led, gates=(f"| 1 | {shas[3]} | owner | DONE | n |\n"
f"| 2 | {shas[1]} | owner | DONE | n |"),
ladder=good_ladder)
check_that("commits running in a different order than the gates is reported",
any(f.rule == "ledger/commit-order" for f in judge(root, led)))
_write(led, gates=rows("12345"), ladder="")
check_that("an empty ladder section is reported as a gate not passed",
any(f.rule == "agents/ponytail-ladder"
for f in judge(root, led)))
_write(led, gates=rows("12345"),
ladder=f"| - | - | reused: something | {shas[0]} |")
check_that("a ladder entry naming no concrete candidate is reported",
any("concrete candidate" in f.detail for f in judge(root, led)))
_write(led, gates=rows("12345"),
ladder=f"| validate.py | an engine | it was fine | {shas[0]} |")
check_that("a ladder outcome without reused:/built: is reported",
any("must begin" in f.detail for f in judge(root, led)))
(root / "broken.md").write_text("no frontmatter here\n", encoding="utf-8")
check_that("an unreadable ledger raises instead of reporting nothing",
_raises(lambda: judge(root, root / "broken.md")))
_write(led, gates=rows("12345"), ladder=good_ladder)
check_that("main() exits 0 on a clean ledger",
main(["--ledger", str(led), "--root", str(root)]) == 0)
_write(led, gates=rows("12345"), ladder=good_ladder, related="")
check_that("a closed size-L ledger naming no design document is reported",
any(f.rule == "workflow/design-doc-for-L"
for f in judge(root, led)))
_write(led, gates=rows("13245"), ladder=good_ladder)
check_that("main() exits 1 on a real violation",
main(["--ledger", str(led), "--root", str(root)]) == 1)
total = passed + len(failed)
print()
if failed:
print(f"selftest: {len(failed)} of {total} assertions FAILED")
for name in failed:
print(f" - {name}")
return 1
print(f"selftest: {total} assertions passed")
return 0
def _has(root: Path, branch: str) -> bool:
return subprocess.run(["git", "-C", str(root), "rev-parse", "--verify", branch],
capture_output=True).returncode == 0
def _raises(fn) -> bool:
try:
fn()
except JudgeError:
return True
except Exception:
return False
return False
def report(findings: list[Finding]) -> None:
if not findings:
print("judge: no findings")
return
print(f"judge: {len(findings)} finding(s)\n")
for bucket in ("framework-gap", "model-failure"):
group = [f for f in findings if f.bucket == bucket]
if not group:
continue
print(f" {bucket} ({len(group)}):")
for f in group:
print(f" [{f.rule}] {f.detail}")
print()
print("A finding is the model not following a clear rule, or a gap in the")
print("framework. Deciding which is the point — see ADR-0010.")
def main(argv: list[str] | None = None) -> int:
args = list(sys.argv[1:] if argv is None else argv)
if "--selftest" in args:
return selftest()
if "--coverage" in args:
return coverage_report()
if "--ledger" not in args:
print("judge: --ledger <file> is required (or --coverage / --selftest)",
file=sys.stderr)
return 1
ledger = Path(args[args.index("--ledger") + 1])
root = Path(args[args.index("--root") + 1]) if "--root" in args else Path(".")
try:
findings = judge(root.resolve(), ledger)
except JudgeError as exc:
print(f"judge: {exc}", file=sys.stderr)
print("unchecked is not passed.", file=sys.stderr)
return 1
report(findings)
return 1 if findings else 0
if __name__ == "__main__":
sys.exit(main())
+4 -2
View File
@@ -4,7 +4,7 @@
Schützt die übernommenen Framework-Dateien vor stillem Umschreiben
(Entscheidung 8 im Design 2026-08-11, Frage von sorb: „wird die
AGENTS.md ggf. durch Agenten umgeschrieben?"). Die Baseline liegt unter
docs/sources/upstream/neckbeard-v0.1.1/ (siehe HERKUNFT.md dort); ein
docs/sources/upstream/neckbeard-v0.3.1/ (siehe HERKUNFT.md dort); ein
Framework-Upgrade aktualisiert Baseline und Arbeitskopie im selben,
bewussten Commit.
@@ -22,7 +22,7 @@ from __future__ import annotations
import sys
from pathlib import Path
BASELINE = "docs/sources/upstream/neckbeard-v0.1.1"
BASELINE = "docs/sources/upstream/neckbeard-v0.3.1"
MARKE = "<!-- projektabschnitt -->"
# (Arbeitskopie, Baseline-Datei) — byte-identisch
@@ -33,6 +33,8 @@ PAARE = [
("docs/design/template.md", "templates/design-template.md"),
("docs/aar/template.md", "templates/aar-template.md"),
("docs/issues/template.md", "templates/issue-template.md"),
("docs/ledger/template.md", "templates/ledger-template.md"),
("docs/verdict/template.md", "templates/verdict-template.md"),
]
+29
View File
@@ -161,6 +161,32 @@ def check_body_links(path: Path, body: str, root: Path, inbound: set) -> None:
inbound.add(target)
def check_vendored_portable(root: Path, vendored: list, marker: str) -> None:
"""Files adopters hold byte-identical must carry no repo-relative link.
They are copied into repositories without this repo's docs/, so such a
link resolves here and nowhere else — and the adoption path in
AGENTS.md §5 asks for a byte-for-byte copy, which makes the dead link
the adopter's problem and unfixable without breaking the copy.
"""
for rel in vendored or []:
path = root / rel
if not path.is_file():
continue
body = path.read_text(encoding="utf-8")
if marker and marker in body:
# Everything below the marker is the adopting project's own
# section: never copied elsewhere, so its links are fine.
body = body.split(marker, 1)[0]
body = re.sub(r"```.*?```", "", body, flags=re.S)
body = re.sub(r"`[^`\n]*`", "", body)
for link in (m.group(1) for m in INLINE_LINK_RE.finditer(body)):
if link.startswith(EXTERNAL_PREFIXES) or link.startswith("#"):
continue
err(path, f"vendored file carries a repo-relative link: {link} "
f"— name the target instead of linking to it")
def apply_rules(path: Path, rel: str, meta: dict, spec: dict) -> None:
for rule in spec.get("rules", []):
if rule == "superseded_requires_pointer":
@@ -184,6 +210,9 @@ def main() -> int:
link_fields = schema.get("link_fields", [])
types = schema.get("types", {})
inbound: set = set()
check_vendored_portable(root, schema.get("vendored", []),
schema.get("project_section_marker", ""))
wiki_pages: list[tuple[Path, dict]] = []
# Root documents: inline links must resolve; no frontmatter required.