diff --git a/AGENTS.md b/AGENTS.md index f204e09..80208f5 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -26,7 +26,10 @@ use judgment — but say so. - Before writing new code, stop at the first rung that holds: needed at all? → codebase already has it? → stdlib? → platform-native? → installed dependency? → one line? → only then: the minimum that works. - (Ladder after ponytail, MIT.) + (Ladder after ponytail, MIT.) **Name the rung that holds, and record the + walk in the run's ledger** — searched, found, outcome. A rung that held + and a rung never tried leave the same diff, so without the record this + rule cannot be followed observably. See `WORKFLOW.md`, The Run Ledger. - Never cut, at any rung: trust-boundary validation, data-loss handling, security, accessibility. - Lazy about the solution, never about reading the code first. @@ -47,6 +50,11 @@ use judgment — but say so. - A task is well-defined only if it names all four: **files, action, verify, done.** Missing one? The task is too vague — say so. +### Exceptions need decisions +- Every **permanent exception** to a rule requires an ADR. An exception + that is documented but never decided is an error. + (Field-proven rule from the first adoption; see docs/issues/0006.) + ### Verification before completion - Never claim something works without evidence: a test run, command output, a rendered result. "Should work" is not a status. @@ -84,7 +92,9 @@ explicitly grants that exception. | `docs/issues/` | In-repo issues, one file each; status lives in frontmatter | | `docs/wiki/` | Wiki areas as folders, created on demand — rules in `docs/wiki/index.md` | | `docs/sources/` | Immutable original sources; wiki pages cite them — read-only for agents | -| `scripts/` | Deterministic tooling: `validate.py`, `gen_status.py` | +| `docs/ledger/` | One trace per session: a row per gate, plus the ladder walk | +| `docs/verdict/` | Judged runs; findings split into model failure vs framework gap | +| `scripts/` | Deterministic tooling: `validate.py`, `gen_status.py`, the check family | Before proposing options (Gate 2), read the relevant ADRs and AARs first — past decisions and learnings are input, not trivia. @@ -93,6 +103,9 @@ past decisions and learnings are input, not trivia. - All artifacts are standard Markdown with YAML frontmatter conforming to `schema.yaml`. Standard links only (`[text](path.md)`), no wikilinks. + **Link targets are files, never directories** — `validate.py` rejects + `[x](dir/)` even where a forge would render it; point at the + directory's README or a concrete file instead. Diagrams as Mermaid. This keeps every artifact portable across LLMs, GitLab, and Obsidian. - Never invent frontmatter fields or status values. `validate.py` is @@ -101,7 +114,13 @@ past decisions and learnings are input, not trivia. - Deterministic jobs (status generation, validation, link checks) are done by scripts, not by you. If a deterministic job lacks a script, propose one instead of doing it by inference. - +- **Adopting projects:** project-specific rules live in a marked section + *appended below* the upstream content of this file (marker line, e.g. + ``) — never woven into it. Vendor the pristine + upstream originals under `docs/sources/upstream//` and check + the upstream part byte-for-byte against them in CI: a silent rewrite + of the framework files becomes a red pipeline, a framework upgrade + becomes a deliberate baseline-plus-copy commit. ## 6. Gruppenregeln (Projekt axion1337.chat) diff --git a/WORKFLOW.md b/WORKFLOW.md index 3b2fe7d..486cbca 100644 --- a/WORKFLOW.md +++ b/WORKFLOW.md @@ -78,15 +78,32 @@ sections. - Each slice ends with verification evidence, a status (`DONE` | `DONE_WITH_CONCERNS` | `NEEDS_CONTEXT` | `BLOCKED`), and a **STOP** for human review before the next slice. +- **A check owes proof that it can fail.** A slice that introduces a + gate, test or check shows it going red with a deliberate break, and + quotes that run. A gate only ever seen green is a hypothesis, not a + result — and a suite that only calls its own functions proves the + functions, not the program, so break the wiring too. The proof has to + come from **where the check actually runs**: a control that passes on a + workstation says nothing about the environment it was wired into. +- **A delivering artifact owes one real result.** Where the artifact's + output *is* the product — a dashboard, a report, a query — acceptance + quotes one observation it actually returned, or states why it is + legitimately empty. "It exists" is not delivery. ### Gate 5 — Closeout - AAR section in the design doc: planned / actual / why the difference / learnings. - Harvest: learnings useful to future readers go to the wiki (FAQ, Stolpersteine) with source links. A missing or wrong framework - rule becomes a framework issue or update. + rule becomes a framework issue or update — generalized first, per + **Harvesting to the Framework** below. - Good analyses produced along the way may be filed as wiki pages (with citations) instead of dying in chat history. +- **Name what this work made false.** Ask it explicitly — which existing + statement, in which artifact, does this undertaking now contradict? — + and record the answer in the closeout, including when it is "none". + Appending is cheap and feels complete; revising an earlier claim costs + attention, so it only happens when someone asks the question. - Move the design doc to `docs/design/done/`. Run `gen_status.py`. ## Debugging Path @@ -113,6 +130,83 @@ For bugs and incidents, any size: - End every working session by answering: "Which choices did I make that I'm least confident about?" File the answer in the design doc. +## The Run Ledger + +A rule that leaves no trace cannot be checked by anyone — not a reviewer, +not a script, not a later reader. Following it and ignoring it produce the +same repository. **Every rule therefore owes a trace, or it is +decoration.** + +The trace of a run is one file in `docs/ledger/`, one per session, with a +row per gate: + +- **`gate`** — the gate number, or the slice for size M. +- **`commit`** — the commit that closed it. A commit, not a timestamp: a + timestamp is written by the same hand as the claim, a commit is not. +- **`approval`** — who released the gate. Empty means the next gate began + without a stop. +- **`status`** — one of `DONE` | `DONE_WITH_CONCERNS` | `NEEDS_CONTEXT` | + `BLOCKED`. +- **`note`** — one line, for a reader. + +A second table records the **ladder walk** (`AGENTS.md` §1): what was +searched, what was found, and an outcome that begins `reused:` or +`built:`. The section is mandatory even when nothing new was built — a +session that built nothing says so in a row, because an absent record and +an ignored rule look identical. + +`scripts/judge.py` reads the ledger against these rules and against git. +`scripts/judge.py --coverage` reports which rules of this framework are +observable at all; a rule at zero is either unobservable or inert, and +both are worth knowing. That report is never a gate. + +## Judging a Run + +Two judges, and they answer different questions. + +**The deterministic one** is `scripts/judge.py`: gate order, gates owed by +the declared size class, approvals, the status vocabulary, and the +cross-checks against git — that every named commit exists, is reachable, +and that the commits run in the same order as the gates claim. That last +group is the only part measured against evidence the ledger's author did +not write. + +**The inferential one** is a ritual, not a script, because no script can +tell a real design document from a plausible one. It is invoked as a +skill or slash command, and its one binding condition is where it runs: + +> ⚠️ **Fresh context, and without the work transcript.** An agent that +> judges its own session justifies rather than checks. The judge receives +> `AGENTS.md`, `WORKFLOW.md`, the diff and the ledger — and nothing else. +> This is a condition, not an optimisation; run in the same context the +> output is worthless. + +Its rubric, five questions, each answered with a quotation from the diff +or the ledger rather than an impression: + +1. **Scope.** Does every changed file trace to the stated undertaking? + Name any that does not. +2. **Substance.** Is the design document load-bearing — do its non-goals + exclude something a reader would otherwise expect, and are the shakiest + calls real risks rather than modesty? +3. **Size.** Was the declared class plausible for what the diff turned out + to be? A size that fits only in hindsight is the finding. +4. **Evidence.** Where a slice claims a result, is a run quoted? Where a + check was added, was it shown going red? +5. **Silence.** Which rule *should* have produced a trace here and did + not? + +The output is a `verdict` artifact under `docs/verdict/`, and every +finding is classified into one of two buckets: **model failure** — a clear +rule was not followed — or **framework gap** — the rule is missing, +unenforceable, or invisible. Keeping those apart is the whole point; a +finding in the wrong bucket either blames a person for a missing rule or +writes a behavioural lapse into the rule set. Binding rule: +ADR-0010. + +⚠️ The ledger is written by the agent it describes and can be wrong. The +threat model is drift, not sabotage — see the ADR. + ## Refinement Session A recurring, human-triggered ritual. Agenda: @@ -121,12 +215,44 @@ A recurring, human-triggered ritual. Agenda: last time. 2. Backlog triage over `docs/issues/`: close, reprioritize, split. 3. AAR harvest: walk recent AARs; update the wiki (FAQ, Stolpersteine); - propose framework changes. + propose framework changes — generalized per **Harvesting to the + Framework** below. 4. Wiki lint (content-level, beyond `validate.py`): contradictions between pages, claims superseded by newer sources, orphan pages, missing cross-references, gaps worth a new page or a web search. 5. STATUS review: anything stale or surprising in `STATUS.md`. +## Harvesting to the Framework + +A harvest is the one artifact that leaves the project: findings travel +from an adopting repo back into the framework so the next version can +cover them. It crosses a trust boundary, and it is written by the people +least likely to notice what is specific about their own project. + +- **What travels:** the failure class, its effect, its cause, and the + framework change it argues for. +- **What never travels:** names of the adopter, its products, domains or + customers; hostnames, URLs and paths; stack components; people; commit + hashes; the adopter's own issue and ticket identifiers. Refer to + evidence by an identifier that resolves only in the private record, and + use roles instead of names — "a steering repo for a multi-component + group", never the group. +- **Before it leaves:** run `scripts/check_harvest.py` over the range + being handed over. It reads file contents, file names and commit + messages, and fails closed: a missing term list or an unresolvable range + is an error, never a pass. +- **The term list belongs to the adopter** and is never committed to the + framework — the framework must not store the names it exists to keep + out. A term the framework itself uses is shared vocabulary, not a + secret; listing it only produces noise that teaches people to skip the + check. + +⚠️ The check supplements review, it never replaces it: a denylist finds +only the nouns somebody thought of, and it deliberately does not look at +commit identity. Read the diff as well. + +Binding rule: ADR-0008. + ## Knowledge Handling (summary) Full rules live in `docs/wiki/index.md`. The short version: diff --git a/docs/ledger/template.md b/docs/ledger/template.md new file mode 100644 index 0000000..98942bd --- /dev/null +++ b/docs/ledger/template.md @@ -0,0 +1,42 @@ +--- +type: ledger +date: YYYY-MM-DD +size: L # S | M | L +status: open # open | closed (closed at Gate 5) +related: [] # the design doc, the issue, whatever this run served +--- + + + +# Ledger: + +## Gates + + + +| gate | commit | approval | status | note | +|---|---|---|---|---| +| 1 | abc1234 | owner | DONE | one line, for a reader | + +## Ladder + + + +| searched | found | outcome | commit | +|---|---|---|---| +| where you looked | what was there | reused: … / built: … | abc1234 | + +## Notes + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/AGENTS.md b/docs/sources/upstream/neckbeard-v0.3.1/AGENTS.md new file mode 100644 index 0000000..5b56235 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/AGENTS.md @@ -0,0 +1,123 @@ +# AGENTS.md — Canonical Agent Instructions + +Canonical instruction set for any coding agent working in this repository +(Claude Code, GPT-OSS harnesses, others). `CLAUDE.md` points here. +This file is loaded into every session — keep it short. Process details +live in `WORKFLOW.md`; read that when a task begins, not preemptively. + +Tradeoff: these rules bias toward caution over speed. For trivial tasks, +use judgment — but say so. + +## 1. Operating Rules + +### Think before coding +- State your assumptions explicitly. If uncertain, ask. +- If multiple interpretations exist, present them — don't pick silently. +- If a simpler approach exists, say so. Push back when warranted. +- If something is unclear, stop. Name what's confusing. Ask. + +### Simplicity first +- Minimum code that solves the problem. Nothing speculative. +- No features beyond what was asked. No abstractions for single-use code. +- No "flexibility" or "configurability" that wasn't requested. +- No error handling for impossible scenarios. +- If you write 200 lines and it could be 50, rewrite it. +- Test: "Would a senior engineer call this overcomplicated?" If yes, simplify. +- Before writing new code, stop at the first rung that holds: + needed at all? → codebase already has it? → stdlib? → platform-native? + → installed dependency? → one line? → only then: the minimum that works. + (Ladder after ponytail, MIT.) **Name the rung that holds, and record the + walk in the run's ledger** — searched, found, outcome. A rung that held + and a rung never tried leave the same diff, so without the record this + rule cannot be followed observably. See `WORKFLOW.md`, The Run Ledger. +- Never cut, at any rung: trust-boundary validation, data-loss handling, + security, accessibility. +- Lazy about the solution, never about reading the code first. + +### Surgical changes +- Touch only what you must. Match existing style, even if you'd differ. +- Don't "improve" adjacent code, comments, or formatting. +- Don't refactor things that aren't broken. +- If you notice unrelated dead code, mention it — don't delete it. +- Remove imports/variables/functions that YOUR changes made unused; + leave pre-existing dead code alone unless asked. +- Every changed line must trace directly to the request. + +### Goal-driven execution +- Transform tasks into verifiable goals: + "fix the bug" → "write a test that reproduces it, then make it pass". +- For multi-step work, state a brief plan: step → verify, step → verify. +- A task is well-defined only if it names all four: + **files, action, verify, done.** Missing one? The task is too vague — say so. + +### Exceptions need decisions +- Every **permanent exception** to a rule requires an ADR. An exception + that is documented but never decided is an error. + (Field-proven rule from the first adoption; see docs/issues/0006.) + +### Verification before completion +- Never claim something works without evidence: a test run, command + output, a rendered result. "Should work" is not a status. +- Report every task/slice with exactly one status: + `DONE` | `DONE_WITH_CONCERNS` | `NEEDS_CONTEXT` | `BLOCKED`. +- Uncertainty is reported, never swallowed. Flag your shakiest calls. + +## 2. Project Initialization (Gate 0) + +At session start, read `PROJECT.md`. If it does not exist, initialization +is your first task: before anything else, ask the Gate 0 questions defined +in `WORKFLOW.md` — response language, size-S gate exception (yes/no), +one-line project purpose, audience — write the answers to `PROJECT.md`, +and have `validate.py` accept it. Never guess these answers; ask. + +## 3. Workflow + +For anything beyond a trivial change, read `WORKFLOW.md` and follow its +gates. At task start, propose a size class (S/M/L); the human confirms +(possibly batched later). **Never advance past a gate without explicit +human approval** — sole exception: size-S tasks, and only if `PROJECT.md` +explicitly grants that exception. + +## 4. Repository Map + +| Path | Purpose | +|---|---| +| `WORKFLOW.md` | Gate 0 (init) + Gates 1–5, size classes, debugging path, session handoff, refinement ritual | +| `PROJECT.md` | Per-project answers from Gate 0: language, size-S exception, purpose, audience | +| `STATUS.md` | Generated overview: open issues, active designs, recent ADRs — do not edit by hand | +| `schema.yaml` | Frontmatter schema — single source of truth for artifact structure | +| `docs/adr/` | Architecture Decision Records — binding; never edited, only superseded | +| `docs/design/` | One design doc per undertaking; completed ones move to `done/` | +| `docs/aar/` | Standalone After Action Reviews (incidents, major deviations only) | +| `docs/issues/` | In-repo issues, one file each; status lives in frontmatter | +| `docs/wiki/` | Wiki areas as folders, created on demand — rules in `docs/wiki/index.md` | +| `docs/sources/` | Immutable original sources; wiki pages cite them — read-only for agents | +| `docs/ledger/` | One trace per session: a row per gate, plus the ladder walk | +| `docs/verdict/` | Judged runs; findings split into model failure vs framework gap | +| `scripts/` | Deterministic tooling: `validate.py`, `gen_status.py`, the check family | + +Before proposing options (Gate 2), read the relevant ADRs and AARs first — +past decisions and learnings are input, not trivia. + +## 5. Artifact Rules + +- All artifacts are standard Markdown with YAML frontmatter conforming to + `schema.yaml`. Standard links only (`[text](path.md)`), no wikilinks. + **Link targets are files, never directories** — `validate.py` rejects + `[x](dir/)` even where a forge would render it; point at the + directory's README or a concrete file instead. + Diagrams as Mermaid. This keeps every artifact portable across LLMs, + GitLab, and Obsidian. +- Never invent frontmatter fields or status values. `validate.py` is + authoritative; if it rejects your artifact, fix the artifact, not the + validator. +- Deterministic jobs (status generation, validation, link checks) are done + by scripts, not by you. If a deterministic job lacks a script, propose + one instead of doing it by inference. +- **Adopting projects:** project-specific rules live in a marked section + *appended below* the upstream content of this file (marker line, e.g. + ``) — never woven into it. Vendor the pristine + upstream originals under `docs/sources/upstream//` and check + the upstream part byte-for-byte against them in CI: a silent rewrite + of the framework files becomes a red pipeline, a framework upgrade + becomes a deliberate baseline-plus-copy commit. diff --git a/docs/sources/upstream/neckbeard-v0.3.1/CLAUDE.md b/docs/sources/upstream/neckbeard-v0.3.1/CLAUDE.md new file mode 100644 index 0000000..9aef12f --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/CLAUDE.md @@ -0,0 +1 @@ +Read AGENTS.md — the canonical instruction file for this repository. All rules live there. diff --git a/docs/sources/upstream/neckbeard-v0.3.1/HERKUNFT.md b/docs/sources/upstream/neckbeard-v0.3.1/HERKUNFT.md new file mode 100644 index 0000000..cde70bd --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/HERKUNFT.md @@ -0,0 +1,55 @@ +# Herkunft dieser Baseline + +Unveränderte Originale aus **neckbeard v0.3.1**, Commit +`82795b667a69761134717b00da762e75e1bd6c60` (`main`, sauber), Origin +`https://git.lab/oss-projekte/ai/neckbeard.git`. Löst die Baseline +v0.1.1 ab, die drei Minor-Stände alt war. + +Zweck: Byte-Baseline für `scripts/pruefe_upstream_drift.py`. Ein +Framework-Upgrade ersetzt diese Dateien bewusst und in einem eigenen +Commit — nie beiläufig. + +| Datei hier | Arbeitskopie | Prüfung | +|---|---|---| +| `AGENTS.md` | `/AGENTS.md` | Präfix bis zur Marke `` | +| `CLAUDE.md` | `/CLAUDE.md` | byte-identisch | +| `WORKFLOW.md` | `/WORKFLOW.md` | byte-identisch | +| `templates/adr-template.md` | `docs/adr/template.md` | byte-identisch | +| `templates/design-template.md` | `docs/design/template.md` | byte-identisch | +| `templates/aar-template.md` | `docs/aar/template.md` | byte-identisch | +| `templates/issue-template.md` | `docs/issues/template.md` | byte-identisch | +| `templates/ledger-template.md` | `docs/ledger/template.md` | byte-identisch **(neu)** | +| `templates/verdict-template.md` | `docs/verdict/template.md` | byte-identisch **(neu)** | +| `scripts/check_harvest.py` | `scripts/check_harvest.py` | byte-identisch **(neu übernommen)** | +| `scripts/judge.py` | `scripts/judge.py` | byte-identisch **(neu übernommen)** | +| `schema.yaml` | `/schema.yaml` | **erklärt projekterweitert** — nur Diff-Referenz | +| `scripts/validate.py` | `scripts/validate.py` | **erklärt projekterweitert** — nur Diff-Referenz | +| `scripts/gen_status.py` | `scripts/gen_status.py` | **erklärt projekterweitert** — nur Diff-Referenz | +| `scripts/check_locked.py` | — | **bewusst nicht übernommen**, siehe unten | + +## Was dieses Upgrade an den erweiterten Dateien nachgezogen hat + +Die drei erklärt-erweiterten Dateien werden nicht byte-verglichen, also +meldet auch nichts, wenn upstream sie ändert. Nachgemessen für v0.1.1 → +v0.3.1 und von Hand übernommen: + +- `scripts/validate.py`: upstream ergänzt `check_vendored_portable` samt + Aufruf. Übernommen; der Rest des Forks (`waiting_requires_reason`, + `slug_matches_filename`) ist unberührt. +- `schema.yaml`: upstream ergänzt `vendored`, `project_section_marker`, + das Feld `judged` und die Typen `ledger` und `verdict`. Übernommen, + Marke auf unsere gesetzt. +- `scripts/gen_status.py`: upstream unverändert. Nichts zu tun. + +⚠️ **Das ist heute ein Handgriff ohne Widerspruch.** Beim nächsten Upgrade +erinnert nichts daran, diesen Vergleich wieder zu fahren. Das ist die +verbleibende Lücke dieser Konstruktion und gehört als Befund zurück ins +Rahmenwerk. + +## check_locked.py: bewusst nicht übernommen + +Das Rahmenwerk liefert seit v0.2.0 `check_locked.py` — dieselbe Aufgabe, +die hier `scripts/pruefe_sperrliste.py` seit dem 2026-08-20 erfüllt und +die bei jedem Push in der CI hängt. Erste Sprosse der Leiter: braucht es +das überhaupt? Nein. Zwei Werkzeuge für eine Regel wären Pflegeaufwand +ohne Gewinn. diff --git a/docs/sources/upstream/neckbeard-v0.3.1/WORKFLOW.md b/docs/sources/upstream/neckbeard-v0.3.1/WORKFLOW.md new file mode 100644 index 0000000..486cbca --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/WORKFLOW.md @@ -0,0 +1,266 @@ +# WORKFLOW.md — Gates, Sizing, and Rituals + +Read this when a task begins, not preemptively. `AGENTS.md` holds the +always-on rules; this file holds the process. + +## Size Classes + +Propose one at task start; the human confirms — individually, or batched +at the next refinement session. + +| Class | Scope | Process | +|---|---|---| +| S | One file / one small change, no design decisions | Direct. AGENTS.md rules only. The one-line go-ahead **before starting is the stop** — waived only if `PROJECT.md` grants the size-S exception. | +| M | Few files, minor decisions, fits one session | Slice plan in chat, no file. **STOP: plan approval before any code.** Then implement; each slice reports evidence and status inline. Gate 5 is a short AAR note in chat, filed to the wiki only if it produced a real learning. | +| L | New feature, multiple files or sessions, real decisions | Full design doc in `docs/design/` following Gates 1–5 below. | + +When in doubt between two classes, pick the larger. + +## Gate 0 — Project Initialization + +Runs once per project, triggered by a missing `PROJECT.md`. Ask, never guess: + +1. Response language? (e.g. de / en) +2. Size-S gate exception granted? (yes / no) +3. One-line project purpose? +4. Audience — who uses this besides the owner? (Drives which wiki areas + become mandatory later; see `docs/wiki/index.md`.) + +Write the answers to `PROJECT.md` (frontmatter per `schema.yaml`), run +`validate.py`, and confirm the result with the human. + +## Gates 1–5 (size L) + +Each gate is a section of the design doc. A gate ends with **STOP**: +present the section, wait for explicit approval. Do not pre-fill later +sections. + +### Gate 1 — Product +- Problem statement: what user problem, for whom. +- Verifiable acceptance criterion. A real number where one exists; + otherwise a concretely checkable outcome. "Works" is not a criterion. +- Non-goals: what this deliberately does not do. +- Announcement paragraph (3–5 sentences): what it is, who it's for, why + it's good. If you can't write it, the product isn't understood yet. +- UI involved? Plain-HTML mockups of the affected screens. + +**STOP.** + +### Gate 2 — Architecture +- Read first: the actual codebase, relevant ADRs, relevant AARs. + Past decisions and learnings are input, not trivia. +- How it fits the real system: endpoints, tables/schemas, query + outlines, the end-to-end flow (Mermaid). +- Constraints: non-functional requirements, proportional to the project. +- Options & trade-offs where more than one viable way exists: pro/contra + each, chosen option, and why. Feature-local decisions stay here. +- Lasting directional decisions discovered here become ADRs (one each), + linked from the design doc. + +**STOP.** + +### Gate 3 — Program Design +- File locations: exact paths, new and touched. +- Types and method signatures — no bodies. +- Call stack for the main flow(s). +- What the tests will assert. +- Boundaries: an explicit DO NOT CHANGE list. +- Shakiest calls: name the decisions you are least confident about. + +**STOP.** + +### Gate 4 — Vertical Slices +- Slice 1 is the tracer bullet: a thin end-to-end path that runs + (mocks and stubs allowed). Only then real logic, one testable slice + at a time. Never build layer-by-layer horizontally. +- Every slice lists its tasks; every task names **files, action, + verify, done**. +- Each slice ends with verification evidence, a status + (`DONE` | `DONE_WITH_CONCERNS` | `NEEDS_CONTEXT` | `BLOCKED`), + and a **STOP** for human review before the next slice. +- **A check owes proof that it can fail.** A slice that introduces a + gate, test or check shows it going red with a deliberate break, and + quotes that run. A gate only ever seen green is a hypothesis, not a + result — and a suite that only calls its own functions proves the + functions, not the program, so break the wiring too. The proof has to + come from **where the check actually runs**: a control that passes on a + workstation says nothing about the environment it was wired into. +- **A delivering artifact owes one real result.** Where the artifact's + output *is* the product — a dashboard, a report, a query — acceptance + quotes one observation it actually returned, or states why it is + legitimately empty. "It exists" is not delivery. + +### Gate 5 — Closeout +- AAR section in the design doc: planned / actual / why the + difference / learnings. +- Harvest: learnings useful to future readers go to the wiki + (FAQ, Stolpersteine) with source links. A missing or wrong framework + rule becomes a framework issue or update — generalized first, per + **Harvesting to the Framework** below. +- Good analyses produced along the way may be filed as wiki pages + (with citations) instead of dying in chat history. +- **Name what this work made false.** Ask it explicitly — which existing + statement, in which artifact, does this undertaking now contradict? — + and record the answer in the closeout, including when it is "none". + Appending is cheap and feels complete; revising an earlier claim costs + attention, so it only happens when someone asks the question. +- Move the design doc to `docs/design/done/`. Run `gen_status.py`. + +## Debugging Path + +For bugs and incidents, any size: + +1. Reproduce first. No reproduction, no fix. +2. Hypothesize the root cause; verify the hypothesis with evidence + before changing anything. +3. Route the failure before fixing (diagnostic failure routing): + - **Intent issue** — we built toward the wrong goal → back to Gate 1. + - **Spec issue** — the design/plan was wrong → fix the spec + (Gate 2/3), then the code. + - **Code issue** — plan right, code wrong → fix in place. +4. Fix, plus a test that would have caught it. +5. Incidents and major misdiagnoses get a standalone AAR in `docs/aar/`. + +## Session Handoff + +- When a slice completes, or context quality degrades, write the current + state into the design doc's **Handoff block** — done slices, open + decisions, next step — then start a fresh session that resumes from + the doc. The doc is the memory; the session is disposable. +- End every working session by answering: "Which choices did I make that + I'm least confident about?" File the answer in the design doc. + +## The Run Ledger + +A rule that leaves no trace cannot be checked by anyone — not a reviewer, +not a script, not a later reader. Following it and ignoring it produce the +same repository. **Every rule therefore owes a trace, or it is +decoration.** + +The trace of a run is one file in `docs/ledger/`, one per session, with a +row per gate: + +- **`gate`** — the gate number, or the slice for size M. +- **`commit`** — the commit that closed it. A commit, not a timestamp: a + timestamp is written by the same hand as the claim, a commit is not. +- **`approval`** — who released the gate. Empty means the next gate began + without a stop. +- **`status`** — one of `DONE` | `DONE_WITH_CONCERNS` | `NEEDS_CONTEXT` | + `BLOCKED`. +- **`note`** — one line, for a reader. + +A second table records the **ladder walk** (`AGENTS.md` §1): what was +searched, what was found, and an outcome that begins `reused:` or +`built:`. The section is mandatory even when nothing new was built — a +session that built nothing says so in a row, because an absent record and +an ignored rule look identical. + +`scripts/judge.py` reads the ledger against these rules and against git. +`scripts/judge.py --coverage` reports which rules of this framework are +observable at all; a rule at zero is either unobservable or inert, and +both are worth knowing. That report is never a gate. + +## Judging a Run + +Two judges, and they answer different questions. + +**The deterministic one** is `scripts/judge.py`: gate order, gates owed by +the declared size class, approvals, the status vocabulary, and the +cross-checks against git — that every named commit exists, is reachable, +and that the commits run in the same order as the gates claim. That last +group is the only part measured against evidence the ledger's author did +not write. + +**The inferential one** is a ritual, not a script, because no script can +tell a real design document from a plausible one. It is invoked as a +skill or slash command, and its one binding condition is where it runs: + +> ⚠️ **Fresh context, and without the work transcript.** An agent that +> judges its own session justifies rather than checks. The judge receives +> `AGENTS.md`, `WORKFLOW.md`, the diff and the ledger — and nothing else. +> This is a condition, not an optimisation; run in the same context the +> output is worthless. + +Its rubric, five questions, each answered with a quotation from the diff +or the ledger rather than an impression: + +1. **Scope.** Does every changed file trace to the stated undertaking? + Name any that does not. +2. **Substance.** Is the design document load-bearing — do its non-goals + exclude something a reader would otherwise expect, and are the shakiest + calls real risks rather than modesty? +3. **Size.** Was the declared class plausible for what the diff turned out + to be? A size that fits only in hindsight is the finding. +4. **Evidence.** Where a slice claims a result, is a run quoted? Where a + check was added, was it shown going red? +5. **Silence.** Which rule *should* have produced a trace here and did + not? + +The output is a `verdict` artifact under `docs/verdict/`, and every +finding is classified into one of two buckets: **model failure** — a clear +rule was not followed — or **framework gap** — the rule is missing, +unenforceable, or invisible. Keeping those apart is the whole point; a +finding in the wrong bucket either blames a person for a missing rule or +writes a behavioural lapse into the rule set. Binding rule: +ADR-0010. + +⚠️ The ledger is written by the agent it describes and can be wrong. The +threat model is drift, not sabotage — see the ADR. + +## Refinement Session + +A recurring, human-triggered ritual. Agenda: + +1. Batched confirmations: size classes and small approvals queued since + last time. +2. Backlog triage over `docs/issues/`: close, reprioritize, split. +3. AAR harvest: walk recent AARs; update the wiki (FAQ, Stolpersteine); + propose framework changes — generalized per **Harvesting to the + Framework** below. +4. Wiki lint (content-level, beyond `validate.py`): contradictions + between pages, claims superseded by newer sources, orphan pages, + missing cross-references, gaps worth a new page or a web search. +5. STATUS review: anything stale or surprising in `STATUS.md`. + +## Harvesting to the Framework + +A harvest is the one artifact that leaves the project: findings travel +from an adopting repo back into the framework so the next version can +cover them. It crosses a trust boundary, and it is written by the people +least likely to notice what is specific about their own project. + +- **What travels:** the failure class, its effect, its cause, and the + framework change it argues for. +- **What never travels:** names of the adopter, its products, domains or + customers; hostnames, URLs and paths; stack components; people; commit + hashes; the adopter's own issue and ticket identifiers. Refer to + evidence by an identifier that resolves only in the private record, and + use roles instead of names — "a steering repo for a multi-component + group", never the group. +- **Before it leaves:** run `scripts/check_harvest.py` over the range + being handed over. It reads file contents, file names and commit + messages, and fails closed: a missing term list or an unresolvable range + is an error, never a pass. +- **The term list belongs to the adopter** and is never committed to the + framework — the framework must not store the names it exists to keep + out. A term the framework itself uses is shared vocabulary, not a + secret; listing it only produces noise that teaches people to skip the + check. + +⚠️ The check supplements review, it never replaces it: a denylist finds +only the nouns somebody thought of, and it deliberately does not look at +commit identity. Read the diff as well. + +Binding rule: ADR-0008. + +## Knowledge Handling (summary) + +Full rules live in `docs/wiki/index.md`. The short version: + +- Original sources live in `docs/sources/`, immutable — agents read + them, never modify them. Wiki pages cite the sources they draw on. +- Contradictions are resolved or explicitly flagged — never left + silently coexisting. +- If the wiki has no confident answer, say so. Never file a + low-confidence synthesis back as knowledge. +- Git is the changelog. No separate log file. diff --git a/docs/sources/upstream/neckbeard-v0.3.1/schema.yaml b/docs/sources/upstream/neckbeard-v0.3.1/schema.yaml new file mode 100644 index 0000000..d97f197 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/schema.yaml @@ -0,0 +1,154 @@ +# schema.yaml — single source of truth for artifact frontmatter. +# Stage 1 of ADR-0004: scripts/validate.py checks generically against this +# file. Extending the framework's metadata means editing THIS file, not code. +# Agents: never invent fields or status values; propose a schema change. + +version: 1 + +scope: + # Files considered artifacts. Templates and raw sources are exempt. + include: + - "PROJECT.md" + - "docs/**/*.md" + exclude: + - "**/template.md" + - "docs/sources/**" + - "vendor/**" + # Files whose inline links are checked, but which need no frontmatter + # (root-level prose: README, AGENTS, WORKFLOW, generated STATUS, ...). + link_only: + - "*.md" + +# Files an adopting project holds byte-identical against its vendored +# baseline (AGENTS.md §5). They are copied into repositories that do not +# have this repo's docs/, so they must carry no repo-relative link — a +# link that resolves here and nowhere else makes the adoption path +# unfollowable. Mention an ADR by its identifier instead. +vendored: + - "AGENTS.md" + - "WORKFLOW.md" + - "CLAUDE.md" + +# An adopting project appends its own always-on rules below a marker in a +# vendored file (AGENTS.md §5). Everything from the marker on belongs to +# that project, is never copied anywhere, and may link freely — the rule +# above applies only to the upstream part. Adopters set their own marker. +project_section_marker: "" + +# Frontmatter fields whose values are links. Values starting with +# http://, https:// or mailto: are treated as external and only +# format-checked; everything else must be a repo-root-relative path +# to an existing file. +link_fields: [related, sources, supersedes, superseded_by, judged] + +types: + project: + dir: "." + filename: "^PROJECT\\.md$" + required: [type, language, size_s_exception, purpose, audience] + fields: + language: { enum: [de, en] } + size_s_exception: { kind: bool } + purpose: { kind: str } + audience: { kind: str } + + adr: + dir: "docs/adr" + filename: "^\\d{4}-[a-z0-9-]+\\.md$" + required: [type, id, status, date] + fields: + id: { pattern: "^\\d{4}$" } + status: { enum: [proposed, accepted, superseded] } + date: { kind: date } + supersedes: { kind: link, nullable: true } + superseded_by: { kind: link, nullable: true } + related: { kind: links } + rules: + # status: superseded requires superseded_by to point at the successor. + - superseded_requires_pointer + + design: + dir: "docs/design" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, status, date, size] + fields: + status: { enum: [gate-1, gate-2, gate-3, gate-4, gate-5, done] } + size: { enum: [L] } + date: { kind: date } + related: { kind: links } + rules: + # status: done if and only if the file lives under docs/design/done/. + - done_iff_in_done_dir + + aar: + dir: "docs/aar" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, status, date] + fields: + status: { enum: [open, harvested] } + date: { kind: date } + related: { kind: links } + + issue: + dir: "docs/issues" + filename: "^\\d{4}-[a-z0-9-]+\\.md$" + required: [type, id, status, created] + fields: + id: { pattern: "^\\d{4}$" } + status: { enum: [open, in-progress, done, rejected] } + created: { kind: date } + related: { kind: links } + + # One per session. The envelope is validated here; the gate rows and + # ladder entries in the body are outside what this engine can express + # (it has no notion of a list of records) and belong to scripts/judge.py. + # That seam is deliberate — see ADR-0010. + ledger: + dir: "docs/ledger" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, date, size, status] + fields: + date: { kind: date } + size: { enum: [S, M, L] } + status: { enum: [open, closed] } + related: { kind: links } + + # The output of a judged run. Categories are the two-bucket + # classification the harvest assessment established: a finding is either + # the model not following a clear rule, or a gap in the framework. + verdict: + dir: "docs/verdict" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, date, outcome, judged] + fields: + date: { kind: date } + outcome: { enum: [clean, model-failure, framework-gap, both] } + related: { kind: links } + + wiki-page: + dir: "docs/wiki" + filename: "^[a-z0-9-]+\\.md$" + required: [type, area] + fields: + area: + enum: + - index + - architecture + - admin + - deployment + - user-guide + - requirements + - faq + - stolpersteine + # Optional, and meaningful on a page that records a recurring failure + # class (area: stolpersteine). `harvested` means a released framework + # version covers the pattern — handed over is not harvested, and + # `harvested_in` names that version (ADR-0009). + status: { enum: [open, partly, harvested] } + harvested_in: { kind: str } + sources: { kind: links } + related: { kind: links } + rules: + # Pages other than the index should be linked from somewhere + # (reported as WARNING, not error — see validate.py). + - warn_if_orphan diff --git a/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_harvest.py b/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_harvest.py new file mode 100644 index 0000000..ac6b94e --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_harvest.py @@ -0,0 +1,461 @@ +#!/usr/bin/env python3 +"""check_harvest.py — refuse a harvest that carries adopter specifics. + +A harvest travels from an adopting project into the framework repository. +Only generalized statements may cross: the failure class, its effect, its +cause, and the framework change they argue for. Names of the adopter, its +products, hosts, stack components, people, paths and identifiers must not. + +Three surfaces are checked, because a real leak used all of them: + * file contents + * file names + * commit messages in the range being handed over + +Fails closed (exit 1, never "clean"): + * term list missing, unreadable, or empty after stripping comments + * a revision range git cannot resolve + * a file that cannot be read as text — a harvest is prose, and a blob + nobody can read is precisely what nobody can review + +Deliberately NOT a fourth surface: the author and committer identity of +those commits. It is a repository-wide property rather than something a +harvest carries in — the framework's own history and LICENSE hold the same +name — so a run over any branch would report it every time. A check that +is permanently red reports nothing, and teaches people to skip it. Commit +identity belongs to the repository's publication decision, not to this +check; verify it there, once, and not in every harvest. + +⚠️ This check supplements human review, it does not replace it. A denylist +finds only the nouns somebody thought of, and the paragraph above names a +surface it does not look at by design. Treating a green run as proof of +absence is the same mistake that made the leak it exists to prevent. + +The term list belongs to the adopter and is never shipped with the +framework — the framework must not store the names it exists to keep out. +Two rules for writing one: + * one term per line, `#` starts a comment. Matching is case-insensitive + and respects word boundaries, so `zephyr` finds "Zephyr-Web" and + "zephyr.example" but not the middle of an unrelated English word. + Wrap a term in asterisks — `*zephyr*` — to match inside words too, + for the rare name that hides in a compound with no separator. + * a term the framework itself uses is shared vocabulary, not a secret. + Listing it only produces noise that teaches people to ignore the check. + +Scope. With `--range`, the lines that range **adds** are read, plus the +names of the files it touches and the messages of its commits. A harvest +answers for what it writes, not for what the repository already carried: +regenerating a shared index would otherwise drag every pre-existing line +of that index into the result. Without `--range`, every tracked file is +read in full — an audit of the repository, which is a different question. + +Findings print the term and a masked excerpt: enough to locate the leak, +without repeating the secret in full wherever the output ends up. + +Usage: + python scripts/check_harvest.py --terms [--range ..] [path ...] + python scripts/check_harvest.py --selftest +""" +from __future__ import annotations + +import io +import os +import re +import subprocess +import sys +import tempfile +from contextlib import redirect_stdout +from functools import lru_cache +from pathlib import Path +from typing import Iterator, NamedTuple + +MASK = "***" +COMMIT_SEP = "\x1e" +FIELD_SEP = "\x1f" + + +class HarvestError(RuntimeError): + """The check could not answer the question. That is never a pass.""" + + +class Finding(NamedTuple): + surface: str # "content" | "filename" | "commit" | "unreadable" + where: str # path, or commit sha + line: int | None # None for filenames, commit subjects, unreadable + term: str + excerpt: str + + +def git(root: Path, *args: str) -> str: + """Run git. A failure raises — it never becomes an empty result. + + ⚠️ This is the core of failing closed. An unresolvable range makes git + exit non-zero; turning that into "" would turn it into "no changes" + and therefore into a silent pass. + """ + done = subprocess.run(["git", "-C", str(root), *args], + capture_output=True, text=True) + if done.returncode != 0: + raise HarvestError(f"git {' '.join(args)}: {done.stderr.strip()[:160]}") + return done.stdout + + +def load_terms(path: Path) -> list[str]: + """Read the adopter's term list. Missing, unreadable or empty is an error.""" + try: + raw = path.read_text(encoding="utf-8") + except OSError as err: + raise HarvestError(f"term list unreadable: {path} ({err})") from err + terms = [z.strip() for z in raw.splitlines()] + terms = [z for z in terms if z and not z.startswith("#")] + if not terms: + raise HarvestError(f"term list is empty: {path}") + return terms + + +@lru_cache(maxsize=None) +def term_pattern(term: str) -> re.Pattern[str]: + """`*x*` matches inside words; a bare term respects word boundaries. + + ⚠️ Word boundaries are the default because a substring match on a + proper noun hits ordinary language: short personal names sit inside + perfectly ordinary English words — "rene" inside "serene", and the + name that forced this inside the word "authored". Measured, not + hypothesised: a substring run flagged every commit trailer here. + """ + if len(term) > 2 and term.startswith("*") and term.endswith("*"): + return re.compile(re.escape(term[1:-1]), re.I) + return re.compile(rf"(? str: + """Replace the matched term, keeping the surrounding context readable.""" + return term_pattern(term).sub(MASK, text).strip()[:120] + + +def scan_text(text: str, terms: list[str], *, surface: str, + where: str, numbered: bool = True) -> list[Finding]: + found: list[Finding] = [] + for nr, zeile in enumerate(text.splitlines() or [text], start=1): + for term in terms: + if term_pattern(term).search(zeile): + found.append(Finding(surface, where, nr if numbered else None, + term, mask(zeile, term))) + return found + + +def iter_files(root: Path, paths: list[str], + rev_range: str | None = None) -> Iterator[Path]: + """Scope: the range's own files, or every tracked file when none given. + + A harvest is checked against what it adds. Scanning the whole repository + instead surfaces pre-existing content that the harvest never touched — + noise that buries the findings that matter. + """ + if not paths: + if rev_range: + roh = git(root, "diff", "--name-only", "--diff-filter=d", rev_range) + else: + roh = git(root, "ls-files").replace("\0", "\n") + for name in roh.splitlines(): + if name and (root / name).is_file(): + yield root / name + return + for roh in paths: + p = Path(roh) + if p.is_dir(): + yield from (q for q in sorted(p.rglob("*")) if q.is_file()) + elif p.is_file(): + yield p + else: + raise HarvestError(f"path does not exist: {p}") + + +def _kurzname(root: Path, datei: Path) -> str: + """Path for the report — never a crash. + + ⚠️ A file handed in explicitly may sit outside the repository, and + `Path.relative_to` raises for those. A check that dies on a path it was + asked to read answers nothing; the reason it is not fatal is that the + only thing wanted here is a label. + """ + try: + return str(datei.relative_to(root)) + except ValueError: + return str(datei) + + +def scan_files(root: Path, paths: list[str], terms: list[str], + rev_range: str | None = None) -> list[Finding]: + found: list[Finding] = [] + for datei in iter_files(root, paths, rev_range): + rel = _kurzname(root, datei) + found += scan_text(rel, terms, surface="filename", where=rel, + numbered=False) + try: + inhalt = datei.read_text(encoding="utf-8") + except (UnicodeDecodeError, OSError): + found.append(Finding("unreadable", rel, None, "-", + "not readable as text — cannot be reviewed")) + continue + found += scan_text(inhalt, terms, surface="content", where=rel) + return found + + +def scan_diff(root: Path, rev_range: str, terms: list[str]) -> list[Finding]: + """Read the lines a range *adds*, not the files it happens to touch. + + ⚠️ Reading whole files reports content the harvest never wrote — a + generated index regenerated by the harvest carries every pre-existing + line of the repository into the result. Those belong to the repository + and its own publication decision; a harvest answers for what it adds. + """ + roh = git(root, "diff", "--unified=0", "--no-color", rev_range) + found: list[Finding] = [] + datei, nr = "", 0 + for z in roh.splitlines(): + if z.startswith("+++ "): + ziel = z[4:].strip() + datei = ziel[2:] if ziel.startswith("b/") else ziel + elif z.startswith("@@"): + treffer = re.search(r"\+(\d+)", z) + nr = int(treffer.group(1)) if treffer else 0 + elif z.startswith("Binary files") and datei and datei != "/dev/null": + found.append(Finding("unreadable", datei, None, "-", + "binary — cannot be reviewed as text")) + elif z.startswith("+") and not z.startswith("+++"): + if datei and datei != "/dev/null": + inhalt = z[1:] + for term in terms: + if term_pattern(term).search(inhalt): + found.append(Finding("content", datei, nr, term, + mask(inhalt, term))) + nr += 1 + for name in git(root, "diff", "--name-only", "--diff-filter=d", + rev_range).splitlines(): + if name: + found += scan_text(name, terms, surface="filename", where=name, + numbered=False) + return found + + +def scan_commits(root: Path, rev_range: str, terms: list[str]) -> list[Finding]: + roh = git(root, "log", f"--format=%H{FIELD_SEP}%B{COMMIT_SEP}", rev_range) + found: list[Finding] = [] + for block in roh.split(COMMIT_SEP): + block = block.strip("\n") + if FIELD_SEP not in block: + continue + sha, nachricht = block.split(FIELD_SEP, 1) + found += scan_text(nachricht, terms, surface="commit", + where=sha[:12], numbered=False) + return found + + +def report(findings: list[Finding]) -> None: + for surface in ("content", "filename", "commit", "unreadable"): + teil = [f for f in findings if f.surface == surface] + if not teil: + continue + print(f"\n {surface} ({len(teil)}):") + for f in teil: + ort = f"{f.where}:{f.line}" if f.line else f.where + print(f" {ort} — term {f.term!r}") + print(f" {f.excerpt}") + + +def selftest() -> int: + """Positive and negative controls. A check only ever seen green is a guess.""" + fehler: list[str] = [] + + def pruefe(name: str, bedingung: bool) -> None: + print(f" {'ok ' if bedingung else 'FAIL'} {name}") + if not bedingung: + fehler.append(name) + + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + git(root, "init", "-q", "-b", "main") + git(root, "config", "user.email", "selftest@invalid") + git(root, "config", "user.name", "selftest") + liste = root / "terms.txt" + liste.write_text("# comment\nzephyr\nrene\n*brand*\n\n", encoding="utf-8") + terms = load_terms(liste) + + (root / "clean.md").write_text("a generalized statement\n", encoding="utf-8") + (root / "body.md").write_text("one\ntwo Zephyr three\n", encoding="utf-8") + (root / "zephyr-notes.md").write_text("clean body\n", encoding="utf-8") + # "serene" contains the name "rene" — the false-positive class that + # forced word boundaries. "xbrandy" is the opposite case, opted + # into with *…*. + (root / "english.md").write_text("a serene afternoon\n", encoding="utf-8") + (root / "touched.md").write_text("old line says Zephyr\n", encoding="utf-8") + (root / "compound.md").write_text("xbrandy\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-q", "-m", "initial, no secret here") + (root / "clean.md").write_text("still generalized\n", encoding="utf-8") + (root / "late.md").write_text("added later, says Zephyr\n", encoding="utf-8") + # Touched by the range, but the term sits in a line the range did + # not write — the generated-index case, in miniature. + # Line 1 is pre-existing and must stay silent; line 2 is added by + # the range and must be reported *at line 2* — which also proves the + # hunk header is read instead of assumed. + (root / "touched.md").write_text( + "old line says Zephyr\nadded line also says Zephyr\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-q", "-m", "mentions Zephyr in the message") + + treffer = scan_files(root, [], terms) + inhalt = [f for f in treffer if f.surface == "content"] + namen = [f for f in treffer if f.surface == "filename"] + commits = scan_commits(root, "HEAD~1..HEAD", terms) + + pruefe("1 term in a file body is found, with its line number", + any(f.where == "body.md" and f.line == 2 for f in inhalt)) + pruefe("2 term in a filename is found while the body is clean", + any(f.where == "zephyr-notes.md" for f in namen) + and not any(f.where == "zephyr-notes.md" for f in inhalt)) + pruefe("3 term in a commit message is found while the tree is clean", + len(commits) == 1) + # The listed term is lower case; the file says "Zephyr". The match + # is only proof of case-insensitivity if the source really differs. + pruefe("4 a lower-case term matches a capitalized occurrence", + "Zephyr" in (root / "body.md").read_text(encoding="utf-8") + and any(f.term == "zephyr" and f.where == "body.md" + for f in inhalt)) + pruefe("5 the excerpt masks the term instead of repeating it", + all(MASK in f.excerpt for f in inhalt)) + + leer = root / "empty.txt" + leer.write_text("# only comments\n", encoding="utf-8") + pruefe("6 an empty term list raises instead of passing", + _raises(lambda: load_terms(leer))) + pruefe("7 a missing term list raises instead of passing", + _raises(lambda: load_terms(root / "nope.txt"))) + pruefe("8 an unresolvable range raises instead of reporting nothing", + _raises(lambda: scan_commits(root, "nosuchref..HEAD", terms))) + pruefe("9 clean input against a non-empty list finds nothing", + scan_files(root, [str(root / "clean.md")], terms) == []) + pruefe("10 a bare term does not match inside an unrelated word", + not any(f.where == "english.md" for f in treffer)) + pruefe("11 a *term* does match inside a word, when opted into", + any(f.where == "compound.md" and f.term == "*brand*" + for f in treffer)) + # Sharp in both directions: the range's own file must be found, and + # the older ones must not. An assertion that only checks for "nothing" + # would also pass if the scoping read no files at all. + bereich = scan_diff(root, "HEAD~1..HEAD", terms) + pruefe("12 with a range, exactly that range's own additions are read", + {f.where for f in bereich} == {"late.md", "touched.md"} + and any(f.where == "body.md" for f in treffer)) + beruehrt = [f for f in bereich if f.where == "touched.md"] + pruefe("13 a pre-existing line in a touched file is not reported", + len(beruehrt) == 1) + pruefe("14 an added line is reported at its real line number", + bool(beruehrt) and beruehrt[0].line == 2) + + # Everything above tests the scanners directly, which leaves the + # dispatch in main() unproven: it could call the whole-file reader, + # or skip commit messages entirely, and every assertion above would + # still pass. Two discriminators that only the wiring can satisfy: + # whole-file mode reports touched.md at line 1 (the pre-existing + # line) where the diff reader reports line 2, and a missing + # scan_commits call removes the commit section from the output. + zuvor = Path.cwd() + puffer = io.StringIO() + try: + os.chdir(root) + with redirect_stdout(puffer): + main(["--terms", str(liste), "--range", "HEAD~1..HEAD"]) + finally: + os.chdir(zuvor) + ausgabe = puffer.getvalue() + pruefe("15 main() routes a range to the diff reader, not whole files", + "touched.md:2" in ausgabe and "touched.md:1" not in ausgabe) + pruefe("16 main() actually reads the commit messages of the range", + "commit (" in ausgabe) + + # A path handed in explicitly may sit outside the repository. That + # crashed with a ValueError until a positive control tried it. + aussen = Path(tempfile.gettempdir()) / "check_harvest_outside.md" + aussen.write_text("mentions Zephyr\n", encoding="utf-8") + try: + draussen = scan_files(root, [str(aussen)], terms) + except ValueError: + draussen = [] + finally: + aussen.unlink(missing_ok=True) + pruefe("17 a path outside the repository is reported, not fatal", + len(draussen) == 1 and draussen[0].term == "zephyr") + + print() + if fehler: + print(f"selftest: {len(fehler)} assertion(s) failed") + return 1 + print("selftest: 17 assertions passed") + return 0 + + +def _raises(fn) -> bool: + try: + fn() + except HarvestError: + return True + return False + + +def main(argv: list[str]) -> int: + if "--selftest" in argv: + return selftest() + + terms_pfad: Path | None = None + rev_range: str | None = None + paths: list[str] = [] + i = 0 + while i < len(argv): + if argv[i] == "--terms" and i + 1 < len(argv): + terms_pfad, i = Path(argv[i + 1]), i + 2 + elif argv[i] == "--range" and i + 1 < len(argv): + rev_range, i = argv[i + 1], i + 2 + elif argv[i].startswith("--"): + print(f"unknown option: {argv[i]}") + return 2 + else: + paths.append(argv[i]) + i += 1 + + if terms_pfad is None: + print(__doc__.strip().splitlines()[-3].strip()) + print("check_harvest: --terms is required. Unchecked is not clean.") + return 2 + + root = Path.cwd() + try: + terms = load_terms(terms_pfad) + if rev_range and not paths: + findings = scan_diff(root, rev_range, terms) + findings += scan_commits(root, rev_range, terms) + else: + findings = scan_files(root, paths, terms) + if rev_range: + findings += scan_commits(root, rev_range, terms) + except HarvestError as err: + print(f"check_harvest: {err}") + print("The question could not be answered — that is not a pass.") + return 1 + + umfang = f"{len(terms)} term(s)" + if rev_range: + umfang += f", commits {rev_range}" + if findings: + print(f"check_harvest: {len(findings)} finding(s) — {umfang}") + report(findings) + print("\nA harvest carries the failure class, its effect and its cause —") + print("never the adopter's names. Generalize, then run this again.") + return 1 + print(f"check_harvest: no findings — {umfang}") + print("⚠️ A denylist finds only what someone listed. Human review still applies.") + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_locked.py b/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_locked.py new file mode 100644 index 0000000..140b6b8 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/scripts/check_locked.py @@ -0,0 +1,351 @@ +#!/usr/bin/env python3 +"""check_locked.py — contradict an edit to an artifact the rules call binding. + +`AGENTS.md` has said since v0.1.0 that accepted decision records are +"binding; never edited, only superseded", and that `docs/sources/` holds +immutable originals. Both rules were prose, and prose is the category of +rule the second field report found broken without exception: an accepted +record was amended, committed and pushed, and the violation was caught by +chance rather than by anything that refused. The generated index carries +the same kind of prohibition and was never once violated, because a script +contradicts it. This is that script for the other two. + +Locked: + * every file under `docs/sources/` — modification, deletion or rename. + Adding a new source is not an edit and stays permitted. + * every ADR under `docs/adr/` that carried `status: accepted` *before* + the range. A newly added ADR is not locked, whatever status it + arrives with. + +Permitted, because the ADR template prescribes exactly this edit: setting +`status:` and `superseded_by:` on an accepted record when a newer one +replaces it. Nothing else about the file may change in that commit. + +Scope is the range it is given, and only that. Checking history would +report violations nobody can undo, which paints the pipeline permanently +red — and a permanently red check reports nothing. That failure class is +issue 0022's subject; walking into it here would be the same mistake with +a different subject. + +Fails closed (exit 1, never "clean"): + * a revision range git cannot resolve + * a blob git cannot produce + +⚠️ This sees one push. A violation that reaches the default branch by some +path this never runs on stays unseen; that is a boundary, not an oversight. + +Usage: + python scripts/check_locked.py [--range ..] + python scripts/check_locked.py --selftest +""" +from __future__ import annotations + +import difflib +import os +import re +import subprocess +import sys +import tempfile +from pathlib import Path + +LOCKED_DIR = "docs/sources/" +ADR_DIR = "docs/adr/" +PERMITTED_FIELD = re.compile(r"^(status|superseded_by):", re.I) +STATUS_ACCEPTED = re.compile(r"^status:\s*accepted\s*$", re.I | re.M) + + +class LockError(RuntimeError): + """The check could not answer the question. Never a pass.""" + + +def git(root: Path, *args: str) -> str: + """Run git, or raise. A failure is never an empty result.""" + proc = subprocess.run( + ["git", "-C", str(root), *args], + capture_output=True, text=True, + ) + if proc.returncode != 0: + raise LockError( + f"git {' '.join(args)} failed: {proc.stderr.strip() or 'no output'}" + ) + return proc.stdout + + +def split_range(rev_range: str) -> tuple[str, str]: + """'a..b' -> ('a', 'b'). An empty right side means HEAD.""" + if ".." not in rev_range: + raise LockError(f"not a revision range: {rev_range!r} (expected a..b)") + base, _, head = rev_range.partition("..") + base, head = base.strip(), head.strip() or "HEAD" + if not base: + raise LockError(f"range has no base: {rev_range!r}") + return base, head + + +def changed_paths(root: Path, base: str, head: str) -> list[tuple[str, str]]: + """[(status, path)] for the range. Status is git's A/M/D/R letter.""" + out = git(root, "diff", "--name-status", "--no-renames", base, head) + entries: list[tuple[str, str]] = [] + for line in out.splitlines(): + if not line.strip(): + continue + parts = line.split("\t") + entries.append((parts[0][0], parts[-1])) + return entries + + +def blob(root: Path, rev: str, path: str) -> str | None: + """File content at a revision, or None if it did not exist there.""" + proc = subprocess.run( + ["git", "-C", str(root), "show", f"{rev}:{path}"], + capture_output=True, text=True, + ) + if proc.returncode != 0: + if "exists on disk, but not in" in proc.stderr or "does not exist" in proc.stderr: + return None + if "path" in proc.stderr and "not in" in proc.stderr: + return None + return None + return proc.stdout + + +def was_accepted(root: Path, base: str, path: str) -> bool: + """Did this ADR carry `status: accepted` before the range?""" + before = blob(root, base, path) + return before is not None and bool(STATUS_ACCEPTED.search(before)) + + +def only_supersede_fields(root: Path, base: str, head: str, path: str) -> bool: + """True if the change touches nothing but status/superseded_by.""" + before = blob(root, base, path) + after = blob(root, head, path) + if before is None or after is None: + return False + diff = difflib.unified_diff( + before.splitlines(), after.splitlines(), n=0, lineterm="" + ) + touched = [ + line[1:].strip() + for line in diff + if line[:1] in "+-" and not line.startswith(("+++", "---")) + ] + if not touched: + return False + return all(PERMITTED_FIELD.match(line) for line in touched) + + +def check(root: Path, base: str, head: str) -> list[str]: + findings: list[str] = [] + for state, path in changed_paths(root, base, head): + if path.startswith(LOCKED_DIR): + if state != "A": + findings.append( + f"{path}: immutable source, {_verb(state)} in this range" + ) + continue + if path.startswith(ADR_DIR): + if not was_accepted(root, base, path): + continue + if state == "D": + findings.append(f"{path}: accepted decision record, deleted") + elif only_supersede_fields(root, base, head, path): + continue + else: + findings.append( + f"{path}: accepted decision record, edited — supersede it " + f"with a new record instead" + ) + return findings + + +def _verb(state: str) -> str: + return {"M": "modified", "D": "deleted", "R": "renamed"}.get(state, "changed") + + +def report(findings: list[str]) -> None: + if not findings: + print("check_locked: no findings") + return + print(f"check_locked: {len(findings)} finding(s)\n") + for f in findings: + print(f" {f}") + print( + "\nAn accepted record is superseded, never edited; a source under " + f"{LOCKED_DIR} is never changed at all." + ) + + +# -------------------------------------------------------------------------- +# Positive control. A gate that has only ever been seen green is a +# hypothesis — WORKFLOW.md, Gate 4. +# -------------------------------------------------------------------------- + +def selftest() -> int: + passed = 0 + failed: list[str] = [] + + def check_that(name: str, condition: bool) -> None: + nonlocal passed + if condition: + passed += 1 + print(f" ok {passed + len(failed)} {name}") + else: + failed.append(name) + print(f" FAIL {passed + len(failed)} {name}") + + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + run = lambda *a: git(root, *a) + git(root, "init", "-q", ".") + run("config", "user.email", "selftest@invalid") + run("config", "user.name", "selftest") + + adr = root / ADR_DIR + adr.mkdir(parents=True) + (adr / "0001-locked.md").write_text( + "---\ntype: adr\nid: \"0001\"\nstatus: accepted\n" + "superseded_by: null\n---\n\n# ADR-0001\n\nBody.\n", + encoding="utf-8", + ) + src = root / LOCKED_DIR + src.mkdir(parents=True) + (src / "digest.md").write_text("original\n", encoding="utf-8") + (root / "README.md").write_text("free\n", encoding="utf-8") + run("add", "-A") + run("commit", "-qm", "base") + base = run("rev-parse", "HEAD").strip() + + # 1 — an accepted record, edited + (adr / "0001-locked.md").write_text( + "---\ntype: adr\nid: \"0001\"\nstatus: accepted\n" + "superseded_by: null\n---\n\n# ADR-0001\n\nBody, amended.\n", + encoding="utf-8", + ) + run("add", "-A"); run("commit", "-qm", "amend an accepted adr") + check_that("an edit to an accepted record is reported", + len(check(root, base, "HEAD")) == 1) + run("reset", "-q", "--hard", base) + + # 2 — an immutable source, edited + (src / "digest.md").write_text("rewritten\n", encoding="utf-8") + run("add", "-A"); run("commit", "-qm", "edit a source") + check_that("an edit under the immutable sources is reported", + len(check(root, base, "HEAD")) == 1) + run("reset", "-q", "--hard", base) + + # 3 — negative control + (root / "README.md").write_text("free, changed\n", encoding="utf-8") + run("add", "-A"); run("commit", "-qm", "ordinary change") + check_that("an ordinary change is not reported", + check(root, base, "HEAD") == []) + run("reset", "-q", "--hard", base) + + # 4 — the one edit the template prescribes + (adr / "0001-locked.md").write_text( + "---\ntype: adr\nid: \"0001\"\nstatus: superseded\n" + "superseded_by: docs/adr/0002-newer.md\n---\n\n# ADR-0001\n\nBody.\n", + encoding="utf-8", + ) + run("add", "-A"); run("commit", "-qm", "supersede") + check_that("setting status and superseded_by stays permitted", + check(root, base, "HEAD") == []) + run("reset", "-q", "--hard", base) + + # 5 — a new record may arrive accepted + (adr / "0002-newer.md").write_text( + "---\ntype: adr\nid: \"0002\"\nstatus: accepted\n---\n\n# ADR-0002\n", + encoding="utf-8", + ) + run("add", "-A"); run("commit", "-qm", "add an accepted adr") + check_that("a newly added record is not locked", + check(root, base, "HEAD") == []) + run("reset", "-q", "--hard", base) + + # 6 — a new source may be added + (src / "second.md").write_text("new digest\n", encoding="utf-8") + run("add", "-A"); run("commit", "-qm", "add a source") + check_that("adding a new source is not an edit", + check(root, base, "HEAD") == []) + run("reset", "-q", "--hard", base) + + # 7 — a source that is deleted + (src / "digest.md").unlink() + run("add", "-A"); run("commit", "-qm", "delete a source") + check_that("deleting an immutable source is reported", + len(check(root, base, "HEAD")) == 1) + run("reset", "-q", "--hard", base) + + # 8 — fails closed on a range git cannot resolve + check_that("an unresolvable range raises instead of reporting nothing", + _raises(lambda: check(root, "no-such-ref", "HEAD"))) + + # 9 — fails closed on a malformed range, through main() + check_that("main() rejects a range that is not a..b", + main(["--range", "HEAD"], root=root) == 1) + + # 10 — end to end through main(): a real violation exits 1 + (adr / "0001-locked.md").write_text( + "---\ntype: adr\nid: \"0001\"\nstatus: accepted\n" + "superseded_by: null\n---\n\n# ADR-0001\n\nBody, amended again.\n", + encoding="utf-8", + ) + run("add", "-A"); run("commit", "-qm", "amend again") + check_that("main() exits 1 on a real violation", + main(["--range", f"{base}..HEAD"], root=root) == 1) + run("reset", "-q", "--hard", base) + + # 11 — end to end through main(): a clean range exits 0 + (root / "README.md").write_text("free, again\n", encoding="utf-8") + run("add", "-A"); run("commit", "-qm", "ordinary") + check_that("main() exits 0 on a clean range", + main(["--range", f"{base}..HEAD"], root=root) == 0) + + total = passed + len(failed) + print() + if failed: + print(f"selftest: {len(failed)} of {total} assertions FAILED") + for name in failed: + print(f" - {name}") + return 1 + print(f"selftest: {total} assertions passed") + return 0 + + +def _raises(fn) -> bool: + try: + fn() + except LockError: + return True + except Exception: + return False + return False + + +def main(argv: list[str] | None = None, *, root: Path | None = None) -> int: + args = list(sys.argv[1:] if argv is None else argv) + if "--selftest" in args: + return selftest() + + rev_range = "origin/main..HEAD" + if "--range" in args: + i = args.index("--range") + if i + 1 >= len(args): + print("check_locked: --range needs a value", file=sys.stderr) + return 1 + rev_range = args[i + 1] + + repo = root or Path(os.environ.get("CHECK_LOCKED_ROOT", ".")).resolve() + try: + base, head = split_range(rev_range) + findings = check(repo, base, head) + except LockError as exc: + print(f"check_locked: {exc}", file=sys.stderr) + print("unchecked is not passed.", file=sys.stderr) + return 1 + + report(findings) + return 1 if findings else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docs/sources/upstream/neckbeard-v0.3.1/scripts/gen_status.py b/docs/sources/upstream/neckbeard-v0.3.1/scripts/gen_status.py new file mode 100644 index 0000000..e829371 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/scripts/gen_status.py @@ -0,0 +1,143 @@ +#!/usr/bin/env python3 +"""gen_status.py — generate STATUS.md deterministically from frontmatter. + +Writes STATUS.md (no timestamps — output depends only on repo content, so +reruns are diff-clean). With --check, regenerates in memory and fails if +the committed STATUS.md is stale; CI uses this mode. + +Usage: + python scripts/gen_status.py [repo-root] # write STATUS.md + python scripts/gen_status.py --check [repo-root] # verify, exit 1 if stale +""" +from __future__ import annotations + +import re +import sys +from pathlib import Path + +try: + import yaml +except ImportError: # pragma: no cover + sys.exit("gen_status.py needs PyYAML: pip install pyyaml") + +H1_RE = re.compile(r"^#\s+(.*)$", re.M) + + +def parse(path: Path): + lines = path.read_text(encoding="utf-8").splitlines() + if not lines or lines[0].strip() != "---": + return None, "" + for j in range(1, len(lines)): + if lines[j].strip() == "---": + meta = yaml.safe_load("\n".join(lines[1:j])) or {} + body = "\n".join(lines[j + 1:]) + return meta, body + return None, "" + + +def title(body: str, fallback: str) -> str: + match = H1_RE.search(body) + return match.group(1).strip() if match else fallback + + +def collect(root: Path, subdir: str, wanted_type: str): + items = [] + base = root / subdir + if not base.is_dir(): + return items + for path in sorted(base.rglob("*.md")): + if path.name == "template.md": + continue + meta, body = parse(path) + if not isinstance(meta, dict) or meta.get("type") != wanted_type: + continue + rel = path.relative_to(root).as_posix() + items.append((rel, meta, title(body, path.stem))) + return items + + +def render(root: Path) -> str: + issues = collect(root, "docs/issues", "issue") + designs = collect(root, "docs/design", "design") + adrs = collect(root, "docs/adr", "adr") + aars = collect(root, "docs/aar", "aar") + + out: list[str] = [] + out.append("# STATUS") + out.append("") + out.append("") + out.append("") + + open_issues = [i for i in issues + if i[1].get("status") in ("open", "in-progress")] + closed = len(issues) - len(open_issues) + out.append(f"## Issues ({len(open_issues)} open, {closed} closed)") + out.append("") + if open_issues: + out.append("| Issue | Status | Title |") + out.append("|---|---|---|") + for rel, meta, name in open_issues: + out.append(f"| [{meta.get('id', '?')}]({rel}) " + f"| {meta.get('status')} | {name} |") + else: + out.append("_none open_") + out.append("") + + active = [d for d in designs if d[1].get("status") != "done"] + out.append(f"## Active design docs ({len(active)})") + out.append("") + if active: + out.append("| Design | Gate | Title |") + out.append("|---|---|---|") + for rel, meta, name in active: + out.append(f"| [{Path(rel).stem}]({rel}) " + f"| {meta.get('status')} | {name} |") + else: + out.append("_none active_") + out.append("") + + out.append(f"## ADRs ({len(adrs)})") + out.append("") + if adrs: + out.append("| ADR | Status | Title |") + out.append("|---|---|---|") + for rel, meta, name in adrs: + out.append(f"| [{meta.get('id', '?')}]({rel}) " + f"| {meta.get('status')} | {name} |") + else: + out.append("_none_") + out.append("") + + open_aars = [a for a in aars if a[1].get("status") == "open"] + out.append(f"## Open AARs ({len(open_aars)})") + out.append("") + if open_aars: + for rel, _meta, name in open_aars: + out.append(f"- [{name}]({rel})") + else: + out.append("_none — nothing awaiting harvest_") + out.append("") + return "\n".join(out) + + +def main() -> int: + args = [a for a in sys.argv[1:] if a != "--check"] + check = "--check" in sys.argv[1:] + root = Path(args[0]) if args else Path.cwd() + content = render(root) + status = root / "STATUS.md" + if check: + current = status.read_text(encoding="utf-8") if status.is_file() else "" + if current != content: + print("gen_status --check: STATUS.md is stale — " + "run scripts/gen_status.py and commit the result") + return 1 + print("gen_status --check: STATUS.md is current") + return 0 + status.write_text(content, encoding="utf-8", newline="\n") + print(f"wrote {status}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docs/sources/upstream/neckbeard-v0.3.1/scripts/judge.py b/docs/sources/upstream/neckbeard-v0.3.1/scripts/judge.py new file mode 100644 index 0000000..25f5b18 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/scripts/judge.py @@ -0,0 +1,564 @@ +#!/usr/bin/env python3 +"""judge.py — read a run's trace against the workflow. Build form A. + +A judge cannot examine behaviour, only traces. This one reads a session +ledger (`docs/ledger/`) against `WORKFLOW.md`'s rules and against git, +and reports where the two disagree. It is deterministic, stdlib-only and +model-agnostic; the inferential half of judging is build form B, a +separate ritual in fresh context (see WORKFLOW.md, "Judging a run"). + +What it checks + * gate rows are ordered, without duplicates, and complete for the + declared size class — S owes no gates, and is not judged as if it did + * every status comes from the four-value vocabulary + * a gate that has a successor carries an approval + * every named commit resolves, is an ancestor of HEAD, and the commits + run in the same order as the gates they belong to + * the ladder section exists, and each entry names what was searched, + what was found, and an outcome beginning `reused:` or `built:` + +The commit checks matter more than the rest: they are the only part that +compares the ledger against evidence the ledger's author did not write. + +What it deliberately does not do + * judge a session that has no ledger. History cannot be instrumented + after the fact, and a check that reports every past session forever is + a check nobody reads. + * judge quality. Whether a design document is real or filler is form B's + question, and no script can answer it. + * resist tampering. The ledger is written by the agent it describes. + The threat model is drift, not sabotage — ADR-0010 records this as an + explicit non-goal, because a judge that suggests otherwise is worse + than none. + +Coverage + `--coverage` reports which rules of the framework can be observed at all + and by what — a script, the ledger, or nothing. Rules at "nothing" are + either unobservable or inert, and both are worth knowing. This is a + report, never a gate: it must not turn red, or it becomes the standing + finding nobody can fix. + +Fails closed (exit 1, never "clean"): a ledger that cannot be read or +parsed, and a commit git cannot resolve. + +Usage: + python scripts/judge.py --ledger docs/ledger/.md [--root .] + python scripts/judge.py --coverage + python scripts/judge.py --selftest +""" +from __future__ import annotations + +import re +import subprocess +import sys +import tempfile +from pathlib import Path +from typing import NamedTuple + +STATUSES = ("DONE", "DONE_WITH_CONCERNS", "NEEDS_CONTEXT", "BLOCKED") +GATES_FOR_SIZE = {"S": set(), "M": set(), "L": {"1", "2", "3", "4", "5"}} +OUTCOME_PREFIX = ("reused:", "built:") +PENDING = "pending" +PLACEHOLDERS = {"", "-", "—", "n/a", "none", "tbd", "todo", "?"} + + +class JudgeError(RuntimeError): + """The check could not answer the question. Never a pass.""" + + +class Row(NamedTuple): + cells: list[str] + + def get(self, i: int) -> str: + return self.cells[i].strip() if i < len(self.cells) else "" + + +class Finding(NamedTuple): + bucket: str # "model-failure" | "framework-gap" + rule: str + detail: str + + +def git(root: Path, *args: str) -> str: + """Run git, or raise. A failure is never an empty result.""" + proc = subprocess.run(["git", "-C", str(root), *args], + capture_output=True, text=True) + if proc.returncode != 0: + raise JudgeError( + f"git {' '.join(args)} failed: {proc.stderr.strip() or 'no output'}") + return proc.stdout + + +# -------------------------------------------------------------------------- +# Reading the ledger. Frontmatter is read with re rather than PyYAML: the +# envelope is already validated by validate.py against schema.yaml, so this +# only needs the two or three scalars it acts on — and adopters copying one +# file should not inherit a dependency. +# -------------------------------------------------------------------------- + +def parse_ledger(path: Path) -> tuple[dict, list[Row], list[Row]]: + try: + text = path.read_text(encoding="utf-8") + except OSError as exc: + raise JudgeError(f"cannot read ledger {path}: {exc}") from exc + + match = re.match(r"^---\n(.*?)\n---\n", text, re.S) + if not match: + raise JudgeError(f"{path}: no frontmatter block") + meta: dict[str, object] = {} + current: str | None = None + for line in match.group(1).splitlines(): + item = re.match(r"^\s+-\s*(.+)$", line) + if item and current: + meta.setdefault(current, []) + if isinstance(meta[current], list): + meta[current].append(item.group(1).strip().strip('"')) + continue + kv = re.match(r"^([a-z_]+):\s*(.*)$", line) + if kv: + value = kv.group(2).strip().strip('"') + current = kv.group(1) + meta[current] = value if value not in ("", "[]") else [] + for field in ("type", "size", "status"): + if not meta.get(field): + raise JudgeError(f"{path}: frontmatter is missing '{field}'") + if meta["type"] != "ledger": + raise JudgeError(f"{path}: type is '{meta['type']}', not 'ledger'") + + body = text[match.end():] + return meta, _table(body, "Gates"), _table(body, "Ladder") + + +def _table(body: str, heading: str) -> list[Row]: + """Rows of the Markdown table under '## ', header excluded.""" + section = re.search(rf"^##\s+{heading}\s*$(.*?)(?=^##\s|\Z)", + body, re.S | re.M) + if not section: + return [] + rows: list[Row] = [] + for line in section.group(1).splitlines(): + line = line.strip() + if not line.startswith("|"): + continue + cells = [c.strip() for c in line.strip("|").split("|")] + if all(re.fullmatch(r":?-{2,}:?", c) for c in cells): + continue # the ---|--- separator + rows.append(Row(cells)) + return rows[1:] if rows else rows # drop the header row + + +# -------------------------------------------------------------------------- +# The checks +# -------------------------------------------------------------------------- + +def check_gate_order(rows: list[Row], size: str, closed: bool) -> list[Finding]: + out: list[Finding] = [] + seen = [r.get(0) for r in rows] + numbers = [s for s in seen if s.isdigit()] + if len(set(numbers)) != len(numbers): + out.append(Finding("model-failure", "workflow/gate-order", + f"a gate is recorded twice: {numbers}")) + if numbers != sorted(numbers, key=int): + out.append(Finding("model-failure", "workflow/gate-order", + f"gates are out of order: {numbers}")) + owed = GATES_FOR_SIZE.get(size.upper(), set()) + if closed: + missing = sorted(owed - set(numbers)) + if missing: + out.append(Finding("model-failure", "workflow/gates-for-size", + f"size {size} owes gates {missing}, not recorded")) + return out + + +def check_status_vocabulary(rows: list[Row]) -> list[Finding]: + out: list[Finding] = [] + for row in rows: + status = row.get(3) + if status and status not in STATUSES: + out.append(Finding("model-failure", "agents/status-vocabulary", + f"gate {row.get(0)}: '{status}' is not one of " + f"{', '.join(STATUSES)}")) + return out + + +def check_approvals(rows: list[Row]) -> list[Finding]: + """A gate with a successor must carry an approval.""" + out: list[Finding] = [] + for i, row in enumerate(rows[:-1]): + if row.get(2).lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "workflow/stop-before-next-gate", + f"gate {row.get(0)} has no approval, but gate " + f"{rows[i + 1].get(0)} was started")) + return out + + +def check_commits(root: Path, rows: list[Row]) -> list[Finding]: + """The only part measured against evidence the author did not write.""" + out: list[Finding] = [] + order: list[tuple[str, int]] = [] + for row in rows: + sha = row.get(1) + if not sha or sha.lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "ledger/commit-required", + f"gate {row.get(0)} names no commit")) + continue + try: + full = git(root, "rev-parse", "--verify", f"{sha}^{{commit}}").strip() + except JudgeError as exc: + raise JudgeError( + f"gate {row.get(0)} names commit {sha}, which does not " + f"resolve — unchecked is not passed ({exc})") from exc + ancestor = subprocess.run( + ["git", "-C", str(root), "merge-base", "--is-ancestor", full, "HEAD"], + capture_output=True, text=True) + if ancestor.returncode != 0: + out.append(Finding("model-failure", "ledger/commit-reachable", + f"gate {row.get(0)}: commit {sha} is not an " + f"ancestor of HEAD")) + continue + depth = int(git(root, "rev-list", "--count", f"{full}..HEAD").strip()) + order.append((row.get(0), depth)) + + ranked = [gate for gate, _ in sorted(order, key=lambda p: -p[1])] + stated = [gate for gate, _ in order] + if ranked != stated: + out.append(Finding("framework-gap", "ledger/commit-order", + f"the commits run in the order {ranked}, the gates " + f"claim {stated}")) + return out + + +def check_design_doc(meta: dict, root: Path, closed: bool) -> list[Finding]: + """Size L owes a design document. The coverage table claimed this + before anything checked it — the inventory drift its own entry warns + about, caught on the first read.""" + if str(meta.get("size", "")).upper() != "L" or not closed: + return [] + related = meta.get("related") or [] + if isinstance(related, str): + related = [related] + if any(r.startswith("docs/design/") for r in related): + return [] + return [Finding("model-failure", "workflow/design-doc-for-L", + "size L, but the ledger names no design document in " + "'related'")] + + +def check_ladder(rows: list[Row], closed: bool) -> list[Finding]: + out: list[Finding] = [] + if not rows: + out.append(Finding("framework-gap", "agents/ponytail-ladder", + "the Ladder section is empty — the gate was not " + "passed. A session that built nothing says so.")) + return out + for i, row in enumerate(rows, start=1): + searched, found, outcome = row.get(0), row.get(1), row.get(2) + if searched.lower() in PLACEHOLDERS or found.lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "agents/ponytail-ladder", + f"ladder entry {i} names no concrete candidate")) + if not outcome.lower().startswith(OUTCOME_PREFIX): + out.append(Finding("model-failure", "agents/ponytail-ladder", + f"ladder entry {i}: outcome must begin " + f"'reused:' or 'built:', got '{outcome[:30]}'")) + if closed and row.get(3).lower() in PLACEHOLDERS | {PENDING}: + out.append(Finding("model-failure", "ledger/commit-required", + f"ladder entry {i} is still 'pending' in a " + f"closed ledger")) + return out + + +def judge(root: Path, ledger: Path) -> list[Finding]: + meta, gates, ladder = parse_ledger(ledger) + closed = meta.get("status") == "closed" + size = meta.get("size", "L") + findings = check_gate_order(gates, size, closed) + findings += check_status_vocabulary(gates) + findings += check_approvals(gates) + findings += check_commits(root, gates) + findings += check_design_doc(meta, root, closed) + findings += check_ladder(ladder, closed) + return findings + + +# -------------------------------------------------------------------------- +# Rule coverage. A report, never a gate. +# +# The inventory is maintained by hand because the rules live in prose and +# nothing derives them mechanically. That is a known weakness: it will drift +# from AGENTS.md and WORKFLOW.md unless someone updates it, and nothing +# contradicts that drift today. +# -------------------------------------------------------------------------- + +COVERAGE = [ + # (rule, where it is written, what observes it) + ("artifact frontmatter and enums", "AGENTS.md §5", "validate.py"), + ("link targets are files, never directories", "AGENTS.md §5", "validate.py"), + ("STATUS.md is generated, never hand-edited", "AGENTS.md §4", "gen_status.py --check"), + ("accepted ADRs are never edited", "AGENTS.md §4", "check_locked.py"), + ("docs/sources is immutable", "AGENTS.md §5", "check_locked.py"), + ("a harvest carries no adopter specifics", "WORKFLOW.md", "check_harvest.py"), + ("PROJECT.md answers Gate 0", "AGENTS.md §2", "validate.py"), + ("a design doc is done iff it sits in done/", "WORKFLOW.md", "validate.py"), + ("the ponytail ladder was walked", "AGENTS.md §1", "ledger"), + ("a size class is declared for the run", "AGENTS.md §3", "ledger"), + ("gates run in order", "WORKFLOW.md", "ledger"), + ("a gate is approved before the next begins", "WORKFLOW.md", "ledger"), + ("every slice reports one of four statuses", "AGENTS.md §1", "ledger"), + ("size L owes a design document", "WORKFLOW.md", "ledger"), + ("simplicity: the minimum that solves it", "AGENTS.md §1", None), + ("surgical changes: every line traces to the request", "AGENTS.md §1", None), + ("slice 1 is a tracer bullet", "WORKFLOW.md", None), + ("a check owes proof it can fail", "WORKFLOW.md", None), + ("a delivering artifact owes one real result", "WORKFLOW.md", None), + ("a closeout names what it made false", "WORKFLOW.md", None), + ("reproduce before fixing", "WORKFLOW.md", None), + ("incidents get an AAR", "WORKFLOW.md", None), + ("a permanent exception owes an ADR", "AGENTS.md §1", None), +] + + +def coverage_report() -> int: + by_script = [r for r in COVERAGE if r[2] and r[2] != "ledger"] + by_ledger = [r for r in COVERAGE if r[2] == "ledger"] + unobserved = [r for r in COVERAGE if r[2] is None] + total = len(COVERAGE) + + print(f"rule coverage: {total} rules inventoried\n") + for title, group in (("observed by a script", by_script), + ("observed by the ledger", by_ledger), + ("not observable today", unobserved)): + print(f" {title} — {len(group)}") + for rule, where, how in group: + suffix = f" [{how}]" if how else "" + print(f" {rule} ({where}){suffix}") + print() + observed = len(by_script) + len(by_ledger) + print(f" {observed}/{total} observable, {len(unobserved)} decoration " + f"until they gain a duty to leave a trace.") + print("\nThis is a report, not a gate. A rule at zero coverage is either") + print("unobservable or inert; neither is a violation of anything.") + return 0 + + +# -------------------------------------------------------------------------- +# Positive control. A gate that has only ever been seen green is a +# hypothesis — WORKFLOW.md, Gate 4. +# -------------------------------------------------------------------------- + +LEDGER_HEAD = """--- +type: ledger +date: 2026-08-21 +size: {size} +status: {status} +related: +{related}--- + +## Gates + +| gate | commit | approval | status | note | +|---|---|---|---|---| +{gates} + +## Ladder + +| searched | found | outcome | commit | +|---|---|---|---| +{ladder} +""" + + +def _write(path: Path, *, size="L", status="closed", gates="", ladder="", + related=' - "docs/design/d.md"\n'): + path.write_text(LEDGER_HEAD.format(size=size, status=status, + gates=gates, ladder=ladder, + related=related), + encoding="utf-8") + + +def selftest() -> int: + passed, failed = 0, [] + + def check_that(name: str, condition: bool) -> None: + nonlocal passed + if condition: + passed += 1 + print(f" ok {passed + len(failed)} {name}") + else: + failed.append(name) + print(f" FAIL {passed + len(failed)} {name}") + + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + git(root, "init", "-q", ".") + git(root, "config", "user.email", "selftest@invalid") + git(root, "config", "user.name", "selftest") + shas = [] + for i in range(5): + (root / f"f{i}.txt").write_text(f"{i}\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-qm", f"c{i}") + shas.append(git(root, "rev-parse", "--short", "HEAD").strip()) + + led = root / "ledger.md" + good_ladder = ("| validate.py | a generic engine | reused: schema entry " + f"only | {shas[0]} |") + rows = lambda gates: "\n".join( + f"| {g} | {shas[i]} | owner | DONE | n |" for i, g in enumerate(gates)) + + _write(led, gates=rows("12345"), ladder=good_ladder) + check_that("a clean ledger produces no findings", + judge(root, led) == []) + + _write(led, gates=rows("13245"), ladder=good_ladder) + check_that("a gate out of order is reported", + any(f.rule == "workflow/gate-order" for f in judge(root, led))) + + _write(led, gates=rows("11234"), ladder=good_ladder) + check_that("a duplicated gate is reported", + any("twice" in f.detail for f in judge(root, led))) + + _write(led, gates=rows("123"), ladder=good_ladder) + check_that("a closed size-L ledger missing gates 4 and 5 is reported", + any(f.rule == "workflow/gates-for-size" + for f in judge(root, led))) + + _write(led, size="S", gates="", ladder=good_ladder) + check_that("size S is not judged against gates it does not owe", + not any(f.rule == "workflow/gates-for-size" + for f in judge(root, led))) + + _write(led, gates=f"| 1 | {shas[0]} | owner | FINISHED | n |", + ladder=good_ladder) + check_that("a status outside the vocabulary is reported", + any(f.rule == "agents/status-vocabulary" + for f in judge(root, led))) + + _write(led, gates=(f"| 1 | {shas[0]} | | DONE | n |\n" + f"| 2 | {shas[1]} | owner | DONE | n |"), + ladder=good_ladder) + check_that("a gate whose successor began without approval is reported", + any(f.rule == "workflow/stop-before-next-gate" + for f in judge(root, led))) + + _write(led, gates="| 1 | deadbee | owner | DONE | n |", + ladder=good_ladder) + check_that("a commit that does not resolve raises instead of passing", + _raises(lambda: judge(root, led))) + + git(root, "checkout", "-q", "-b", "side", shas[0]) + (root / "side.txt").write_text("x\n", encoding="utf-8") + git(root, "add", "-A"); git(root, "commit", "-qm", "side") + off = git(root, "rev-parse", "--short", "HEAD").strip() + git(root, "checkout", "-q", "main" if _has(root, "main") else "master") + _write(led, gates=f"| 1 | {off} | owner | DONE | n |", ladder=good_ladder) + check_that("a commit that is not an ancestor of HEAD is reported", + any(f.rule == "ledger/commit-reachable" + for f in judge(root, led))) + + _write(led, gates=(f"| 1 | {shas[3]} | owner | DONE | n |\n" + f"| 2 | {shas[1]} | owner | DONE | n |"), + ladder=good_ladder) + check_that("commits running in a different order than the gates is reported", + any(f.rule == "ledger/commit-order" for f in judge(root, led))) + + _write(led, gates=rows("12345"), ladder="") + check_that("an empty ladder section is reported as a gate not passed", + any(f.rule == "agents/ponytail-ladder" + for f in judge(root, led))) + + _write(led, gates=rows("12345"), + ladder=f"| - | - | reused: something | {shas[0]} |") + check_that("a ladder entry naming no concrete candidate is reported", + any("concrete candidate" in f.detail for f in judge(root, led))) + + _write(led, gates=rows("12345"), + ladder=f"| validate.py | an engine | it was fine | {shas[0]} |") + check_that("a ladder outcome without reused:/built: is reported", + any("must begin" in f.detail for f in judge(root, led))) + + (root / "broken.md").write_text("no frontmatter here\n", encoding="utf-8") + check_that("an unreadable ledger raises instead of reporting nothing", + _raises(lambda: judge(root, root / "broken.md"))) + + _write(led, gates=rows("12345"), ladder=good_ladder) + check_that("main() exits 0 on a clean ledger", + main(["--ledger", str(led), "--root", str(root)]) == 0) + + _write(led, gates=rows("12345"), ladder=good_ladder, related="") + check_that("a closed size-L ledger naming no design document is reported", + any(f.rule == "workflow/design-doc-for-L" + for f in judge(root, led))) + + _write(led, gates=rows("13245"), ladder=good_ladder) + check_that("main() exits 1 on a real violation", + main(["--ledger", str(led), "--root", str(root)]) == 1) + + total = passed + len(failed) + print() + if failed: + print(f"selftest: {len(failed)} of {total} assertions FAILED") + for name in failed: + print(f" - {name}") + return 1 + print(f"selftest: {total} assertions passed") + return 0 + + +def _has(root: Path, branch: str) -> bool: + return subprocess.run(["git", "-C", str(root), "rev-parse", "--verify", branch], + capture_output=True).returncode == 0 + + +def _raises(fn) -> bool: + try: + fn() + except JudgeError: + return True + except Exception: + return False + return False + + +def report(findings: list[Finding]) -> None: + if not findings: + print("judge: no findings") + return + print(f"judge: {len(findings)} finding(s)\n") + for bucket in ("framework-gap", "model-failure"): + group = [f for f in findings if f.bucket == bucket] + if not group: + continue + print(f" {bucket} ({len(group)}):") + for f in group: + print(f" [{f.rule}] {f.detail}") + print() + print("A finding is the model not following a clear rule, or a gap in the") + print("framework. Deciding which is the point — see ADR-0010.") + + +def main(argv: list[str] | None = None) -> int: + args = list(sys.argv[1:] if argv is None else argv) + if "--selftest" in args: + return selftest() + if "--coverage" in args: + return coverage_report() + + if "--ledger" not in args: + print("judge: --ledger is required (or --coverage / --selftest)", + file=sys.stderr) + return 1 + ledger = Path(args[args.index("--ledger") + 1]) + root = Path(args[args.index("--root") + 1]) if "--root" in args else Path(".") + + try: + findings = judge(root.resolve(), ledger) + except JudgeError as exc: + print(f"judge: {exc}", file=sys.stderr) + print("unchecked is not passed.", file=sys.stderr) + return 1 + report(findings) + return 1 if findings else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docs/sources/upstream/neckbeard-v0.3.1/scripts/validate.py b/docs/sources/upstream/neckbeard-v0.3.1/scripts/validate.py new file mode 100644 index 0000000..acb8bc0 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/scripts/validate.py @@ -0,0 +1,283 @@ +#!/usr/bin/env python3 +"""validate.py — deterministic artifact validation against schema.yaml. + +Checks (errors, exit 1): + * frontmatter present, parseable, `type` known + * file location and filename match the type's rules + * required fields, enums, patterns, dates + * link fields: repo-root-relative targets exist (http/https/mailto skipped) + * inline markdown links in bodies resolve (relative to the file) + * per-type rules: superseded_requires_pointer, done_iff_in_done_dir + +Warnings (exit 0): + * wiki pages (except index) with no inbound link anywhere + +Usage: python scripts/validate.py [repo-root] +""" +from __future__ import annotations + +import datetime +import fnmatch +import re +import sys +from pathlib import Path + +try: + import yaml +except ImportError: # pragma: no cover + sys.exit("validate.py needs PyYAML: pip install pyyaml") + +DATE_RE = re.compile(r"^\d{4}-\d{2}-\d{2}$") +INLINE_LINK_RE = re.compile(r"\]\(([^)\s]+)\)") +HTML_SRC_RE = re.compile(r"(?:src|srcset)=\"([^\"]+)\"") +EXTERNAL_PREFIXES = ("http://", "https://", "mailto:") + +errors: list[str] = [] +warnings: list[str] = [] + + +def err(path: Path, msg: str) -> None: + errors.append(f"ERROR {path}: {msg}") + + +def warn(path: Path, msg: str) -> None: + warnings.append(f"WARN {path}: {msg}") + + +def parse_frontmatter(text: str): + lines = text.splitlines() + if not lines or lines[0].strip() != "---": + return None, text + for j in range(1, len(lines)): + if lines[j].strip() == "---": + fm = "\n".join(lines[1:j]) + body = "\n".join(lines[j + 1:]) + return yaml.safe_load(fm) or {}, body + return None, text # unterminated + + +def is_date(value) -> bool: + if isinstance(value, datetime.date): + return True + return isinstance(value, str) and bool(DATE_RE.match(value)) + + +def as_links(value): + """Normalize a link field's value to a list of strings.""" + if value is None: + return [] + if isinstance(value, str): + return [value] + if isinstance(value, list): + return [v for v in value if isinstance(v, str)] + return None # wrong shape + + +def discover(root: Path, scope: dict) -> list[Path]: + files: set[Path] = set() + for pattern in scope.get("include", []): + files.update(root.glob(pattern)) + result = [] + for f in sorted(files): + rel = f.relative_to(root).as_posix() + if any(fnmatch.fnmatch(rel, pat) for pat in scope.get("exclude", [])): + continue + if f.is_file(): + result.append(f) + return result + + +def check_fields(path: Path, meta: dict, spec: dict, root: Path) -> None: + for field in spec.get("required", []): + if field not in meta or meta[field] is None: + err(path, f"missing required field '{field}'") + for field, rule in (spec.get("fields") or {}).items(): + if field not in meta: + continue + value = meta[field] + if value is None: + if not rule.get("nullable"): + # required-check already covers required fields; + # a present-but-null optional field is fine unless typed link + pass + continue + if "enum" in rule and value not in rule["enum"]: + err(path, f"'{field}: {value}' not in enum {rule['enum']}") + if "pattern" in rule and not re.match(rule["pattern"], str(value)): + err(path, f"'{field}: {value}' does not match {rule['pattern']}") + kind = rule.get("kind") + if kind == "date" and not is_date(value): + err(path, f"'{field}: {value}' is not a YYYY-MM-DD date") + if kind == "bool" and not isinstance(value, bool): + err(path, f"'{field}: {value}' is not a boolean") + if kind == "str" and not isinstance(value, str): + err(path, f"'{field}' must be a string") + + +def check_links(path: Path, meta: dict, link_fields: list, root: Path, + inbound: set) -> None: + for field in link_fields: + if field not in meta: + continue + links = as_links(meta[field]) + if links is None: + err(path, f"'{field}' must be a string or list of strings") + continue + for link in links: + if link.startswith(EXTERNAL_PREFIXES): + continue + target = (root / link) + if not target.is_file(): + err(path, f"'{field}' link target missing: {link}") + else: + inbound.add(target.resolve()) + + +def check_body_links(path: Path, body: str, root: Path, inbound: set) -> None: + # strip fenced code blocks and inline code spans so mermaid, code + # samples, and literal link examples in backticks aren't scanned + body = re.sub(r"```.*?```", "", body, flags=re.S) + body = re.sub(r"`[^`\n]*`", "", body) + candidates = [m.group(1) for m in INLINE_LINK_RE.finditer(body)] + for raw in (m.group(1) for m in HTML_SRC_RE.finditer(body)): + # srcset may list "path 2x, path2 1x" pairs — take each path token + for part in raw.split(","): + candidates.append(part.strip().split()[0]) + for link in candidates: + if link.startswith(EXTERNAL_PREFIXES) or link.startswith("#"): + continue + link = link.split("#", 1)[0] + if not link: + continue + target = (path.parent / link).resolve() + if not target.is_file(): + err(path, f"inline link target missing: {link}") + else: + inbound.add(target) + + +def check_vendored_portable(root: Path, vendored: list, marker: str) -> None: + """Files adopters hold byte-identical must carry no repo-relative link. + + They are copied into repositories without this repo's docs/, so such a + link resolves here and nowhere else — and the adoption path in + AGENTS.md §5 asks for a byte-for-byte copy, which makes the dead link + the adopter's problem and unfixable without breaking the copy. + """ + for rel in vendored or []: + path = root / rel + if not path.is_file(): + continue + body = path.read_text(encoding="utf-8") + if marker and marker in body: + # Everything below the marker is the adopting project's own + # section: never copied elsewhere, so its links are fine. + body = body.split(marker, 1)[0] + body = re.sub(r"```.*?```", "", body, flags=re.S) + body = re.sub(r"`[^`\n]*`", "", body) + for link in (m.group(1) for m in INLINE_LINK_RE.finditer(body)): + if link.startswith(EXTERNAL_PREFIXES) or link.startswith("#"): + continue + err(path, f"vendored file carries a repo-relative link: {link} " + f"— name the target instead of linking to it") + + +def apply_rules(path: Path, rel: str, meta: dict, spec: dict) -> None: + for rule in spec.get("rules", []): + if rule == "superseded_requires_pointer": + if meta.get("status") == "superseded" and not meta.get("superseded_by"): + err(path, "status 'superseded' requires 'superseded_by'") + elif rule == "done_iff_in_done_dir": + in_done = "/done/" in f"/{rel}" + if (meta.get("status") == "done") != in_done: + err(path, "status 'done' <-> file in docs/design/done/ mismatch") + + +def main() -> int: + root = Path(sys.argv[1]) if len(sys.argv) > 1 else Path.cwd() + schema = yaml.safe_load((root / "schema.yaml").read_text(encoding="utf-8")) + link_fields = schema.get("link_fields", []) + types = schema.get("types", {}) + inbound: set = set() + wiki_pages: list[tuple[Path, dict]] = [] + + check_vendored_portable(root, schema.get("vendored", []), + schema.get("project_section_marker", "")) + + # Root documents: inline links must resolve; no frontmatter required. + for rel in schema.get("scope", {}).get("link_only", []): + path = root / rel + if not path.is_file(): + continue # e.g. STATUS.md before first generation + text = path.read_text(encoding="utf-8") + _meta, body = parse_frontmatter(text) + check_body_links(path, body if _meta is not None else text, + root, inbound) + seen_ids: dict[tuple[str, str], Path] = {} + artifacts = discover(root, schema.get("scope", {})) + + for path in artifacts: + rel = path.relative_to(root).as_posix() + meta, body = parse_frontmatter(path.read_text(encoding="utf-8")) + if meta is None: + err(path, "missing or unterminated YAML frontmatter") + continue + if not isinstance(meta, dict) or "type" not in meta: + err(path, "frontmatter has no 'type'") + continue + t = meta["type"] + if t not in types: + err(path, f"unknown type '{t}'") + continue + spec = types[t] + expected_dir = spec.get("dir", ".") + actual_dir = str(Path(rel).parent.as_posix()) + if expected_dir == ".": + if actual_dir != ".": + err(path, f"type '{t}' must live in repo root") + elif not (actual_dir == expected_dir + or actual_dir.startswith(expected_dir + "/")): + err(path, f"type '{t}' must live under {expected_dir}/") + fn_pattern = spec.get("filename") + if fn_pattern and not re.match(fn_pattern, path.name): + err(path, f"filename does not match {fn_pattern}") + check_fields(path, meta, spec, root) + if "id" in (spec.get("fields") or {}) and meta.get("id") is not None: + artifact_id = str(meta["id"]) + if not path.name.startswith(f"{artifact_id}-"): + err(path, f"id '{artifact_id}' does not match filename prefix") + key = (t, artifact_id) + if key in seen_ids: + err(path, f"duplicate {t} id '{artifact_id}' " + f"(also in {seen_ids[key].name})") + else: + seen_ids[key] = path + check_links(path, meta, link_fields, root, inbound) + check_body_links(path, body, root, inbound) + apply_rules(path, rel, meta, spec) + if t == "wiki-page" and meta.get("area") != "index": + wiki_pages.append((path, meta)) + + # link-only files: inline links are checked, frontmatter not required + already = {p.resolve() for p in artifacts} + for pattern in schema.get("scope", {}).get("link_only", []): + for path in sorted(root.glob(pattern)): + if not path.is_file() or path.resolve() in already: + continue + meta, body = parse_frontmatter(path.read_text(encoding="utf-8")) + if meta is None: + body = path.read_text(encoding="utf-8") + check_body_links(path, body, root, inbound) + + for path, _meta in wiki_pages: + if path.resolve() not in inbound: + warn(path, "orphan wiki page — nothing links to it") + + for line in errors + warnings: + print(line) + print(f"validate: {len(errors)} error(s), {len(warnings)} warning(s)") + return 1 if errors else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/aar-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/aar-template.md new file mode 100644 index 0000000..b23ff59 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/aar-template.md @@ -0,0 +1,34 @@ +--- +type: aar +status: open # open | harvested +date: YYYY-MM-DD +related: [] # design docs, issues, ADRs involved +--- + + + +# AAR: Title + +## What was planned / expected + +## What happened + + + +## Why the difference + + + +## Learnings + + + +## Actions + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/adr-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/adr-template.md new file mode 100644 index 0000000..9a9d73e --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/adr-template.md @@ -0,0 +1,37 @@ +--- +type: adr +id: "0000" +status: proposed # proposed | accepted | superseded +date: YYYY-MM-DD +supersedes: null # path to older ADR, e.g. docs/adr/0002-old.md +superseded_by: null # filled in on the OLD adr when a new one replaces it +related: [] # optional: paths to design docs / issues +--- + + + +# ADR-0000: Title + +## Context + + + +## Options Considered + + + +## Decision + + + +## Consequences + + + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/design-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/design-template.md new file mode 100644 index 0000000..32e1508 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/design-template.md @@ -0,0 +1,110 @@ +--- +type: design +status: gate-1 # gate-1 | gate-2 | gate-3 | gate-4 | gate-5 | done +date: YYYY-MM-DD +size: L # this template is for size L +related: [] # issues, ADRs spawned or read +--- + + + +# Design: Title + +## Gate 1 — Product + +**Problem.** + +**Acceptance criterion.** + +**Non-goals.** + +**Announcement.** + +**Mockups.** + +> **STOP — awaiting Gate 1 approval.** + +## Gate 2 — Architecture + +**Inputs read.** + +**System fit.** + +**Constraints.** + +**Options & trade-offs.** + +**New ADRs.** + +> **STOP — awaiting Gate 2 approval.** + +## Gate 3 — Program Design + +**Files.** + +**Signatures.** + +**Call stack.** + +**Test assertions.** + +**Boundaries — DO NOT CHANGE.** + +**Shakiest calls.** + +> **STOP — awaiting Gate 3 approval.** + +## Gate 4 — Vertical Slices + + + +### Slice 1 — Tracer bullet +- [ ] Task: … — files: … — action: … — verify: … — done: … + +**Evidence:** +**Status:** + +> **STOP — slice review.** + +### Slice 2 — … + +### Handoff + + + +## Gate 5 — Closeout (AAR) + +**Planned vs. actual.** + +**Why the difference.** + +**Learnings.** + +**Harvested.** + +**Open uncertainties.** + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/issue-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/issue-template.md new file mode 100644 index 0000000..8adbf4b --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/issue-template.md @@ -0,0 +1,29 @@ +--- +type: issue +id: "0000" +status: open # open | in-progress | done | rejected +created: YYYY-MM-DD +related: [] # design docs, ADRs, other issues +--- + + + +# Issue-0000: Title + +## Problem / Motivation + + + +## Acceptance + + + +## Notes + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/ledger-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/ledger-template.md new file mode 100644 index 0000000..98942bd --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/ledger-template.md @@ -0,0 +1,42 @@ +--- +type: ledger +date: YYYY-MM-DD +size: L # S | M | L +status: open # open | closed (closed at Gate 5) +related: [] # the design doc, the issue, whatever this run served +--- + + + +# Ledger: + +## Gates + + + +| gate | commit | approval | status | note | +|---|---|---|---|---| +| 1 | abc1234 | owner | DONE | one line, for a reader | + +## Ladder + + + +| searched | found | outcome | commit | +|---|---|---|---| +| where you looked | what was there | reused: … / built: … | abc1234 | + +## Notes + + diff --git a/docs/sources/upstream/neckbeard-v0.3.1/templates/verdict-template.md b/docs/sources/upstream/neckbeard-v0.3.1/templates/verdict-template.md new file mode 100644 index 0000000..0a60bc6 --- /dev/null +++ b/docs/sources/upstream/neckbeard-v0.3.1/templates/verdict-template.md @@ -0,0 +1,60 @@ +--- +type: verdict +date: YYYY-MM-DD +outcome: clean # clean | model-failure | framework-gap | both +judged: docs/ledger/YYYY-MM-DD-slug.md +related: [] +--- + + + +# Verdict: + +⚠️ **Written in fresh context, without the work transcript.** An agent that +judges its own session justifies rather than checks. The judge receives +`AGENTS.md`, `WORKFLOW.md`, the diff and the ledger — nothing else. A +verdict produced in the working session is void, whatever it says. + +## Deterministic pass + + + +``` +judge: … +``` + +## Rubric + + + +**1. Scope.** Does every changed file trace to the stated undertaking? + +**2. Substance.** Is the design document load-bearing — do the non-goals +exclude something a reader would otherwise expect, are the shakiest calls +real risks? + +**3. Size.** Was the declared class plausible for what the diff became? + +**4. Evidence.** Where a slice claims a result, is a run quoted? Where a +check was added, was it shown going red? + +**5. Silence.** Which rule should have left a trace here and did not? + +## Findings + + + +| bucket | rule | finding | +|---|---|---| +| model-failure | | a clear rule was not followed | +| framework-gap | | the rule is missing, unenforceable, or invisible | + +## What follows + + diff --git a/docs/verdict/template.md b/docs/verdict/template.md new file mode 100644 index 0000000..0a60bc6 --- /dev/null +++ b/docs/verdict/template.md @@ -0,0 +1,60 @@ +--- +type: verdict +date: YYYY-MM-DD +outcome: clean # clean | model-failure | framework-gap | both +judged: docs/ledger/YYYY-MM-DD-slug.md +related: [] +--- + + + +# Verdict: + +⚠️ **Written in fresh context, without the work transcript.** An agent that +judges its own session justifies rather than checks. The judge receives +`AGENTS.md`, `WORKFLOW.md`, the diff and the ledger — nothing else. A +verdict produced in the working session is void, whatever it says. + +## Deterministic pass + + + +``` +judge: … +``` + +## Rubric + + + +**1. Scope.** Does every changed file trace to the stated undertaking? + +**2. Substance.** Is the design document load-bearing — do the non-goals +exclude something a reader would otherwise expect, are the shakiest calls +real risks? + +**3. Size.** Was the declared class plausible for what the diff became? + +**4. Evidence.** Where a slice claims a result, is a run quoted? Where a +check was added, was it shown going red? + +**5. Silence.** Which rule should have left a trace here and did not? + +## Findings + + + +| bucket | rule | finding | +|---|---|---| +| model-failure | | a clear rule was not followed | +| framework-gap | | the rule is missing, unenforceable, or invisible | + +## What follows + + diff --git a/schema.yaml b/schema.yaml index b6c202b..7f8c4c6 100644 --- a/schema.yaml +++ b/schema.yaml @@ -3,8 +3,8 @@ # file. Extending the framework's metadata means editing THIS file, not code. # Agents: never invent fields or status values; propose a schema change. # -# PROJEKTERWEITERUNGEN gegenüber neckbeard v0.1.1 (Original: -# docs/sources/upstream/neckbeard-v0.1.1/schema.yaml; Design: +# PROJEKTERWEITERUNGEN gegenüber neckbeard v0.3.1 (Original: +# docs/sources/upstream/neckbeard-v0.3.1/schema.yaml; Design: # docs/design/done/2026-08-11-neckbeard-migration.md, ADR-0012/0013): # * issue: Pflichtfelder milestone (M1–M5) + priority; Status-Enum um # next/waiting erweitert; due/host/area/wartegrund/gitlab_iid; @@ -15,6 +15,8 @@ # Fehlt das Feld, ist management gemeint. # * component: neuer Typ unter docs/components/ (Dateiname = Slug). # * wiki-page: Area-Enum um "vision" erweitert. +# * project_section_marker auf unsere Marke gesetzt; ledger und +# verdict unverändert von upstream übernommen (v0.3.1). version: 1 @@ -32,11 +34,27 @@ scope: link_only: - "*.md" +# Files an adopting project holds byte-identical against its vendored +# baseline (AGENTS.md §5). They are copied into repositories that do not +# have this repo's docs/, so they must carry no repo-relative link — a +# link that resolves here and nowhere else makes the adoption path +# unfollowable. Mention an ADR by its identifier instead. +vendored: + - "AGENTS.md" + - "WORKFLOW.md" + - "CLAUDE.md" + +# An adopting project appends its own always-on rules below a marker in a +# vendored file (AGENTS.md §5). Everything from the marker on belongs to +# that project, is never copied anywhere, and may link freely — the rule +# above applies only to the upstream part. Adopters set their own marker. +project_section_marker: "" + # Frontmatter fields whose values are links. Values starting with # http://, https:// or mailto: are treated as external and only # format-checked; everything else must be a repo-root-relative path # to an existing file. -link_fields: [related, sources, supersedes, superseded_by] +link_fields: [related, sources, supersedes, superseded_by, judged] types: project: @@ -128,6 +146,32 @@ types: # The canonical slug is the filename — no second naming scheme. - slug_matches_filename + # One per session. The envelope is validated here; the gate rows and + # ladder entries in the body are outside what this engine can express + # (it has no notion of a list of records) and belong to scripts/judge.py. + # That seam is deliberate — see ADR-0010. + ledger: + dir: "docs/ledger" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, date, size, status] + fields: + date: { kind: date } + size: { enum: [S, M, L] } + status: { enum: [open, closed] } + related: { kind: links } + + # The output of a judged run. Categories are the two-bucket + # classification the harvest assessment established: a finding is either + # the model not following a clear rule, or a gap in the framework. + verdict: + dir: "docs/verdict" + filename: "^\\d{4}-\\d{2}-\\d{2}-[a-z0-9-]+\\.md$" + required: [type, date, outcome, judged] + fields: + date: { kind: date } + outcome: { enum: [clean, model-failure, framework-gap, both] } + related: { kind: links } + wiki-page: dir: "docs/wiki" filename: "^[a-z0-9-]+\\.md$" diff --git a/scripts/check_harvest.py b/scripts/check_harvest.py new file mode 100644 index 0000000..ac6b94e --- /dev/null +++ b/scripts/check_harvest.py @@ -0,0 +1,461 @@ +#!/usr/bin/env python3 +"""check_harvest.py — refuse a harvest that carries adopter specifics. + +A harvest travels from an adopting project into the framework repository. +Only generalized statements may cross: the failure class, its effect, its +cause, and the framework change they argue for. Names of the adopter, its +products, hosts, stack components, people, paths and identifiers must not. + +Three surfaces are checked, because a real leak used all of them: + * file contents + * file names + * commit messages in the range being handed over + +Fails closed (exit 1, never "clean"): + * term list missing, unreadable, or empty after stripping comments + * a revision range git cannot resolve + * a file that cannot be read as text — a harvest is prose, and a blob + nobody can read is precisely what nobody can review + +Deliberately NOT a fourth surface: the author and committer identity of +those commits. It is a repository-wide property rather than something a +harvest carries in — the framework's own history and LICENSE hold the same +name — so a run over any branch would report it every time. A check that +is permanently red reports nothing, and teaches people to skip it. Commit +identity belongs to the repository's publication decision, not to this +check; verify it there, once, and not in every harvest. + +⚠️ This check supplements human review, it does not replace it. A denylist +finds only the nouns somebody thought of, and the paragraph above names a +surface it does not look at by design. Treating a green run as proof of +absence is the same mistake that made the leak it exists to prevent. + +The term list belongs to the adopter and is never shipped with the +framework — the framework must not store the names it exists to keep out. +Two rules for writing one: + * one term per line, `#` starts a comment. Matching is case-insensitive + and respects word boundaries, so `zephyr` finds "Zephyr-Web" and + "zephyr.example" but not the middle of an unrelated English word. + Wrap a term in asterisks — `*zephyr*` — to match inside words too, + for the rare name that hides in a compound with no separator. + * a term the framework itself uses is shared vocabulary, not a secret. + Listing it only produces noise that teaches people to ignore the check. + +Scope. With `--range`, the lines that range **adds** are read, plus the +names of the files it touches and the messages of its commits. A harvest +answers for what it writes, not for what the repository already carried: +regenerating a shared index would otherwise drag every pre-existing line +of that index into the result. Without `--range`, every tracked file is +read in full — an audit of the repository, which is a different question. + +Findings print the term and a masked excerpt: enough to locate the leak, +without repeating the secret in full wherever the output ends up. + +Usage: + python scripts/check_harvest.py --terms [--range ..] [path ...] + python scripts/check_harvest.py --selftest +""" +from __future__ import annotations + +import io +import os +import re +import subprocess +import sys +import tempfile +from contextlib import redirect_stdout +from functools import lru_cache +from pathlib import Path +from typing import Iterator, NamedTuple + +MASK = "***" +COMMIT_SEP = "\x1e" +FIELD_SEP = "\x1f" + + +class HarvestError(RuntimeError): + """The check could not answer the question. That is never a pass.""" + + +class Finding(NamedTuple): + surface: str # "content" | "filename" | "commit" | "unreadable" + where: str # path, or commit sha + line: int | None # None for filenames, commit subjects, unreadable + term: str + excerpt: str + + +def git(root: Path, *args: str) -> str: + """Run git. A failure raises — it never becomes an empty result. + + ⚠️ This is the core of failing closed. An unresolvable range makes git + exit non-zero; turning that into "" would turn it into "no changes" + and therefore into a silent pass. + """ + done = subprocess.run(["git", "-C", str(root), *args], + capture_output=True, text=True) + if done.returncode != 0: + raise HarvestError(f"git {' '.join(args)}: {done.stderr.strip()[:160]}") + return done.stdout + + +def load_terms(path: Path) -> list[str]: + """Read the adopter's term list. Missing, unreadable or empty is an error.""" + try: + raw = path.read_text(encoding="utf-8") + except OSError as err: + raise HarvestError(f"term list unreadable: {path} ({err})") from err + terms = [z.strip() for z in raw.splitlines()] + terms = [z for z in terms if z and not z.startswith("#")] + if not terms: + raise HarvestError(f"term list is empty: {path}") + return terms + + +@lru_cache(maxsize=None) +def term_pattern(term: str) -> re.Pattern[str]: + """`*x*` matches inside words; a bare term respects word boundaries. + + ⚠️ Word boundaries are the default because a substring match on a + proper noun hits ordinary language: short personal names sit inside + perfectly ordinary English words — "rene" inside "serene", and the + name that forced this inside the word "authored". Measured, not + hypothesised: a substring run flagged every commit trailer here. + """ + if len(term) > 2 and term.startswith("*") and term.endswith("*"): + return re.compile(re.escape(term[1:-1]), re.I) + return re.compile(rf"(? str: + """Replace the matched term, keeping the surrounding context readable.""" + return term_pattern(term).sub(MASK, text).strip()[:120] + + +def scan_text(text: str, terms: list[str], *, surface: str, + where: str, numbered: bool = True) -> list[Finding]: + found: list[Finding] = [] + for nr, zeile in enumerate(text.splitlines() or [text], start=1): + for term in terms: + if term_pattern(term).search(zeile): + found.append(Finding(surface, where, nr if numbered else None, + term, mask(zeile, term))) + return found + + +def iter_files(root: Path, paths: list[str], + rev_range: str | None = None) -> Iterator[Path]: + """Scope: the range's own files, or every tracked file when none given. + + A harvest is checked against what it adds. Scanning the whole repository + instead surfaces pre-existing content that the harvest never touched — + noise that buries the findings that matter. + """ + if not paths: + if rev_range: + roh = git(root, "diff", "--name-only", "--diff-filter=d", rev_range) + else: + roh = git(root, "ls-files").replace("\0", "\n") + for name in roh.splitlines(): + if name and (root / name).is_file(): + yield root / name + return + for roh in paths: + p = Path(roh) + if p.is_dir(): + yield from (q for q in sorted(p.rglob("*")) if q.is_file()) + elif p.is_file(): + yield p + else: + raise HarvestError(f"path does not exist: {p}") + + +def _kurzname(root: Path, datei: Path) -> str: + """Path for the report — never a crash. + + ⚠️ A file handed in explicitly may sit outside the repository, and + `Path.relative_to` raises for those. A check that dies on a path it was + asked to read answers nothing; the reason it is not fatal is that the + only thing wanted here is a label. + """ + try: + return str(datei.relative_to(root)) + except ValueError: + return str(datei) + + +def scan_files(root: Path, paths: list[str], terms: list[str], + rev_range: str | None = None) -> list[Finding]: + found: list[Finding] = [] + for datei in iter_files(root, paths, rev_range): + rel = _kurzname(root, datei) + found += scan_text(rel, terms, surface="filename", where=rel, + numbered=False) + try: + inhalt = datei.read_text(encoding="utf-8") + except (UnicodeDecodeError, OSError): + found.append(Finding("unreadable", rel, None, "-", + "not readable as text — cannot be reviewed")) + continue + found += scan_text(inhalt, terms, surface="content", where=rel) + return found + + +def scan_diff(root: Path, rev_range: str, terms: list[str]) -> list[Finding]: + """Read the lines a range *adds*, not the files it happens to touch. + + ⚠️ Reading whole files reports content the harvest never wrote — a + generated index regenerated by the harvest carries every pre-existing + line of the repository into the result. Those belong to the repository + and its own publication decision; a harvest answers for what it adds. + """ + roh = git(root, "diff", "--unified=0", "--no-color", rev_range) + found: list[Finding] = [] + datei, nr = "", 0 + for z in roh.splitlines(): + if z.startswith("+++ "): + ziel = z[4:].strip() + datei = ziel[2:] if ziel.startswith("b/") else ziel + elif z.startswith("@@"): + treffer = re.search(r"\+(\d+)", z) + nr = int(treffer.group(1)) if treffer else 0 + elif z.startswith("Binary files") and datei and datei != "/dev/null": + found.append(Finding("unreadable", datei, None, "-", + "binary — cannot be reviewed as text")) + elif z.startswith("+") and not z.startswith("+++"): + if datei and datei != "/dev/null": + inhalt = z[1:] + for term in terms: + if term_pattern(term).search(inhalt): + found.append(Finding("content", datei, nr, term, + mask(inhalt, term))) + nr += 1 + for name in git(root, "diff", "--name-only", "--diff-filter=d", + rev_range).splitlines(): + if name: + found += scan_text(name, terms, surface="filename", where=name, + numbered=False) + return found + + +def scan_commits(root: Path, rev_range: str, terms: list[str]) -> list[Finding]: + roh = git(root, "log", f"--format=%H{FIELD_SEP}%B{COMMIT_SEP}", rev_range) + found: list[Finding] = [] + for block in roh.split(COMMIT_SEP): + block = block.strip("\n") + if FIELD_SEP not in block: + continue + sha, nachricht = block.split(FIELD_SEP, 1) + found += scan_text(nachricht, terms, surface="commit", + where=sha[:12], numbered=False) + return found + + +def report(findings: list[Finding]) -> None: + for surface in ("content", "filename", "commit", "unreadable"): + teil = [f for f in findings if f.surface == surface] + if not teil: + continue + print(f"\n {surface} ({len(teil)}):") + for f in teil: + ort = f"{f.where}:{f.line}" if f.line else f.where + print(f" {ort} — term {f.term!r}") + print(f" {f.excerpt}") + + +def selftest() -> int: + """Positive and negative controls. A check only ever seen green is a guess.""" + fehler: list[str] = [] + + def pruefe(name: str, bedingung: bool) -> None: + print(f" {'ok ' if bedingung else 'FAIL'} {name}") + if not bedingung: + fehler.append(name) + + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + git(root, "init", "-q", "-b", "main") + git(root, "config", "user.email", "selftest@invalid") + git(root, "config", "user.name", "selftest") + liste = root / "terms.txt" + liste.write_text("# comment\nzephyr\nrene\n*brand*\n\n", encoding="utf-8") + terms = load_terms(liste) + + (root / "clean.md").write_text("a generalized statement\n", encoding="utf-8") + (root / "body.md").write_text("one\ntwo Zephyr three\n", encoding="utf-8") + (root / "zephyr-notes.md").write_text("clean body\n", encoding="utf-8") + # "serene" contains the name "rene" — the false-positive class that + # forced word boundaries. "xbrandy" is the opposite case, opted + # into with *…*. + (root / "english.md").write_text("a serene afternoon\n", encoding="utf-8") + (root / "touched.md").write_text("old line says Zephyr\n", encoding="utf-8") + (root / "compound.md").write_text("xbrandy\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-q", "-m", "initial, no secret here") + (root / "clean.md").write_text("still generalized\n", encoding="utf-8") + (root / "late.md").write_text("added later, says Zephyr\n", encoding="utf-8") + # Touched by the range, but the term sits in a line the range did + # not write — the generated-index case, in miniature. + # Line 1 is pre-existing and must stay silent; line 2 is added by + # the range and must be reported *at line 2* — which also proves the + # hunk header is read instead of assumed. + (root / "touched.md").write_text( + "old line says Zephyr\nadded line also says Zephyr\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-q", "-m", "mentions Zephyr in the message") + + treffer = scan_files(root, [], terms) + inhalt = [f for f in treffer if f.surface == "content"] + namen = [f for f in treffer if f.surface == "filename"] + commits = scan_commits(root, "HEAD~1..HEAD", terms) + + pruefe("1 term in a file body is found, with its line number", + any(f.where == "body.md" and f.line == 2 for f in inhalt)) + pruefe("2 term in a filename is found while the body is clean", + any(f.where == "zephyr-notes.md" for f in namen) + and not any(f.where == "zephyr-notes.md" for f in inhalt)) + pruefe("3 term in a commit message is found while the tree is clean", + len(commits) == 1) + # The listed term is lower case; the file says "Zephyr". The match + # is only proof of case-insensitivity if the source really differs. + pruefe("4 a lower-case term matches a capitalized occurrence", + "Zephyr" in (root / "body.md").read_text(encoding="utf-8") + and any(f.term == "zephyr" and f.where == "body.md" + for f in inhalt)) + pruefe("5 the excerpt masks the term instead of repeating it", + all(MASK in f.excerpt for f in inhalt)) + + leer = root / "empty.txt" + leer.write_text("# only comments\n", encoding="utf-8") + pruefe("6 an empty term list raises instead of passing", + _raises(lambda: load_terms(leer))) + pruefe("7 a missing term list raises instead of passing", + _raises(lambda: load_terms(root / "nope.txt"))) + pruefe("8 an unresolvable range raises instead of reporting nothing", + _raises(lambda: scan_commits(root, "nosuchref..HEAD", terms))) + pruefe("9 clean input against a non-empty list finds nothing", + scan_files(root, [str(root / "clean.md")], terms) == []) + pruefe("10 a bare term does not match inside an unrelated word", + not any(f.where == "english.md" for f in treffer)) + pruefe("11 a *term* does match inside a word, when opted into", + any(f.where == "compound.md" and f.term == "*brand*" + for f in treffer)) + # Sharp in both directions: the range's own file must be found, and + # the older ones must not. An assertion that only checks for "nothing" + # would also pass if the scoping read no files at all. + bereich = scan_diff(root, "HEAD~1..HEAD", terms) + pruefe("12 with a range, exactly that range's own additions are read", + {f.where for f in bereich} == {"late.md", "touched.md"} + and any(f.where == "body.md" for f in treffer)) + beruehrt = [f for f in bereich if f.where == "touched.md"] + pruefe("13 a pre-existing line in a touched file is not reported", + len(beruehrt) == 1) + pruefe("14 an added line is reported at its real line number", + bool(beruehrt) and beruehrt[0].line == 2) + + # Everything above tests the scanners directly, which leaves the + # dispatch in main() unproven: it could call the whole-file reader, + # or skip commit messages entirely, and every assertion above would + # still pass. Two discriminators that only the wiring can satisfy: + # whole-file mode reports touched.md at line 1 (the pre-existing + # line) where the diff reader reports line 2, and a missing + # scan_commits call removes the commit section from the output. + zuvor = Path.cwd() + puffer = io.StringIO() + try: + os.chdir(root) + with redirect_stdout(puffer): + main(["--terms", str(liste), "--range", "HEAD~1..HEAD"]) + finally: + os.chdir(zuvor) + ausgabe = puffer.getvalue() + pruefe("15 main() routes a range to the diff reader, not whole files", + "touched.md:2" in ausgabe and "touched.md:1" not in ausgabe) + pruefe("16 main() actually reads the commit messages of the range", + "commit (" in ausgabe) + + # A path handed in explicitly may sit outside the repository. That + # crashed with a ValueError until a positive control tried it. + aussen = Path(tempfile.gettempdir()) / "check_harvest_outside.md" + aussen.write_text("mentions Zephyr\n", encoding="utf-8") + try: + draussen = scan_files(root, [str(aussen)], terms) + except ValueError: + draussen = [] + finally: + aussen.unlink(missing_ok=True) + pruefe("17 a path outside the repository is reported, not fatal", + len(draussen) == 1 and draussen[0].term == "zephyr") + + print() + if fehler: + print(f"selftest: {len(fehler)} assertion(s) failed") + return 1 + print("selftest: 17 assertions passed") + return 0 + + +def _raises(fn) -> bool: + try: + fn() + except HarvestError: + return True + return False + + +def main(argv: list[str]) -> int: + if "--selftest" in argv: + return selftest() + + terms_pfad: Path | None = None + rev_range: str | None = None + paths: list[str] = [] + i = 0 + while i < len(argv): + if argv[i] == "--terms" and i + 1 < len(argv): + terms_pfad, i = Path(argv[i + 1]), i + 2 + elif argv[i] == "--range" and i + 1 < len(argv): + rev_range, i = argv[i + 1], i + 2 + elif argv[i].startswith("--"): + print(f"unknown option: {argv[i]}") + return 2 + else: + paths.append(argv[i]) + i += 1 + + if terms_pfad is None: + print(__doc__.strip().splitlines()[-3].strip()) + print("check_harvest: --terms is required. Unchecked is not clean.") + return 2 + + root = Path.cwd() + try: + terms = load_terms(terms_pfad) + if rev_range and not paths: + findings = scan_diff(root, rev_range, terms) + findings += scan_commits(root, rev_range, terms) + else: + findings = scan_files(root, paths, terms) + if rev_range: + findings += scan_commits(root, rev_range, terms) + except HarvestError as err: + print(f"check_harvest: {err}") + print("The question could not be answered — that is not a pass.") + return 1 + + umfang = f"{len(terms)} term(s)" + if rev_range: + umfang += f", commits {rev_range}" + if findings: + print(f"check_harvest: {len(findings)} finding(s) — {umfang}") + report(findings) + print("\nA harvest carries the failure class, its effect and its cause —") + print("never the adopter's names. Generalize, then run this again.") + return 1 + print(f"check_harvest: no findings — {umfang}") + print("⚠️ A denylist finds only what someone listed. Human review still applies.") + return 0 + + +if __name__ == "__main__": + sys.exit(main(sys.argv[1:])) diff --git a/scripts/judge.py b/scripts/judge.py new file mode 100644 index 0000000..25f5b18 --- /dev/null +++ b/scripts/judge.py @@ -0,0 +1,564 @@ +#!/usr/bin/env python3 +"""judge.py — read a run's trace against the workflow. Build form A. + +A judge cannot examine behaviour, only traces. This one reads a session +ledger (`docs/ledger/`) against `WORKFLOW.md`'s rules and against git, +and reports where the two disagree. It is deterministic, stdlib-only and +model-agnostic; the inferential half of judging is build form B, a +separate ritual in fresh context (see WORKFLOW.md, "Judging a run"). + +What it checks + * gate rows are ordered, without duplicates, and complete for the + declared size class — S owes no gates, and is not judged as if it did + * every status comes from the four-value vocabulary + * a gate that has a successor carries an approval + * every named commit resolves, is an ancestor of HEAD, and the commits + run in the same order as the gates they belong to + * the ladder section exists, and each entry names what was searched, + what was found, and an outcome beginning `reused:` or `built:` + +The commit checks matter more than the rest: they are the only part that +compares the ledger against evidence the ledger's author did not write. + +What it deliberately does not do + * judge a session that has no ledger. History cannot be instrumented + after the fact, and a check that reports every past session forever is + a check nobody reads. + * judge quality. Whether a design document is real or filler is form B's + question, and no script can answer it. + * resist tampering. The ledger is written by the agent it describes. + The threat model is drift, not sabotage — ADR-0010 records this as an + explicit non-goal, because a judge that suggests otherwise is worse + than none. + +Coverage + `--coverage` reports which rules of the framework can be observed at all + and by what — a script, the ledger, or nothing. Rules at "nothing" are + either unobservable or inert, and both are worth knowing. This is a + report, never a gate: it must not turn red, or it becomes the standing + finding nobody can fix. + +Fails closed (exit 1, never "clean"): a ledger that cannot be read or +parsed, and a commit git cannot resolve. + +Usage: + python scripts/judge.py --ledger docs/ledger/.md [--root .] + python scripts/judge.py --coverage + python scripts/judge.py --selftest +""" +from __future__ import annotations + +import re +import subprocess +import sys +import tempfile +from pathlib import Path +from typing import NamedTuple + +STATUSES = ("DONE", "DONE_WITH_CONCERNS", "NEEDS_CONTEXT", "BLOCKED") +GATES_FOR_SIZE = {"S": set(), "M": set(), "L": {"1", "2", "3", "4", "5"}} +OUTCOME_PREFIX = ("reused:", "built:") +PENDING = "pending" +PLACEHOLDERS = {"", "-", "—", "n/a", "none", "tbd", "todo", "?"} + + +class JudgeError(RuntimeError): + """The check could not answer the question. Never a pass.""" + + +class Row(NamedTuple): + cells: list[str] + + def get(self, i: int) -> str: + return self.cells[i].strip() if i < len(self.cells) else "" + + +class Finding(NamedTuple): + bucket: str # "model-failure" | "framework-gap" + rule: str + detail: str + + +def git(root: Path, *args: str) -> str: + """Run git, or raise. A failure is never an empty result.""" + proc = subprocess.run(["git", "-C", str(root), *args], + capture_output=True, text=True) + if proc.returncode != 0: + raise JudgeError( + f"git {' '.join(args)} failed: {proc.stderr.strip() or 'no output'}") + return proc.stdout + + +# -------------------------------------------------------------------------- +# Reading the ledger. Frontmatter is read with re rather than PyYAML: the +# envelope is already validated by validate.py against schema.yaml, so this +# only needs the two or three scalars it acts on — and adopters copying one +# file should not inherit a dependency. +# -------------------------------------------------------------------------- + +def parse_ledger(path: Path) -> tuple[dict, list[Row], list[Row]]: + try: + text = path.read_text(encoding="utf-8") + except OSError as exc: + raise JudgeError(f"cannot read ledger {path}: {exc}") from exc + + match = re.match(r"^---\n(.*?)\n---\n", text, re.S) + if not match: + raise JudgeError(f"{path}: no frontmatter block") + meta: dict[str, object] = {} + current: str | None = None + for line in match.group(1).splitlines(): + item = re.match(r"^\s+-\s*(.+)$", line) + if item and current: + meta.setdefault(current, []) + if isinstance(meta[current], list): + meta[current].append(item.group(1).strip().strip('"')) + continue + kv = re.match(r"^([a-z_]+):\s*(.*)$", line) + if kv: + value = kv.group(2).strip().strip('"') + current = kv.group(1) + meta[current] = value if value not in ("", "[]") else [] + for field in ("type", "size", "status"): + if not meta.get(field): + raise JudgeError(f"{path}: frontmatter is missing '{field}'") + if meta["type"] != "ledger": + raise JudgeError(f"{path}: type is '{meta['type']}', not 'ledger'") + + body = text[match.end():] + return meta, _table(body, "Gates"), _table(body, "Ladder") + + +def _table(body: str, heading: str) -> list[Row]: + """Rows of the Markdown table under '## ', header excluded.""" + section = re.search(rf"^##\s+{heading}\s*$(.*?)(?=^##\s|\Z)", + body, re.S | re.M) + if not section: + return [] + rows: list[Row] = [] + for line in section.group(1).splitlines(): + line = line.strip() + if not line.startswith("|"): + continue + cells = [c.strip() for c in line.strip("|").split("|")] + if all(re.fullmatch(r":?-{2,}:?", c) for c in cells): + continue # the ---|--- separator + rows.append(Row(cells)) + return rows[1:] if rows else rows # drop the header row + + +# -------------------------------------------------------------------------- +# The checks +# -------------------------------------------------------------------------- + +def check_gate_order(rows: list[Row], size: str, closed: bool) -> list[Finding]: + out: list[Finding] = [] + seen = [r.get(0) for r in rows] + numbers = [s for s in seen if s.isdigit()] + if len(set(numbers)) != len(numbers): + out.append(Finding("model-failure", "workflow/gate-order", + f"a gate is recorded twice: {numbers}")) + if numbers != sorted(numbers, key=int): + out.append(Finding("model-failure", "workflow/gate-order", + f"gates are out of order: {numbers}")) + owed = GATES_FOR_SIZE.get(size.upper(), set()) + if closed: + missing = sorted(owed - set(numbers)) + if missing: + out.append(Finding("model-failure", "workflow/gates-for-size", + f"size {size} owes gates {missing}, not recorded")) + return out + + +def check_status_vocabulary(rows: list[Row]) -> list[Finding]: + out: list[Finding] = [] + for row in rows: + status = row.get(3) + if status and status not in STATUSES: + out.append(Finding("model-failure", "agents/status-vocabulary", + f"gate {row.get(0)}: '{status}' is not one of " + f"{', '.join(STATUSES)}")) + return out + + +def check_approvals(rows: list[Row]) -> list[Finding]: + """A gate with a successor must carry an approval.""" + out: list[Finding] = [] + for i, row in enumerate(rows[:-1]): + if row.get(2).lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "workflow/stop-before-next-gate", + f"gate {row.get(0)} has no approval, but gate " + f"{rows[i + 1].get(0)} was started")) + return out + + +def check_commits(root: Path, rows: list[Row]) -> list[Finding]: + """The only part measured against evidence the author did not write.""" + out: list[Finding] = [] + order: list[tuple[str, int]] = [] + for row in rows: + sha = row.get(1) + if not sha or sha.lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "ledger/commit-required", + f"gate {row.get(0)} names no commit")) + continue + try: + full = git(root, "rev-parse", "--verify", f"{sha}^{{commit}}").strip() + except JudgeError as exc: + raise JudgeError( + f"gate {row.get(0)} names commit {sha}, which does not " + f"resolve — unchecked is not passed ({exc})") from exc + ancestor = subprocess.run( + ["git", "-C", str(root), "merge-base", "--is-ancestor", full, "HEAD"], + capture_output=True, text=True) + if ancestor.returncode != 0: + out.append(Finding("model-failure", "ledger/commit-reachable", + f"gate {row.get(0)}: commit {sha} is not an " + f"ancestor of HEAD")) + continue + depth = int(git(root, "rev-list", "--count", f"{full}..HEAD").strip()) + order.append((row.get(0), depth)) + + ranked = [gate for gate, _ in sorted(order, key=lambda p: -p[1])] + stated = [gate for gate, _ in order] + if ranked != stated: + out.append(Finding("framework-gap", "ledger/commit-order", + f"the commits run in the order {ranked}, the gates " + f"claim {stated}")) + return out + + +def check_design_doc(meta: dict, root: Path, closed: bool) -> list[Finding]: + """Size L owes a design document. The coverage table claimed this + before anything checked it — the inventory drift its own entry warns + about, caught on the first read.""" + if str(meta.get("size", "")).upper() != "L" or not closed: + return [] + related = meta.get("related") or [] + if isinstance(related, str): + related = [related] + if any(r.startswith("docs/design/") for r in related): + return [] + return [Finding("model-failure", "workflow/design-doc-for-L", + "size L, but the ledger names no design document in " + "'related'")] + + +def check_ladder(rows: list[Row], closed: bool) -> list[Finding]: + out: list[Finding] = [] + if not rows: + out.append(Finding("framework-gap", "agents/ponytail-ladder", + "the Ladder section is empty — the gate was not " + "passed. A session that built nothing says so.")) + return out + for i, row in enumerate(rows, start=1): + searched, found, outcome = row.get(0), row.get(1), row.get(2) + if searched.lower() in PLACEHOLDERS or found.lower() in PLACEHOLDERS: + out.append(Finding("model-failure", "agents/ponytail-ladder", + f"ladder entry {i} names no concrete candidate")) + if not outcome.lower().startswith(OUTCOME_PREFIX): + out.append(Finding("model-failure", "agents/ponytail-ladder", + f"ladder entry {i}: outcome must begin " + f"'reused:' or 'built:', got '{outcome[:30]}'")) + if closed and row.get(3).lower() in PLACEHOLDERS | {PENDING}: + out.append(Finding("model-failure", "ledger/commit-required", + f"ladder entry {i} is still 'pending' in a " + f"closed ledger")) + return out + + +def judge(root: Path, ledger: Path) -> list[Finding]: + meta, gates, ladder = parse_ledger(ledger) + closed = meta.get("status") == "closed" + size = meta.get("size", "L") + findings = check_gate_order(gates, size, closed) + findings += check_status_vocabulary(gates) + findings += check_approvals(gates) + findings += check_commits(root, gates) + findings += check_design_doc(meta, root, closed) + findings += check_ladder(ladder, closed) + return findings + + +# -------------------------------------------------------------------------- +# Rule coverage. A report, never a gate. +# +# The inventory is maintained by hand because the rules live in prose and +# nothing derives them mechanically. That is a known weakness: it will drift +# from AGENTS.md and WORKFLOW.md unless someone updates it, and nothing +# contradicts that drift today. +# -------------------------------------------------------------------------- + +COVERAGE = [ + # (rule, where it is written, what observes it) + ("artifact frontmatter and enums", "AGENTS.md §5", "validate.py"), + ("link targets are files, never directories", "AGENTS.md §5", "validate.py"), + ("STATUS.md is generated, never hand-edited", "AGENTS.md §4", "gen_status.py --check"), + ("accepted ADRs are never edited", "AGENTS.md §4", "check_locked.py"), + ("docs/sources is immutable", "AGENTS.md §5", "check_locked.py"), + ("a harvest carries no adopter specifics", "WORKFLOW.md", "check_harvest.py"), + ("PROJECT.md answers Gate 0", "AGENTS.md §2", "validate.py"), + ("a design doc is done iff it sits in done/", "WORKFLOW.md", "validate.py"), + ("the ponytail ladder was walked", "AGENTS.md §1", "ledger"), + ("a size class is declared for the run", "AGENTS.md §3", "ledger"), + ("gates run in order", "WORKFLOW.md", "ledger"), + ("a gate is approved before the next begins", "WORKFLOW.md", "ledger"), + ("every slice reports one of four statuses", "AGENTS.md §1", "ledger"), + ("size L owes a design document", "WORKFLOW.md", "ledger"), + ("simplicity: the minimum that solves it", "AGENTS.md §1", None), + ("surgical changes: every line traces to the request", "AGENTS.md §1", None), + ("slice 1 is a tracer bullet", "WORKFLOW.md", None), + ("a check owes proof it can fail", "WORKFLOW.md", None), + ("a delivering artifact owes one real result", "WORKFLOW.md", None), + ("a closeout names what it made false", "WORKFLOW.md", None), + ("reproduce before fixing", "WORKFLOW.md", None), + ("incidents get an AAR", "WORKFLOW.md", None), + ("a permanent exception owes an ADR", "AGENTS.md §1", None), +] + + +def coverage_report() -> int: + by_script = [r for r in COVERAGE if r[2] and r[2] != "ledger"] + by_ledger = [r for r in COVERAGE if r[2] == "ledger"] + unobserved = [r for r in COVERAGE if r[2] is None] + total = len(COVERAGE) + + print(f"rule coverage: {total} rules inventoried\n") + for title, group in (("observed by a script", by_script), + ("observed by the ledger", by_ledger), + ("not observable today", unobserved)): + print(f" {title} — {len(group)}") + for rule, where, how in group: + suffix = f" [{how}]" if how else "" + print(f" {rule} ({where}){suffix}") + print() + observed = len(by_script) + len(by_ledger) + print(f" {observed}/{total} observable, {len(unobserved)} decoration " + f"until they gain a duty to leave a trace.") + print("\nThis is a report, not a gate. A rule at zero coverage is either") + print("unobservable or inert; neither is a violation of anything.") + return 0 + + +# -------------------------------------------------------------------------- +# Positive control. A gate that has only ever been seen green is a +# hypothesis — WORKFLOW.md, Gate 4. +# -------------------------------------------------------------------------- + +LEDGER_HEAD = """--- +type: ledger +date: 2026-08-21 +size: {size} +status: {status} +related: +{related}--- + +## Gates + +| gate | commit | approval | status | note | +|---|---|---|---|---| +{gates} + +## Ladder + +| searched | found | outcome | commit | +|---|---|---|---| +{ladder} +""" + + +def _write(path: Path, *, size="L", status="closed", gates="", ladder="", + related=' - "docs/design/d.md"\n'): + path.write_text(LEDGER_HEAD.format(size=size, status=status, + gates=gates, ladder=ladder, + related=related), + encoding="utf-8") + + +def selftest() -> int: + passed, failed = 0, [] + + def check_that(name: str, condition: bool) -> None: + nonlocal passed + if condition: + passed += 1 + print(f" ok {passed + len(failed)} {name}") + else: + failed.append(name) + print(f" FAIL {passed + len(failed)} {name}") + + with tempfile.TemporaryDirectory() as tmp: + root = Path(tmp) + git(root, "init", "-q", ".") + git(root, "config", "user.email", "selftest@invalid") + git(root, "config", "user.name", "selftest") + shas = [] + for i in range(5): + (root / f"f{i}.txt").write_text(f"{i}\n", encoding="utf-8") + git(root, "add", "-A") + git(root, "commit", "-qm", f"c{i}") + shas.append(git(root, "rev-parse", "--short", "HEAD").strip()) + + led = root / "ledger.md" + good_ladder = ("| validate.py | a generic engine | reused: schema entry " + f"only | {shas[0]} |") + rows = lambda gates: "\n".join( + f"| {g} | {shas[i]} | owner | DONE | n |" for i, g in enumerate(gates)) + + _write(led, gates=rows("12345"), ladder=good_ladder) + check_that("a clean ledger produces no findings", + judge(root, led) == []) + + _write(led, gates=rows("13245"), ladder=good_ladder) + check_that("a gate out of order is reported", + any(f.rule == "workflow/gate-order" for f in judge(root, led))) + + _write(led, gates=rows("11234"), ladder=good_ladder) + check_that("a duplicated gate is reported", + any("twice" in f.detail for f in judge(root, led))) + + _write(led, gates=rows("123"), ladder=good_ladder) + check_that("a closed size-L ledger missing gates 4 and 5 is reported", + any(f.rule == "workflow/gates-for-size" + for f in judge(root, led))) + + _write(led, size="S", gates="", ladder=good_ladder) + check_that("size S is not judged against gates it does not owe", + not any(f.rule == "workflow/gates-for-size" + for f in judge(root, led))) + + _write(led, gates=f"| 1 | {shas[0]} | owner | FINISHED | n |", + ladder=good_ladder) + check_that("a status outside the vocabulary is reported", + any(f.rule == "agents/status-vocabulary" + for f in judge(root, led))) + + _write(led, gates=(f"| 1 | {shas[0]} | | DONE | n |\n" + f"| 2 | {shas[1]} | owner | DONE | n |"), + ladder=good_ladder) + check_that("a gate whose successor began without approval is reported", + any(f.rule == "workflow/stop-before-next-gate" + for f in judge(root, led))) + + _write(led, gates="| 1 | deadbee | owner | DONE | n |", + ladder=good_ladder) + check_that("a commit that does not resolve raises instead of passing", + _raises(lambda: judge(root, led))) + + git(root, "checkout", "-q", "-b", "side", shas[0]) + (root / "side.txt").write_text("x\n", encoding="utf-8") + git(root, "add", "-A"); git(root, "commit", "-qm", "side") + off = git(root, "rev-parse", "--short", "HEAD").strip() + git(root, "checkout", "-q", "main" if _has(root, "main") else "master") + _write(led, gates=f"| 1 | {off} | owner | DONE | n |", ladder=good_ladder) + check_that("a commit that is not an ancestor of HEAD is reported", + any(f.rule == "ledger/commit-reachable" + for f in judge(root, led))) + + _write(led, gates=(f"| 1 | {shas[3]} | owner | DONE | n |\n" + f"| 2 | {shas[1]} | owner | DONE | n |"), + ladder=good_ladder) + check_that("commits running in a different order than the gates is reported", + any(f.rule == "ledger/commit-order" for f in judge(root, led))) + + _write(led, gates=rows("12345"), ladder="") + check_that("an empty ladder section is reported as a gate not passed", + any(f.rule == "agents/ponytail-ladder" + for f in judge(root, led))) + + _write(led, gates=rows("12345"), + ladder=f"| - | - | reused: something | {shas[0]} |") + check_that("a ladder entry naming no concrete candidate is reported", + any("concrete candidate" in f.detail for f in judge(root, led))) + + _write(led, gates=rows("12345"), + ladder=f"| validate.py | an engine | it was fine | {shas[0]} |") + check_that("a ladder outcome without reused:/built: is reported", + any("must begin" in f.detail for f in judge(root, led))) + + (root / "broken.md").write_text("no frontmatter here\n", encoding="utf-8") + check_that("an unreadable ledger raises instead of reporting nothing", + _raises(lambda: judge(root, root / "broken.md"))) + + _write(led, gates=rows("12345"), ladder=good_ladder) + check_that("main() exits 0 on a clean ledger", + main(["--ledger", str(led), "--root", str(root)]) == 0) + + _write(led, gates=rows("12345"), ladder=good_ladder, related="") + check_that("a closed size-L ledger naming no design document is reported", + any(f.rule == "workflow/design-doc-for-L" + for f in judge(root, led))) + + _write(led, gates=rows("13245"), ladder=good_ladder) + check_that("main() exits 1 on a real violation", + main(["--ledger", str(led), "--root", str(root)]) == 1) + + total = passed + len(failed) + print() + if failed: + print(f"selftest: {len(failed)} of {total} assertions FAILED") + for name in failed: + print(f" - {name}") + return 1 + print(f"selftest: {total} assertions passed") + return 0 + + +def _has(root: Path, branch: str) -> bool: + return subprocess.run(["git", "-C", str(root), "rev-parse", "--verify", branch], + capture_output=True).returncode == 0 + + +def _raises(fn) -> bool: + try: + fn() + except JudgeError: + return True + except Exception: + return False + return False + + +def report(findings: list[Finding]) -> None: + if not findings: + print("judge: no findings") + return + print(f"judge: {len(findings)} finding(s)\n") + for bucket in ("framework-gap", "model-failure"): + group = [f for f in findings if f.bucket == bucket] + if not group: + continue + print(f" {bucket} ({len(group)}):") + for f in group: + print(f" [{f.rule}] {f.detail}") + print() + print("A finding is the model not following a clear rule, or a gap in the") + print("framework. Deciding which is the point — see ADR-0010.") + + +def main(argv: list[str] | None = None) -> int: + args = list(sys.argv[1:] if argv is None else argv) + if "--selftest" in args: + return selftest() + if "--coverage" in args: + return coverage_report() + + if "--ledger" not in args: + print("judge: --ledger is required (or --coverage / --selftest)", + file=sys.stderr) + return 1 + ledger = Path(args[args.index("--ledger") + 1]) + root = Path(args[args.index("--root") + 1]) if "--root" in args else Path(".") + + try: + findings = judge(root.resolve(), ledger) + except JudgeError as exc: + print(f"judge: {exc}", file=sys.stderr) + print("unchecked is not passed.", file=sys.stderr) + return 1 + report(findings) + return 1 if findings else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/pruefe_upstream_drift.py b/scripts/pruefe_upstream_drift.py index 6152cef..81a9301 100644 --- a/scripts/pruefe_upstream_drift.py +++ b/scripts/pruefe_upstream_drift.py @@ -4,7 +4,7 @@ Schützt die übernommenen Framework-Dateien vor stillem Umschreiben (Entscheidung 8 im Design 2026-08-11, Frage von sorb: „wird die AGENTS.md ggf. durch Agenten umgeschrieben?"). Die Baseline liegt unter -docs/sources/upstream/neckbeard-v0.1.1/ (siehe HERKUNFT.md dort); ein +docs/sources/upstream/neckbeard-v0.3.1/ (siehe HERKUNFT.md dort); ein Framework-Upgrade aktualisiert Baseline und Arbeitskopie im selben, bewussten Commit. @@ -22,7 +22,7 @@ from __future__ import annotations import sys from pathlib import Path -BASELINE = "docs/sources/upstream/neckbeard-v0.1.1" +BASELINE = "docs/sources/upstream/neckbeard-v0.3.1" MARKE = "" # (Arbeitskopie, Baseline-Datei) — byte-identisch @@ -33,6 +33,8 @@ PAARE = [ ("docs/design/template.md", "templates/design-template.md"), ("docs/aar/template.md", "templates/aar-template.md"), ("docs/issues/template.md", "templates/issue-template.md"), + ("docs/ledger/template.md", "templates/ledger-template.md"), + ("docs/verdict/template.md", "templates/verdict-template.md"), ] diff --git a/scripts/validate.py b/scripts/validate.py index d411618..1ef89dc 100644 --- a/scripts/validate.py +++ b/scripts/validate.py @@ -161,6 +161,32 @@ def check_body_links(path: Path, body: str, root: Path, inbound: set) -> None: inbound.add(target) +def check_vendored_portable(root: Path, vendored: list, marker: str) -> None: + """Files adopters hold byte-identical must carry no repo-relative link. + + They are copied into repositories without this repo's docs/, so such a + link resolves here and nowhere else — and the adoption path in + AGENTS.md §5 asks for a byte-for-byte copy, which makes the dead link + the adopter's problem and unfixable without breaking the copy. + """ + for rel in vendored or []: + path = root / rel + if not path.is_file(): + continue + body = path.read_text(encoding="utf-8") + if marker and marker in body: + # Everything below the marker is the adopting project's own + # section: never copied elsewhere, so its links are fine. + body = body.split(marker, 1)[0] + body = re.sub(r"```.*?```", "", body, flags=re.S) + body = re.sub(r"`[^`\n]*`", "", body) + for link in (m.group(1) for m in INLINE_LINK_RE.finditer(body)): + if link.startswith(EXTERNAL_PREFIXES) or link.startswith("#"): + continue + err(path, f"vendored file carries a repo-relative link: {link} " + f"— name the target instead of linking to it") + + def apply_rules(path: Path, rel: str, meta: dict, spec: dict) -> None: for rule in spec.get("rules", []): if rule == "superseded_requires_pointer": @@ -184,6 +210,9 @@ def main() -> int: link_fields = schema.get("link_fields", []) types = schema.get("types", {}) inbound: set = set() + + check_vendored_portable(root, schema.get("vendored", []), + schema.get("project_section_marker", "")) wiki_pages: list[tuple[Path, dict]] = [] # Root documents: inline links must resolve; no frontmatter required.