# Argus extraction v1 — independent checkpoint review, EVIDENCE / DISCIPLINE / SECURITY

Reviewer: second of two, working separately. Branch `ops/argus-extraction-v1-red`, head `8c2fdda`, repository
`D:\fileStorage\repos\GOTT.Apollo`. Reviewed the committed state only. The other session's commits
(`c54e198`, `6b12c43`, `c576fdd`, `b984afa`, `e0f40bb`) and its uncommitted `src/Sibyla.Web/**` /
`tests/Sibyla.Tests.Browser/**` work are out of scope. Date 2026-09-08.

## Verdict

**REVISE — Critical 1 / High 4 / Medium 6 / Low 2.**

The build is not dishonest. I looked hard for a tuned measure and did not find one: every relaxation I could
test is documented, dated, attributed, and demonstrably conservative or neutral on the scored cells, and the
one scoring rule invented after the answers were read (S-12) is corroborated by an oracle written at RED,
before any answer existed. What fails is not the honesty of the record but the **integrity of the one control
the whole measurement rests on**. By the spec's own design (§5.7) sittings 1 and 2 are *tuning* sittings; the
reserve is the only blind measurement in the slice. That set was read twice, its keys and flags were amended
after its first answers had been read, and the first run's evidence was destroyed. On top of that, the gate
arithmetic the reports headline is not the gate the owner ruled, and the blind reserve rates are stated
nowhere.

---

## Findings

| id | sev | claim audited | evidence | what is wrong | what would fix it |
|---|---|---|---|---|---|
| **E-1** | **Critical** | "3, reserve, held out: 10 documents, movements 156/156 = 100 %, no miss on any cell" — a held-out control. §5.6: the reserve is "the **third, single sitting** … under the frozen tuple"; S-10: "The reserve set (§5.6) is **run once**". | Two reserve runs exist: `evidence/extract-v2/score-20260907-1613-dff90f9-a558523.md` (answers paid 16:13, scored 19:09) and `…score-20260907-2250-e2f02e3-a558523.md` (`run.json`: `{"set":"Reserve","startedAt":"2026-09-07T22:50:48.5323442Z"}`, ten documents `answered`, 49–138 s each). `docs/apollo-argus-extraction-v1-key-appeals-260907.md` lines 145–152 state the re-run. `ls -d tests/Sibyla.Tests.Argus/evidence/extract-v2/score-*/` returns three folders; the `…-1613-…` run folder is absent entirely. Between the two runs, `41c5081` added three flag rows (`keyLinesSummarised` on I26070007 and I26080031, `dateDuePrinted:false` on I26010042) and one correction row (I26070026 `documentId`) **to reserve documents, on the strength of having read their answers**, plus a new scoring rule S-12. | The held-out control does not hold in substance. The set was read twice; the reported result is the second read; its key and flags were fitted to the first read's answers; and the blind read's answers and traces no longer exist, so the two cannot be compared answer-for-answer. §5.6's seal and S-10's "run once" are both broken. Disclosure in one document does not restore a control. | The reserve is spent. Either (a) treat the 100 % as *not* a held-out result and say so in the spec, the appeals sheet and any release note, resting the checkpoint on sittings 1–2 plus the pre-ruling reserve matrix; or (b) export a fresh reserve of comparable shape under the same frozen tuple, keep its keys sealed, and run it once. Do not carry "held out, 100 %, no miss" forward as it stands. |
| **E-2** | **High** | Each report's "Per kind" table — `Header 427/429 = 99.53 %`, `Lines 322/327 = 98.47 %`, `Movements 1334/1389 = 96.04 %` — presented as the S-10 gate. | `tools/extraction-bench/Sibyla.Tools.ExtractionBench/Scorer.cs:477` `var gated = fields.Where(f => f.FloorApplies)…` and `:481` `cells.Where(c => c.Kind == kind && gated.Contains(...))`: **every field with fewer than 20 scored cells is dropped from its kind's tally**, not merely exempted from the floor. Recomputed from the reports' own `cells` arrays: sitting 1 Header 445/447 (report 427/429), Lines 342/347 (report 322/327); statements Movements 1349/1404 (report 1334/1389); reserve Movements 158/158 (report 156/156). The per-field tables in the same reports sum to 447 / 347 / 1404 / 158 — they do not reconcile with the per-kind tables and no line explains the gap. | Q-EX-7 (ruled by the owner in D-EX-4) defines the kinds as "invoices+credit notes header cells; line cells; **movement cells including the statement cells** — three gates, each ≥ 95 %; a per-field floor of 85 % applies only to fields with ≥ 20 scored cells, smaller fields are reported". "Smaller fields are reported" exempts them from the *floor*; it does not remove their cells from the kind. The implementation always excludes `movement_count`, `opening_balance` and `closing_balance` (5 cells each) — precisely the statement cells Q-EX-7 names — so the movements gate as ruled was never computed. The reports do not disclose the divergence. | Either compute the kind rate over all of the kind's cells (the ruled reading — all three sittings still pass: 99.55 / 98.56 / 96.08 %), or amend the spec and say in every report header that the kind rate covers only fields with ≥ 20 cells, with the excluded count. Reconcile the two tables on the page. |
| **E-3** | **High** | "the reserve … PASSED"; "movements 156/156 = 100.00 %; header and lines reported, not gated" (appeals sheet); "not one of the 27 missed cells is an extraction error" (`e8e1227`). | Recomputed from `score-20260907-1613-dff90f9-a558523.json` (the surviving blind run): 289 cells, **28 missed** — Header 76/80 = **95.00 %**, Lines 27/51 = **52.94 %**, Movements 158/158. Under the ruled kind definition (E-2) the blind reserve's **lines rate is 52.94 %**, a hard fail of the 95 % lines gate; under the implementation the Lines kind is simply "not gated". That report nonetheless headlines "**Gate (S-10): PASSED**". Neither rate appears in the appeals sheet the owner ruled from, in spec §1.1 D-EX-6, or in the checkpoint claim table. | The record presents the held-out set as a clean pass and never states the blind numbers. The only reader who learns that 24 of 51 blind line cells missed is one who recomputes the JSON. Separately, S-10 says the reserve is "reported beside", not gated, so "Gate PASSED" on a reserve report overstates what the reserve is for. | State the blind reserve rates (Header 76/80 = 95.00 %, Lines 27/51 = 52.94 %, Movements 158/158 = 100 %) beside the post-ruling ones wherever the reserve result is claimed, with the explanation. Make a reserve report print "reported, not gated" instead of "Gate PASSED". |
| **E-4** | **High** | Every report's "Files read" column: `0` for all 55 document rows across the four sittings; §5.7 requires "per-document `ProcessingEvidence` with `modelUsage`, `num_turns`, **files read**, duration", and §5.7 makes a usage-limit signal a stop. | `tools/extraction-bench/Sibyla.Tools.ExtractionBench/BenchRun.cs:213` — `var stdout = result.RawStdoutPath is { } raw && File.Exists(raw) ? File.ReadAllText(raw) : "";` — but `src/Sibyla.Worker.Documents/ExtractionRunner.cs:20-29` documents `RawStdoutPath` as **relative to the evidence root** (`"e2e8631894014069b90c75ebf185c1d9/attempt-1-stdout.txt"` in the reports), so `File.Exists` is false and `stdout` is always `""`. Measured across all four report JSONs: `filesRead` total = **0** in every document of every sitting; document-level `numTurns`/`modelUsage` null. The raw trace contradicts it: `grep -c '"name":"Read"'` on the reserve re-run's `…/e2e86318…/attempt-1-stdout.txt` = **7**, and the same file carries `"num_turns":8`. | A printed evidence column states a fact the raw trace disproves. Worse for this angle: "which files did the model read" is exactly the evidence the §4.5 permission design is meant to leave behind, and no sitting has it. `BenchTrace`'s usage-limit stop (§5.7) is dead code in practice for the same reason. | Join `RawStdoutPath` to the run directory before `File.Exists`, re-derive the trace summaries from the retained raw stdout of the three surviving run folders, and re-issue the reports. Add a test that a bench record's `FilesRead` is non-empty for a run that read a file. |
| **E-5** | **High** | "The committed pre-ruling report survives in full, **so nothing about the first measurement is lost**", and that "a careless glob deleted that folder's `actual/` directory" (`docs/apollo-argus-extraction-v1-key-appeals-260907.md`, "After a ruling" section). | The whole run folder `score-20260907-1613-dff90f9-a558523/` is gone, not just `actual/` — `ls -d …/score-*/` lists only `…-1321-…`, `…-1546-…`, `…-2250-…`; `find D:/fileStorage/tmp/apollo-extraction -name "*1613*"` returns nothing; `.gitignore` (`score-*/`) kept run folders out of git, so nothing is recoverable. Measured from the surviving report JSON: **0 evidence elements** for all ten documents — no attempt records, no `promptSha256`, no `cliVersion`, no `modelUsage`, no `rawStdout` hashes, no durations. | Materially more was destroyed than the account says: the blind run's raw CLI traces, per-attempt evidence and answers are all gone. "Nothing about the first measurement is lost" is not true — what survives is the scored cell matrix (289 cells with expected/actual), which is real but is not the run. This is the sentence that makes a reader comfortable about a destroyed control. | Correct the account: name what is gone (the whole run folder: answers, traces, per-attempt evidence) and what survives (the cell matrix in the committed `.json`). Take run folders out of the blast radius of cleanup — never glob-delete under `evidence/extract-v2/`. |
| **M-1** | Medium | The four `score-*.md` / `.json` are "the evidence that belongs in the history" (`5ef58e6`) and §5.7's per-document evidence record. | `score-20260907-1321-c54e198-a558523.json` at HEAD: 35 documents, all `status: resumed`, `attempts: 0`, `durationMs: 0`, **0 evidence elements**; header reads "Set Fiscal; 2026-09-07 **22:50 → 22:50** UTC" for a sitting whose documents were read 13:21–15:0x. `git show a50e5d5:…json` has 60 evidence elements and the real attempt table (five `exhausted`, the V-4 reasons, 43–395 s per document); `git show ce7d4cf:…json` has 6. Same pattern on the pre-ruling reserve report (0 elements). The re-scores overwrote the reports in place. | The artifact a checkpoint reader opens for the largest sitting carries no evidence of the run at all, and nothing on the page says it is a re-score of an earlier run. Only `git log` holds the record. | Have `-Resume` re-scoring preserve the resumed documents' evidence elements and original timings, or print a "re-scored from run of <stamp>; original report at <sha>" banner. Keep the superseded report beside the new one, as is already done for the reserve. |
| **M-2** | Medium | §5.5: "**More than 5 % of scored cells corrected stops the run for the owner's look**." | `grep -rn "0.05\|5 %\|FivePercent" tools/extraction-bench/**/*.cs` — no match. The scorer counts corrections (`Scorer.cs`, `CorrectionNotePrefix`) and reports them, and applies only the 95 % kind rates and the 85 % floor. | A spec-mandated safeguard against key-editing is simply not implemented. It never fired (8/756 = 1.1 % fiscal, 16/1389 = 1.2 % statements, 1/268 reserve), so nothing is wrong with these numbers — but the guard that exists to catch exactly the abuse this review is looking for is absent. | Implement it as a gate reason, RED-first, and re-run the scorer over the three retained runs to record that it does not fire. |
| **M-3** | Medium | `41c5081`: "Two defects found while applying it, **both RED first**"; `a50e5d5` amends three caps and V-17 with its tests in the same commit. | `RED-red.txt` sections end at "D-EX-5 register branch — RED"; `RED-green.txt` sections end at "R-EX-2 — the permission probes". Neither file has a section for the R-EX-3 amendment or for D-EX-6. `git show a50e5d5 --stat` and `git show 41c5081 --stat` each carry the tests and the production change in one commit; no failing-run output is recorded anywhere for either. | The RED-first discipline is evidenced meticulously for R0–G3 and D-EX-5 and then stops for the two post-GREEN changes — including the one that introduced a new scoring rule after the answers were read, which is where the evidence matters most. The claim "both RED first" is unsupported by any artifact. | Append an "R-EX-3 amendment" and a "D-EX-6" section to `RED-green.txt` with the failing run for each new test (reproducible by reverting the production change locally), or withdraw the "RED first" claim from the commit record. |
| **M-4** | Medium | The clean-export fingerprints as the binding of the reviewed tree. | I reproduced both exactly, on a clean `git archive` export, with the committed script — RED `62cd08b`: `def4cd1f273226bc050c01001e95dc3a1526a0dbfcf50eedaa89a44d2ba9bc49 over 529 files`; G3 `aba0ad5`: `60660617550fc4fa97c755c399b221d99f3aeb7acd5ae09a20d78a6da106c2fb over 561 files`. Both match `RED-red.txt` and `RED-green.txt` to the character. But the script's roots (`New-ExtractionFingerprint.ps1`) are `tests/Sibyla.Tests.{Platform,Argus,Browser}`, `src`, `tools`, `Sibyla.slnx` — **`docs/` is not among them** (0 `docs/` entries in either file list), and no fingerprint was taken at the reviewed head. | The fingerprints bind the oracles, the keys (41 answer-key files at RED), the flags, the corrections and the skill package (13 files at G3) — that part is excellent and verified. They do **not** bind the accepted specification, which was amended twice after RED (`a50e5d5`, `41c5081`), and the last fingerprint predates the register branch, the cap amendment and D-EX-6. The reviewed tree has no fingerprint. | Add `docs/apollo-argus-extraction-v1-spec.md` to the roots and record a third fingerprint at the checkpoint head, in `RED-green.txt`, before the release. |
| **M-5** | Medium | `permissions.json`'s enumerated denies protect `local\secrets` among the §8 restricted roots; R-EX-2 settled that "every unprefixed form binds". | `PermissionsFile.Build` (`src/Sibyla.Worker.Documents/ClaudeCli.cs:234`) emits `Read(local/**)`. The deny matrix (`deny-matrix-20260907-114606-production/MATRIX.md`, `…-114717-production-dir/MATRIX.md`) proved ten forms: absolute `dir/**` in both slash styles, `**/dir/**`, `dir/*.json`, absolute file paths, `**/name`, bare `name`. A **bare relative `dir/**`** was not among them. | One rule in the deny list has an unproven form, and it is the one covering `local\secrets`. Relative to the CLI's cwd (the sandbox, per the probe traces' `"cwd"`) it would not name the repository's `local\` in any case. The restriction is in fact held by `--restricted` and the two `--add-dir` roots, not by this rule — but the record presents the rule as a control. | Write the rule as an absolute path (or drop it and say the confinement is what holds `local\`), and add the bare relative `dir/**` form to the matrix the next time the probes run. |
| **M-6** | Medium | §8's R-EX-3 gate list. | §8 requires, for R-EX-3: "hostile-document run … recorded"; "**a real intake row's evidence shows the tree hash**". The first does not exist (known, below). The second cannot exist before the release, which `docs/apollo-argus-extraction-v1-owner-guide-260907.md` places at step 7, **after** the checkpoint review at step 6. | Two of the checkpoint's own gate rows are unmet and the record does not name them as deferred anywhere a reviewer or the owner would see; the guide silently reorders the gate. | List the deferred R-EX-3 rows explicitly (with the reason and the point at which each closes) in the owner guide and in the release note, so that "R-EX-3 passed" is never read as "every row of §8 passed". |
| **L-1** | Low | "not one of the **27** missed cells is an extraction error" (`e8e1227`); "Cost 22 of the reserve's **27** missed cells" (appeals A-4). | The pre-ruling reserve JSON has **28** missed cells: Header 4 (I26010042 `date_due`, I26070026 `document_id`, I26070007 `line_count`, I26080031 `line_count`) + Lines 24 (I26070007 21, I26080031 3). A-4's own 22 is right. | Off by one in two places in the primary narrative. | Correct to 28 in the appeals sheet; the commit message stands as history. |
| **L-2** | Low | `RED-green.txt`, "Owner items before the release": "appsettings `Worker:{SkillPackageSha256 = b85d09956ec59a8f540d760426a57a57662f6ebcb3fa0af5ed2f4295c1c0ed51`, …}". | The package was re-pinned during the shakedown to `083cbd796ca37b4ea73d70b36a874e33cee50762cd754553d8d6bb3a8e4c53c5` (`2acf0b4`; the golden manifest's `pinnedAt` records both hashes). The owner guide was corrected (`c4fd799`); the evidence file was not. | A worker configured from that paragraph fails its own start-up check. Harmless because the guide is the operative document, but the evidence file is the one this checkpoint reads. | Append a dated line to that paragraph naming the superseding hash. |

### Known and out of scope

`evidence/extract-v2/hostile-documents-run.md` does not exist; §4.5's hostile-document run against the real CLI
was never done. Recorded as known and being closed in parallel (`69a30e3`, `f8dc163`, both after the reviewed
head). Not counted above. It is, on its own, a blocking R-EX-3 row.

---

## Was the measure tuned to pass?

**My answer: no — with one qualification that is not about tuning but about what the reserve result is worth.**

I attacked this in the order given.

**1. The three length caps and V-17.** The claim is that our own bounds were refusing right answers. I could
test that exactly, because the superseded report survives in git.

* The five refusals are named with their measured lengths in `git show a50e5d5:…score-…-1321-….md`:
  `I25120001`, `I26080006`, `I26080007` (`vat_exemption_text` 100 chars), `I26040053` (133),
  `I26030031` (`payment_terms_text` 147), plus second attempts on `I26050067` / `I26070005` / `I26080017`
  (`evidence.notes` 302 / 335 / 329). Every one of those three fields is free text and **none is a scored
  cell** — none appears in any per-field table in any report. Raising a cap on an unscored field cannot turn
  a wrong scored answer into a right one.
* I diffed the cell matrices before and after the amendment. `a50e5d5 → ce7d4cf`: **82 cells added, of which
  zero missed; zero flips among the 712 cells already scored.** The five re-run documents contributed only
  matching cells. The commit's claims "none of them added a miss" and "a relaxed bound cannot turn an
  accepted answer into a refused one, so the sitting's 30 scored documents stand" are both literally true.
* The rate did rise, 98.90 → 99.07 % header, because 64 perfect header cells joined a 365-cell denominator.
  That is the amendment's only effect on the number, and both figures clear 95 % comfortably.
* V-17 is a genuine loosening — from "null under **every** index base that resolves" to "null under **at
  least one**" (`ExtractionContractV2.cs`, the `not_printed` block). It lets a false `not_printed` claim
  through under the base we did not mean. But `not_printed` is not scored either, and the direction of
  effect is to accept answers that would otherwise be refused — and a refused document contributes **no**
  cells at all, so relaxing V-17 can only add cells that may miss. It is conservative with respect to the
  gate, not favourable. (It also has independent support: both statement answers wrote their `uncertain`
  paths 0-based, which is the ambiguity the amendment names.)
* The rejection corpus was **not** touched: `git log --name-status main..HEAD -- tests/Sibyla.Tests.Platform/golden/extract-v2/`
  shows the 50 reject files added once at RED (`fac7902`) and never modified. `V-17-not-printed-non-null.json`
  and `V-17-not-printed-unresolvable.json` still stand and still pass (Platform extraction 174/174).
* The one change the shakedown made to the *scorer* made it **stricter**: `2acf0b4` added the "a gate over
  nothing is not a pass" reason. `Similarity.cs` — the S-11 0.9 threshold and the prefix rule — was never
  touched after G1.

**2. S-12, the rule introduced after the answers were read.** This is where a benchmark gets gamed, and I
expected to find the damage here. I did not, and the reason is unusually strong.

* The flag is set on exactly two documents — `grep -rn keyLinesSummarised` returns only `I26070007` and
  `I26080031` in `answer-key.flags.json`, each carrying `"ruling": "D-EX-6 / A-4"` / `"A-5"`. `I26070005`,
  the one real extraction error of the slice (a merge of two printed lines), does **not** carry it, is still
  scored line by line, and still misses in the report at HEAD.
* The rule ships with a negative test: `ExtractionScoringTests.AnAnswerThatMergesLinesStillFailsUnderTheSummarisedRule`
  strips six of the seven answer lines and asserts all three sum cells fail.
* The arithmetic checks out: the answer's seven lines sum to **exactly** 260.41 / 50.82 / 311.23 — the key's
  single line and the printed header, to the cent — with the VAT split 41.04 + 0.25 + 7.69 + 0.06 + 0 + 0 +
  1.78 = 50.82 and two M07-exempt components at zero.
* **The decisive point.** `I26080018` — the Locarent credit note that annuls `I26070007` — is one of the six
  goldens **hand-authored at RED** (`fac7902`, before a single bench run, before any answer existed). It
  carries **seven lines** with exactly those amounts: 178.45/41.04/219.49, 33.43/7.69/41.12, 38.43/0/38.43,
  1.07/0.25/1.32, 7.75/1.78/9.53, 1.04/0/1.04, 0.24/0.06/0.30. And the FDR's own measured key for that
  credit note, `golden/extract-v2/answer-keys/I26080018.json`, copied unedited at RED, holds `lineCount: 7`
  with the same seven components as `Anulacao …` rows. The register stores seven lines for the credit note
  and one rolled-up line for the invoice it reverses. The answer matched the register's own seven-line
  reading of the same lease contract. That is not a rule bent to fit an answer; it is a defect in one key,
  provable from an oracle that predates the answer. (`RED-red.txt` deviation 7 records the same seven VAT
  figures at 02:16 on 2026-09-07.)
* `I26080031` (A-5) is the same shape and smaller: one printed zero-value row `Mês Agosto` that the key
  dropped; both sides read 100.00 net.
* A-1 was **refused** — the key stood and the miss was kept. Sitting 2's 55 missed cells were **not**
  appealed at all, and I confirmed the answers had self-declared exactly those rows:
  `BPI-DO-CRF-USD_202604` `evidence.uncertain = ["movements[1].posting_date","movements[2].posting_date"]`,
  `BPI-DO-GOT-EUR_202507` `= ["movements[30].posting_date","movements[31].posting_date","movements[32].posting_date"]`,
  each with a note explaining the page-break reasoning. The model was wrong, and it was scored wrong.
* The corrections mechanism is disciplined: the 40 measured keys and the 10 reserve keys were **never
  modified** (`git log --name-status` on `golden/extract-v2/` shows adds only); every change is a row in
  `answer-key.corrections.csv` or a flag, each carrying `decided_by`; and the RED data guards were made
  **stricter** by the ruling, not looser — `DataGuard_TheCorrectionsFileHolds23PreDeclaredRowsAndTheThreeDEx6Rulings`
  now asserts which three rows were added and by whom, and the flags guard asserts the three `ruling` fields.
* And the re-score changed nothing beyond the ruling: `ce7d4cf → HEAD` is exactly **two** cell flips,
  `I26030019 document_id` and `I26050003 fiscal_no` — the two D-EX-6 corrections that touch the fiscal set —
  and nothing else in 794 cells.

**3. The reserve seal.** In substance, partly. In form, no.

* What held: the reserve keys were exported by the owner at `a4311ad` (11:13) before the reserve ran
  (16:13); the configuration tuple was frozen and identical across all three sittings; nothing about the
  extractor changed between the two reserve runs — the commits in between (`e2f02e3`, `41c5081`) touch the
  appeals sheet, the flags, the corrections and the scorer, never the skill package, model, effort or CLI.
  So run 2 is a re-sample of the same extractor, not a re-tuned one, and it produced no new miss.
* What did not hold: the set was read twice; three flag rows and one correction row were written **about the
  held-out set on the strength of having read its answers**; and the blind run's evidence was destroyed. The
  flags file's own `rulings` block states the tension honestly ("§5.6 seals the held-out set: nothing may be
  written about it on the strength of having read its answers, so this is dated and attributed") — the right
  way to record a breach, but still a breach.
* This matters more than it would elsewhere, because §5.7 makes sittings 1 and 2 explicitly *tuning*
  sittings ("two tuning sittings … while the package, `EXTRACT.md`, model or effort are tuned"), and the
  skill package was in fact re-pinned mid-shakedown on documents from the measured fiscal set. So the
  reserve is the **only** blind measurement in the slice, and it is the one that was compromised.

**Conclusion.** I believe the gate was met honestly. Every scored number in the four reports reconciles to
the committed cell matrices; the two arithmetic rules I disagree with (E-2) move the published numbers
*down*, not up; and the post-hoc scoring rule survives the hardest test I could devise. What I do not
believe is that the slice currently holds a **held-out** result. It holds three measured sittings, one of
which was re-read after its keys were amended, and whose blind figures (Header 95.00 %, Lines 52.94 %,
Movements 100 %) appear nowhere in the record. That is what E-1 and E-3 ask to be fixed — by wording and by
a fresh control, not by re-running anything that has already been read.

---

## Verified as sound

* **Both clean-export fingerprints reproduce exactly.** I ran `git archive` of `62cd08b` and `aba0ad5` into
  fresh temp trees and executed the committed `New-ExtractionFingerprint.ps1` against each:
  `def4cd1f…bc49 over 529 files` and `60660617…c2fb over 561 files` — character-for-character the values in
  `RED-red.txt` and `RED-green.txt`. The framing is honest (per file `<rel>\0<len>\0<bytes>`, ordinal sort,
  `bin`/`obj`/`evidence` excluded) and it binds what matters: the oracles, the 40 answer keys, the flags, the
  corrections and (at G3) the 13 skill-package files.
* **RED counts reconcile exactly.** 242 `Failed` + 34 `Passed` = 276 rows in the `RED-red.txt` per-test
  table, matching the stated 242 = Platform 186 + Browser 5 + Argus 51. The 34 passes decompose exactly as
  claimed: 25 pre-existing (AccountPeriodTests 11, CompanyMatcherTests 10, ProcessingEvidenceTests 4 —
  confirmed against `git show main:…ProcessingEvidenceTests.cs`, which has exactly those four) and 9
  GREEN-by-design guards. I read the bodies of all nine at `62cd08b`; every one reads committed files or a
  scaffold-declared record shape, each is labelled as such **in the test's own comment written at RED**, and
  none is a behaviour test. The HARD STOP was genuinely not reached.
* **GREEN turned green only where it should.** `RED-green.txt` per-test diffs: G1 170 green all inside the
  five G1 classes, 72 still red; G2 +59, 229 cumulative, 13 still red all G3; G3 +13, **242 of 242**, with
  0 regressions at every step and "tests not present at RED" being only the `PreviewParityTests` set. The
  D-EX-5 register branch has its own RED section (17 failing + 1 green-by-construction) before its GREEN
  commits.
* **The suites are green now.** `dotnet test tests/Sibyla.Tests.Argus` → **204/204**;
  `dotnet test tests/Sibyla.Tests.Platform --filter "…Extraction|ClaudeCli|SkillPackage|CliProcessRunner"` →
  **174/174**. Both match `41c5081`'s claim ("Argus 204, Platform extraction 174") exactly.
* **Every rate in every report recomputes** from that report's own `cells` array, under the implementation's
  rule: 427/429 = 99.53 %, 322/327 = 98.47 %, 1334/1389 = 96.04 %, 156/156 = 100 %. Corrections applied
  (8 / 16 / 0 / 1) reconcile with the correction rows that touch each set, including the +2 the `document_id`
  fix unlocked. The 85 % floor is never approached by a gated field (lowest 95.05 %, four statement fields).
  No field was excluded other than by the documented ≥ 20-cell rule (E-2), and every excluded cell in all
  four sittings **matched** — so the exclusion made the published numbers worse, not better.
* **Q-EX-14 tuple integrity holds.** All four reports carry an identical tuple — CLI `2.1.259`, model
  `claude-opus-5`, effort `high`, `SkillTreeSha256 083cbd79…c53c5`, contract `sibyla.extract.v2` — and an
  identical bench identity (`D:\fileStorage\tmp\apollo-extraction\bench-config`, explicitly not part of the
  tuple). The re-pin from `b85d0995…` is commit `2acf0b4`, committed **13:11**; the first measured sitting
  started **13:21** (`startedAt 2026-09-07T13:21…`). The shakedown reports show the transition mid-shakedown
  (four at `b85d0995…`, two at `083cbd79…`). The re-pin is dated and attributed in
  `skill-package.manifest.json` (`pinnedAt`) with the superseded hash recorded, and the `EXTRACT.md`
  paragraph it added is a contract-shape instruction ("these members always carry a value") that teaches
  nothing document-specific.
* **The R-EX-2 defect and its fix are real, and the evidence is first-class.** The deny matrix is raw
  `stream-json`: the canary reads, `deny-a.txt` (the `//D:/…` form) **reads and its planted token appears in
  the trace**, and forms B–E refuse with `<tool_use_error>File is in a directory that is denied by your
  permission settings.</tool_use_error>`. `PermissionsFile.Rule` (`ClaudeCli.cs:256`) now writes the
  unprefixed form; `ClaudeCliPermissionTests` asserts `Read(D:/ApolloData/staging/**)` and
  `DoesNotContain(rule => rule.StartsWith("Read(//"))`; sitting 3's trace shows the planted `deny-probe.txt`
  refused while `document.pdf` and the package files read. `--settings` under `--safe-mode` is proved rather
  than inferred, and the init events show `"tools":["Read"]` with empty `mcp_servers`.
* **The standing restrictions hold, as far as code and evidence can show.** The bench refuses a
  `--config-dir` under `D:\ApolloData\worker-claude` (`BenchRun.cs:49-52`) and refuses any path under the
  forbidden roots or a `secrets` segment (`BenchRun.cs:302-317`); all four runs record the bench config dir,
  never the worker's. `invoice-skill-build` is at `a558523f3f0ad97e6c3f60707d635b3d257d3392`, `git status
  --porcelain` empty, last commit 2026-09-03 — before this slice began; the bench asserts the clean tree and
  the pinned HEAD at start and the post-condition at end, and all four reports record "clean at end — yes".
  My grep for secret markers over the whole committed evidence tree returned only test names
  (`sk-ant-api03-abcdefghijklmnop`, `Host=localhost;…Database=gott_sibyla`) inside the `RED-*.txt` per-test
  tables — synthetic by construction. `answer-key.meta.json` attributes the key read to 2026-09-06 and the
  `syncRunId` to the owner's own SELECT, noting "No key value depends on it". The branch is unmerged, so no
  v2 code is deployed — the "no production deployment before R-EX-3" restriction holds.
* **The Q-EX-22 purge record** is an append-only log that opens its own loose end (`sessions\`, `backups\`),
  chases it, and closes it with an explanation that fits the facts (five identical 41,575-byte
  `.claude.json.backup.<ms>` copies, one afternoon, an order of magnitude too small for a transcript). It
  states plainly that every figure came from the owner's survey and that nothing in the work read the path,
  and it leaves the repeat purge after the release as an open obligation.

## Could not verify, and why

* **That the owner ruled D-EX-6.** The record attributes it ("the owner: 'accept all seven
  recommendations'"; `decided_by` on the three correction rows; the `rulings` block in the flags file), and
  the timing is consistent (appeals sheet `e2f02e3` 22:23, applied `41c5081` 22:52). I have no independent
  channel to the owner. The same applies to D-EX-5.
* **The contents of `D:\ApolloData\worker-claude`.** Forbidden to me by the review's own rules, so the
  housekeeping record's figures — before/after counts, the folder listing, the five backup files and their
  exact size — rest entirely on the owner's restatement. No raw command output is pasted, so nothing in that
  document is independently checkable. That is inherent to the restriction, not a defect of the write-up.
* **The pages of the reserve and sitting documents.** I did not open the source PDFs (`poppler` is absent on
  this host, per the R-EX-2 record). Every appeal that turns on "the document prints X" is therefore
  verified indirectly — for A-4 decisively, through the RED-authored golden and key of the annulling credit
  note; for A-2, A-3, A-5, A-6 and A-7 through internal consistency, the answers' own notes and the flags
  file's page references, but not by eye.
* **The blind reserve run.** Its answers, traces and per-attempt evidence are destroyed (E-5). I could audit
  its 289-cell matrix and did; I cannot compare its answers to the re-run's, so I cannot say whether the
  second read was better, worse or identical — only that it produced no new miss.
* **Sitting 1's and the pre-ruling reserve's run conditions.** The committed reports carry no evidence
  elements (M-1); I reconstructed the attempt table, the five refusals and the durations from
  `git show a50e5d5:…`, which is history rather than the artifact the checkpoint names.
* **The hostile-document run** (§4.5 / §8): does not exist at the reviewed head. Known, being closed in
  parallel, not assessed.
