# Argus extraction v1 — checkpoint round 3, CLOSURE (2026-09-08)

Reviewer: independent reviewer, round 3, **closure angle** — does every round-two finding close, and what did
the round-two fixes break. Branch `ops/argus-extraction-v1-red`, head **`81ab4fc`**, reviewed in its
**committed** state (`git archive`, never the working tree; `src/Sibyla.Web/**` and
`tests/Sibyla.Tests.Browser/**` are another session's and out of scope, as are the three files `21097cc`
carries whole). Round-two reports: `…-checkpoint-round2-closure-260908.md` (N-n) and
`…-checkpoint-round2-regression-260908.md` (R2-n). Closure commits: `c78e790`, `2f07eb9`, `24e4a52`,
`397afb9`, `3b15b5c`, `81ab4fc`. I did not read the other round-3 reviewer's work.

## Verdict

**REVISE — Critical 0 / High 2 / Medium 8 / Low 5.**

**Closure tally: 27 CLOSED, 8 PARTLY, 4 NOT CLOSED (39 round-two findings).**

The code half of round two is the strongest work in this slice. Every round-two reproduction I ran now behaves
the other way: `1348/1419` prints FAILED, `GR…` matches a register `EL…` through the gate *and* the register
lookup, a signed "Verified genuine repeat" survives a later contrary ruling, a ruling on `printedOpeningBalance`
now moves 48 cells that all carry it, and the R2-8 bound admits the sixteen real column slices. I recomputed all
three published rates from the answers and keys on disk with my own scorer, written from §5.3, and reproduced
**movements and lines cell for cell and 411 of the 447 header cells**, including both published misses. The
fingerprint reproduces character-for-character. The blind reserve banner's rates recompute exactly from the
surviving matrix.

Two things stop this from being an ACCEPT. **The R2-6 fix was written into the specification and never into the
code**: S-6 now says `service_period` "is reported, not scored … it moves to S-8", the commit message says the
same, `RED-green.txt` says the opposite ("no code changed"), and the scorer still scores those twelve cells —
so the published header denominator of 447 is not the denominator the specification at this head defines. And
the R2-5 rewrite, which correctly stopped destroying declarations, wrote a new false statement into the same
register of record: every second ruling on an entry claims to supersede a declaration about a **different
file**, in the permanent evidence text and in the audit detail, with no test over it. The pattern the brief
warned about held for a third round.

---

## Closure table — every round-two finding

Evidence: `H` = harness against the built `Sibyla.Platform.Infrastructure.dll` (console project under
`D:\fileStorage\tmp\rev3\probe`, nothing written in the repository); `T` = suite run under the shared build
lock; `R` = my own re-implementation of §5.3 over the committed answers; `C` = code/diff read; `G` =
`git archive` / `git show` against history.

| id | sev | state | evidence | what remains |
|---|---|---|---|---|
| **N-1** | **High** | **PARTLY** | C: `spec.md` §5.6 now reads "the reserve **was to be** the third, single sitting. **It was not: see D-EX-7**", carries the blind rates inline, and keeps the intended sentence visible rather than deleting it. The change-log citation at `:11` ("§1.1 D-EX-7, §5.6, §5.7, S-10") is now true. The self-contradiction the Critical named is gone from §5.6 | Q-EX-9 at `spec.md:72` still says the reserve was "fixed now, **opened once** — as a third, single bench sitting", unannotated → **R3-9** |
| **N-2** | Med | **PARTLY** | C: five of the six amendments (N-3, V-5, V-13, V-15, V-17) now sit inside their row's last cell; I re-ran a GFM cell-count over the whole file | **S-11's row (`:450`) is still truncated**, by unescaped pipes inside a code span rather than by a trailing cell → **R3-3** |
| **N-3** | Med | **CLOSED** | C: `RED-green.txt:1934-1991` is a RED record for `21097cc` covering C-1, C-3, C-4, C-10–C-12, C-14–C-16, C-21–C-25, C-27, C-28 and M-5, with the two reverting cases and the four first-run-green tests declared. G: spot-checks hold — at `21097cc^`, `NormalizeTaxId` has no call site in `src/` or `tools/` (C-24), and `HasMaxLength(20)` on the entry code is present (C-28) | — |
| **N-4** | Med | **CLOSED** | C: the owner guide gains "A sitting now prints FAILED, and that is correct", the `Score-Combined.ps1` command, the combined report's name and its rates, and the honest note that the margin is one cell on five movements | — |
| **N-5** | Med | **CLOSED** | C: fiscal and statements re-scored and now print **FAILED** with the right reason; reserve-2 PASSED; the blind run carries a superseding banner. R: I recomputed the banner's rates from the 289-cell matrix in `…1613….json` — **76/80 = 95.0000 %, 27/51 = 52.9412 %, 158/158** — exactly as printed | the recovered fiscal report's `status` column → **R3-6** |
| **N-6** | Med | **CLOSED** | C: `Scorer.WithTheRulingOnEveryCellItMoved` re-scores the document with `NoFlags` and marks every cell whose expected value, existence or verdict differs. T: `ARulingOnAStatementFlagIsInsideTheFivePercentGuard` (48 cells, none of which any old call site touched) and `ARulingOnATaxIdOrDueDateFlag…` (company, date_due, fiscal_no, origin_class). RED recorded with the exact pre-fix collections (`[]` and `["date_due"]`) | a ruling that **removes** a cell is still invisible; disclosed in the test's doc comment but not in the commit or the spec → **R3-15** |
| **N-7** | Low | **NOT CLOSED** | G: nothing in `RED-green.txt` or any later commit corrects `21097cc`'s "the split would not compile" claim. The N-8/R2-24 appendix corrects the test counts only | one line in the record → **R3-13** |
| **N-8** | Low | **CLOSED** | C: `RED-green.txt:1992-2011` names the filter and sets the rule. T: I ran the documented filter — Platform **250 passed / 0 failed**; Argus **231/231** — exactly what `397afb9` claims | — |
| **N-9** | Low | **NOT CLOSED** | C: `IngestionServiceHoldTests.cs:391` — the `/// <summary>The register entry I26010021 … (AccountPeriodTests' seeding shape).</summary>` is still immediately followed by `[Theory] DiscardPreservesEvidenceAndAuditsOnce` at `:392-397` → **R3-12** |
| **N-10** | Low | **PARTLY** | C: the comment now says "a label **separated from** the number" — accurate | it added a claim that is false: "`CompanyMatcher` strips the glued form when it compares, so the gate still matches the register" → **R3-5** |
| **R2-1** | **High** | **PARTLY** | Same as N-1; `Score-Combined.ps1:29` now reads "NOT a held-out result: it was read, re-keyed on what was read, and read again" | six of R2-1's named sites still assert the seal — `Rescore.cs:27`, `AnswerKeys.cs:32,51,85`, `ExtractionScoringTests.cs:249-251`, the `answer-key.flags.json` rulings block, `key-appeals:95`, `spec:72` → **R3-9** |
| **R2-2** | **High** | **CLOSED** | C: both reserve reports carry a status banner; the blind one is headed "**SUPERSEDED, AND ITS VERDICT IS WRONG**" before any figure, gives the corrected rates, says why it cannot be re-scored, and points at the combined report as the gate. R: its rates verified | — |
| **R2-3** | **High** | **CLOSED** | C: `ExtractionScorer.Meets(matched, scored, bar) => matched >= bar * scored`; `Rate` is display-only; both gate loops and both reason strings use `Meets`/`PercentExact`. Arithmetic checked: `0.95m × 1419 = 1348.05m` exactly, so 1348 fails and 1349 passes. T: `TheGateComparesTheUnroundedRateSoOneCellUnderNinetyFiveFails` runs the exact 1348/1419 reproduction; `TheFieldFloorAlsoComparesTheUnroundedRate` the 16999/20000 one; the RED transcript records both failing first | — |
| **R2-4** | **High** | **CLOSED** | H: I ran the reproduction against the built assembly. `Gate("GR123456789", …, "EL123456789") → MATCH`; `Gate("XI432525179", …, "GB432525179") → MATCH`; `SameTaxId` true for both; `Gate("PT500940231", …, "ES500940231") → no match`, so C-11 is preserved; `NormalizeTaxId("123 456 789", "GR") = "GR123456789"` and that value matches `EL123456789`. Fixed at the cause (`CountryOf`), not at the call site | — |
| **R2-5** | **High** | **CLOSED for the destruction** | C: `IngestionService.cs:668-683` — `ON CONFLICT … DO NOTHING` under a filename scoped to the intake, a five-attempt retry, the prior row serialised whole into the audit detail. T: `ASecondRulingIsRecordedBesideTheFirstAndDestroysNothing` asserts both rows survive with their own clause, evidence, user and `declared_on`; `ARulingLeavesADeclarationTheLegacySyncOwnsExactlyAsItWas` asserts a sync-owned row byte-identical. C: `SyncEngine.Existing<T>` / `IsHeldByNativeRow` already excluded native rows, so the sync leg holds by construction | **the fix introduced a false supersession claim → R3-2** |
| **R2-6** | Med | **NOT CLOSED** | C: `Scorer.cs:293` still emits `Cell("service_period", …)`, a scored header cell, and nothing appears in `ReportedOnly`. The published report still prints `service_period \| Header \| 12 \| 12` inside a denominator of 447 | the spec, the commit message and `RED-green.txt` give three different accounts → **R3-1** |
| **R2-7** | Med | **CLOSED** | C: the combined report prints "192 movement(s) paired; 5 key and 5 answer unpaired. Cells of paired movements: 1334/1334 = 100.00 % … the movement fields are not independent checks but the pairing reported once per field" | — |
| **R2-8** | Med | **CLOSED** | C: `MinimumPrefixTokens = 3` **and** `PrefixCharShare = 0.25` of the longer side's characters. R: I implemented the new rule in my own scorer and the statements set still scores **1349/1419**; the sixteen corrected slices are unaffected. T: `TheS11PrefixBoundSeparatesARealColumnSliceFromATruncation` asserts the band, not the constant | — |
| **R2-9** | Med | **CLOSED** | C: `RescoreOptions.CheckedOutName` refuses a path; `Rescore.Execute` refuses an existing `.json`/`.md` without `--force`. T: `AReScoreRefusesAnOutNameThatIsAPathOrThatWouldReplaceAReport` | — |
| **R2-10** | Med | **PARTLY** | C: the completeness check compares the answered set with the set's keys and refuses without `--partial`. T: `AReScoreRefusesASetWhoseKeysNoRunAnswered` | `ExpectedKeysOf`'s `_ => answered` fallback makes the guard inert for a run whose set is unknown → **R3-7** |
| **R2-11** | Med | **CLOSED** | C: `Rescore` builds `tupleReasons` and passes them to `Report(…, alreadyFailing)`, which seeds `reasons`. T: `ACombinedReportOverTwoConfigurationTuplesIsAGateFailure` covers mixed, same and unknown tuples | — |
| **R2-12** | Med | **CLOSED** | C: every answer goes through `ExtractionContractV2.Validate` before scoring and its SHA-256 is recorded. R: I recomputed all 40 hashes from the files on disk and **all 40 match** the committed report, so the published number is bound to the bytes I scored. T: `AReScoreRevalidatesEveryAnswerAndRecordsItsHash` | the `.md` carries no hash → folded into **R3-6** |
| **R2-13** | Med | **CLOSED** | R: I counted the file — 26 rows = 3 header pre-declared + 16 movement (15 Revolut + 1 BPI) + 4 line rows for #31 + 3 D-EX-6. §5.5 and §8 now say exactly that; the withdrawn I24120001 entry is struck through with its reason | — |
| **R2-14** | Med | **CLOSED** | C: §2.6 now reads "**counterparty** fiscal number — the issuer on a payable, the recipient on a receivable, as §4.7 says and `CompanyGate` does", with the correction dated | — |
| **R2-15** | Med | **CLOSED** | C: the six-substring test is replaced by `AReleaseWithoutTheTupleOrDisagreeingWithThePinnedPackageIsRefused`, which exercises `ReleaseFacts.Read` / `VerifyAgainstGolden` on three bad fixtures, checks the repository's own worker `appsettings.json` against the pinned golden, and pins the script by expressions (`-ne $golden.treeSha256`, `has no Worker:$key`) that exist only inside the enforcement block | — |
| **R2-16** | Med | **CLOSED** | C+T: `AKindWithNoCellBesideAPerfectlyScoredKindIsStillNotAPass` — 1400 movement cells at 100 % with no header or line cell ⇒ `GatePassed == false` with both named reasons, exactly the shape R2-16 asked for | — |
| **R2-17** | Med | **CLOSED** | Same as N-6 | see R3-15 |
| **R2-18** | Med | **CLOSED** | C: `ReleaseFacts` reads the release's own `appsettings.json` as the authority and `WorkerStartupChecks.CheckReleaseFacts` refuses to start on any difference, naming the key, both values and the three override paths. T: `AHostOverrideOfTheReleaseTupleStopsTheWorker` over Model, Effort and SkillCommit, plus a passing case. Guide §8.2 now names `worker.json`, `appsettings.Production.json` and the `Worker__*` variables and marks it OWNER ACTION | the operational consequence is the other reviewer's; the closure is real |
| **R2-19** | Med | **PARTLY** | H: `NormalizeTaxId("12345678 (VAT2)") = "12345678"` — the named case is fixed | the class is not, and the amended N-3 describes a rule the code does not implement → **R3-4** |
| **R2-20** | Low | **CLOSED** | C: nothing is updated, so no row's `created_at`/`declared_on` can disagree and no clause is erased. T: `Assert.Equal(DateOnly.FromDateTime(later.CreatedAt.UtcDateTime), later.DeclaredOn)` | — |
| **R2-21** | Low | **CLOSED (code)** | C: `AssemblyInfo.cs` no longer carries `InternalsVisibleTo("Sibyla.Worker.Documents")` and `SameTaxId` is public. T: 250/250 on the documented filter, so the worker builds without it | §4.3 still specifies the attribute → **R3-8** |
| **R2-22** | Low | **CLOSED** | C: `ReadAnswerMovements` keeps a movement with no `posting_date`, flags `PostingDatePrinted=false` and charges it as unpaired. T: `AnAnswerMovementWithNoPostingDateIsChargedNotDropped` | — |
| **R2-23** | Low | **CLOSED** | C: `movement_count` scores `statement.movements_printed_count`, with the array length on the note and a "disagrees with itself" wording when they differ. R: I scored it that way in my own implementation and the cell is still 5/5 — **the semantic change does not move the published movements figure** | — |
| **R2-24** | Low | **CLOSED** | Same as N-8; both new counts reproduce | — |
| **R2-25** | Low | **CLOSED** | C: `XXX`, `XTS`, `XBA`–`XBD` added. T: `TheActiveListAndTheWithdrawnCodesAFiledDocumentStillPrintsAreAccepted` covers all six plus the eight withdrawn codes that had no test at all; `ACodeIsoNeverAssignedIsStillRefused` keeps the list closed | — |
| **R2-26** | Low | **CLOSED** | C: the revision-9 log now says "`false` at RED … **it is `true` since R-EX-2 of 2026-09-07** … this line records what revision 9 changed, not the value today" | — |
| **R2-27** | Low | **NOT CLOSED** | C: §4.5 still reads "`**\local\**` is one of the **ten forms the matrix did prove**". G: both `MATRIX.md` files list ten forms, of which **eight** bound — A and G read "the rule did NOT bind" → **R3-11** |
| **R2-28** | Low | **PARTLY** | G: attempts and durations are genuinely restored — I diffed the re-scored `…1321….json` against `git show a50e5d5:` and the 35 documents, 60 attempts, 60 evidence elements and the 13:21:11 → 14:45:22 window are **identical**, sums included (5,050,898 ms both sides) | statuses, `jobId` and `modelUsage` are not → **R3-6** |
| **R2-29** | Low | **PARTLY** | C: the quantity `AdjustedCells` is now asserted by the two new N-6 tests (`Assert.Equal(48, …)`, `Assert.Equal(4, …)`) | the two tests R2-29 named are unchanged: `ARuledDueDateFlagIsAnAdjustedCellAndTheHandChecksAreNot` still never reads `report.AdjustedCells`, and `Assert.Contains("35", markdown)` is still a bare substring search |

---

## New findings

| id | sev | file:line | what is wrong | why it matters | what would fix it |
|---|---|---|---|---|---|
| **R3-1** | **High** | `docs/apollo-argus-extraction-v1-spec.md:445` (S-6) vs `tools/extraction-bench/Sibyla.Tools.ExtractionBench/Scorer.cs:285-294`; `RED-green.txt:2267-2278`; commit `397afb9` | S-6 carries "**Amended 2026-09-08 (R2-6): `service_period` is reported, not scored** … It moves to S-8, reported beside the score." The commit message repeats it: "It moves to S-8, reported not scored." The code does neither — `Cell("service_period", expectedPeriod, statedPeriod, …)` at `Scorer.cs:293` is a scored header cell and no `ReportedOnly` entry for it exists (`BenchReport.FillRates` emits five entries, none of them `service_period`). `RED-green.txt` says the opposite of both: "the scorer implements the rule as written, **so no code changed**". The published gate report still prints `service_period \| Header \| 12 \| 12` inside header 445/447 | The header rate that gates the release is computed under a rule the specification at this head repudiates. With the amendment applied the header is **433/435 = 99.5402 %** — still a pass, so nothing operational turns on it — but a published gate number whose denominator disagrees with its own specification is the class this checkpoint exists to catch, and three records at one head give three different accounts of the same fix | Either implement the amendment (drop the cell, add a `ReportedOnly` entry, re-issue the combined report at 433/435) or withdraw it and keep the disclosure paragraph the report already prints, which is what the code actually does. Correct the commit record either way |
| **R3-2** | **High** | `src/Sibyla.Platform.Infrastructure/Ingestion/IngestionService.cs:641-646` (the prior-row SELECT), `:651` (`prior = priorRows[^1]`), `:655-658` (the `evidence +=`), `:699-700` (the audit detail) | The SELECT that finds "the declaration on file" filters on `owner_id` and `entry_code` **only** — there is no filename predicate — and `prior` is simply the newest row on that entry. The evidence text then appends "; supersedes the declaration on file for this entry (…) — that row stands as it was declared" **unconditionally** whenever any prior row exists, and `supersedesPriorDeclaration: true` goes into the audit detail with that unrelated row serialised whole. But two intakes with **different filenames** on the same register entry are independent declarations: "file X is a verified genuine repeat of I26010021" is not superseded by "file Y is a duplicate of I26010021". Both tests collide on the same filename (`doc.pdf`), so neither exercises the non-colliding path | R2-5 was raised because `duplicate_classification` is documented in this codebase as authoritative human input. The fix stopped destroying it and started writing a false claim into it — permanently, in the field that names the declaring user, on every second and subsequent ruling on an entry. An auditor reading declaration B would conclude a valid ruling on file X had been overturned. Reachable on the commonest shape: the same invoice re-uploaded under two names | Scope the prior-row lookup to the same `source_filename` — the only case that is genuinely a supersession — or drop the word "supersedes" and the boolean and say "a declaration is also on file for this entry". Add a test with two intakes, two filenames, one entry |
| **R3-3** | Medium | `docs/apollo-argus-extraction-v1-spec.md:450` (S-11), `:442` (S-3), `:270`–`:271` (V-8, V-9), `:68` (Q-EX-5) | N-2's class of error is not gone, only its trailing-cell form. GFM's tables extension splits cells on an unescaped pipe "including inside other inline spans", so a `\|` inside backticks is a separator. S-11's row contains `` `\|A∩B\| ÷ \|A∪B\|` `` and therefore has **6 cells in a 2-column table**: everything from "the N-4 token overlap (" onward — *including the C-26 amendment and this round's own R2-8 amendment* — is dropped by every GFM renderer. S-3's row (`sign_by_kind(doc, line) × \|extraction\|`) has 4 cells and loses its entire rule text; V-8, V-9 and Q-EX-5 likewise. `git log -L 450,450` shows `:450` was last written by `397afb9` and before it by `c78e790`, the commit that claims "All six now sit inside their row's last cell" | A reader of the rendered spec — which is how it is read on GitLab and in a preview pane — still sees no bound on S-11's prefix rule and no statement of S-3 at all. I detected these mechanically over the whole file; with the pre-existing `:508` they are all of them | Escape the pipes as `\|` in the five rows. A `\|`-aware over-wide-row check belongs beside the spec: this is the second round the class has bitten |
| **R3-4** | Medium | `src/Sibyla.Platform.Infrastructure/Ingestion/ExtractionContractV2.cs:536-550`; `spec.md:253` (N-3) | N-3 removes spaces, dots, slashes and hyphens **before** segmenting, so "the longest digit-bearing segment" is computed over a string in which the commonest separator no longer exists. H, against the built assembly: `"NIF: 500940231 Cap. Social: 250000000"` → **`PT500940231CAPSOCIAL`**; `"VAT GB 123456789 (Company no. 12465777)"` → **`COMPANYNO12465777`** (the company registration number, which §2.2 says is not a tax id); `"CIF B12345678, Reg. Merc. 12345679"` → **`REGMERC12345679`**; `"NIF 123456789 / RC 987654321"` → `NIF123456789RC987654321`. The tie rule (`segment.Length >= id.Length`, "a tie keeps the later segment") reintroduces R2-19's own shape at equal length: `"1234 (V123)"` → **`V123`**. All pass V-4's s(32) and reach `ExtractionResult`, the company gate and the register-duplicate key | C-24 made N-3 a *control* applied by the reader — "the canonical form … carries the N-3 shape whatever the answer wrote" — so it must be robust to what a model writes, and a two-number line is ordinary on a Portuguese or Spanish invoice. The failures are safe (`SameTaxId` then finds no match and the row triages) but the stored canonical id is wrong, and the amendment states the rule as solved. R2-19 also asked that the segmenting be described in N-3; it is not | Treat whitespace as a separator instead of deleting it; prefer the segment after a known label, then the longest, then the **earliest** on a tie; describe the segmenting in N-3 |
| **R3-5** | Medium | `ExtractionContractV2.Schema.cs:265-269`, `ExtractionContractV2.cs:531-534`, and `ExtractionContractV2Tests.N3DropsALabelBeforeASeparatorAndNeverOneGluedToTheDigits` | Both comments now assert, citing N-10, that "`CompanyMatcher` strips the glued form when it compares, **so the gate still matches the register**". H: it does not. `CompanyMatcher.SameTaxId` strips it (via `Printed`); `CompanyMatcher.Gate` and `Match` normalise with `Normalize` only. `Gate("NIPC504615947", null, [Candidate(c, "504615947")])` → **no match**, and against `"PT504615947"` → **no match**, while `SameTaxId` is `true` for both. The test's own summary makes the gate claim and then asserts only `SameTaxId` | The company gate is what assigns a document to a company. A document printing a Portuguese NIPC glued to its label — N-10's own example, which N-3 now deliberately preserves — is triaged rather than assigned, and the register lookup and the company gate disagree on the same value. It fails safe, so it is not unsafe; the documentation is false and the oracle does not test what it says it tests | Either route `Gate`/`Match` through `Printed` as `SameTaxId` does, or correct the two comments and the test summary to say the register lookup rather than the gate |
| **R3-6** | Medium | `evidence/extract-v2/combined-20260908-fiscal-statements-a558523.md:153` and its Documents table; `score-20260907-1321-…json` `documentRecords[].status`, `.jobId`, `evidence[].modelUsage` | The combined report — the gating artefact — introduces its Documents table as "Attempts, durations and **statuses as each run recorded them**" and prints `resumed` on all 40 rows. G: the fiscal run recorded `answered` on 30 and **`exhausted` on 5** (I25120001, I26030031, I26040053, I26080006, I26080007); the re-score overwrote the status. The disclosure exists in the JSON's per-document `note` and in the 1321 sitting report's Note column, but the combined `.md` has **no Note column and zero occurrences of "recorded it as"**. Two further losses in the recovered record: `evidence[].modelUsage` is dropped from all 60 elements, and the document-level `jobId`s are regenerated GUIDs while the evidence elements keep the run's real ones. The per-answer SHA-256 that R2-12 records is likewise in the JSON only | The five exhausted documents are the operative fact behind the outstanding owner work, and the gate report states the opposite of the record in the very sentence that vouches for the table. R2-28 asked precisely that these rows stop reading as measurements | Carry `status` forward as the durations now are, or add the Note column to the combined `.md` and change the sentence; carry `modelUsage` and the original `jobId` in `RestoreFrom`; print the answer hashes on the page |
| **R3-7** | Medium | `tools/extraction-bench/Sibyla.Tools.ExtractionBench/Rescore.cs:340-347` | `ExpectedKeysOf` falls back to `_ => answered` when a run's set is not Fiscal, Statements or Reserve. `SourceOf`'s own doc comment says "a run that never reached its report contributes what it can" — i.e. `Set` is null exactly when a run has no `run.json` and no report — and in that case the expected set collapses to the answered set, `missing` is empty, and R2-10's completeness guard passes silently over a subset | R2-10 exists because a folder holding 39 of the 40 answers can pass three gates over a set it did not measure. The fix closes that for a run that names its set and leaves it open for the one shape the code documents as possible | Refuse an unknown set without `--partial`, or fall back to the whole loaded key set rather than to the answered set |
| **R3-8** | Medium | `docs/apollo-argus-extraction-v1-spec.md:339` (§4.3) vs `src/Sibyla.Platform.Infrastructure/Properties/AssemblyInfo.cs` | §4.3 still specifies, "amended 2026-09-08 (C-10)", that "`Sibyla.Platform.Infrastructure` gains `[InternalsVisibleTo("Sibyla.Worker.Documents")]`, because §4.7's register lookup is specified in terms of `CompanyMatcher.SameTaxId` and the worker could not reach it". R2-21's fix deleted that attribute and made `SameTaxId` public | The seam clause of the specification describes an attribute the code does not have and a reason that no longer holds, in a paragraph amended one commit series earlier. Same class as R2-14, which was Medium | Rewrite the clause to say `SameTaxId` is public, and record that the hatch was withdrawn |
| **R3-9** | Medium | `spec.md:72` (Q-EX-9); `Rescore.cs:27`; `AnswerKeys.cs:32,51,85`; `ExtractionScoringTests.cs:249-251`; `golden/extract-v2/answer-key.flags.json:1858`; `key-appeals-260907.md:95` | R2-1 asked that every site asserting the seal be struck or annotated. §5.6 and `Score-Combined.ps1` were. **Q-EX-9 still reads "The ten named in §5.6 …, fixed now, opened once — as a third, single bench sitting under the frozen tuple at R-EX-3"**, unannotated, in the table of ruled owner questions. `Rescore.cs`'s `--reserve` help still says "also load the **held-out** keys of section 5.6" — R2-1 named it by line. `ExtractionScoringTests.cs:250` still says "**The seal is what makes the reserve set evidence that the two measured sittings were not tuned into**". The flags file's own D-EX-6 rulings block still says "because section 5.6 **seals** the held-out set" | The spec still contradicts itself about the reserve — §5.6 and S-10 say it was read twice and is spent, Q-EX-9 says it was opened once — which is N-1's finding at a different line. Code written after the ruling still teaches the reader the retracted claim | Annotate Q-EX-9 the way §5.6 was; rename the `--reserve` help text; correct the two comments and the flags block |
| **R3-10** | Medium | `docs/apollo-argus-extraction-v1-spec.md:243` (§2.6, step C) vs `src/Sibyla.Platform.Infrastructure/Ingestion/StampDutySplit.cs:107` | §2.6 step C says, in bold, "**`item.Net += other`**", and both worked traces (#25, #31) quote only net values. The code writes `target with { Net = Round(target.Net + surcharge), Total = Round(target.Total + surcharge) }`. R: implementing step C exactly as the spec writes it gives **lines 340/347 = 97.98 %** with `total_amount` at 105/109 and misses on I25040004 and I26080026; adding the undocumented `Total` fold reproduces the published **342/347** exactly | The published lines rate is not reproducible from the specification as written — an independent implementation lands two cells short. The corrected key rows for #31 set both `netAmount` and `totalAmount`, so the code is right and the spec is incomplete | Add "and `item.Total += other`" to step C and to the two traces |
| **R3-11** | Low | `spec.md:349` (§4.5) | "`**\local\**` is one of the **ten forms the matrix did prove**". Both `deny-matrix-*/MATRIX.md` list ten forms; eight are "BOUND (the read was refused)" and forms **A and G** read "the rule did NOT bind". R2-27 named this and it is unchanged | A permission claim in the section about the sandbox | "one of the eight forms the matrix proved, of ten tried" |
| **R3-12** | Low | `tests/Sibyla.Tests.Platform/IngestionServiceHoldTests.cs:391-397` | N-9 unchanged: the `SeedRegisterEntryAsync` doc comment still documents `DiscardPreservesEvidenceAndAuditsOnce` | Committed noise from the entanglement | move the comment |
| **R3-13** | Low | commit `21097cc`, "ONE THING THIS COMMIT CARRIES THAT IS NOT MINE" | N-7 unchanged: nothing in the record corrects the claim that the split "would not compile". A pushed message cannot be rewritten, but `RED-green.txt` took exactly that route for N-8 one commit earlier | The record corrects one untrue commit-message claim and leaves the neighbouring one standing | one line in `RED-green.txt` |
| **R3-14** | Low | `RED-green.txt:2010` vs `:2310`; owner guide §8 lead | The file states the rule "A count quoted in this slice from here on names the filter that produced it", and 300 lines later prints "FINAL RUNS (whole tree built, both suites) … Passed: 334" with no filter named. T: 334 **is** reproducible — with the wider filter written at the transcript's own BASELINE line — so nothing is hidden, but the rule the file just set is broken on its next page. Separately, guide §8 now has four subsections and its lead still says "**Three** of them leave you something to do" | The value of the record is that its numbers can be re-run without hunting for the command | name the filter beside the count; say four |
| **R3-15** | Low | commit `397afb9` ("every cell whose **existence**, expected value or verdict differs carries the ruling"); `Scorer.cs:120-132` | The mechanism iterates the **ruled** cell list, so a cell a ruling *adds* carries the ruling and a cell a ruling *removes* — `printsIssuerTaxId` / `printsRecipientTaxId` set to false, which deletes `fiscal_no` and the two gate cells — is counted nowhere. R2-17 named that direction explicitly. It is disclosed accurately in `ARulingOnATaxIdOrDueDateFlag…`'s doc comment and nowhere else | The commit message and §5.5's amendment claim a completeness the mechanism does not have. Nothing today needs it | count removed cells into `AdjustedCells`, or say in the spec which direction the guard sees |

---

## Independent recomputation of the three published rates

Method: I wrote a scorer from §5.3 in Python (`D:\fileStorage\tmp\rev3\{mov,hdr2,hdr3,gate,dt}.py`) over the 40
answers under `evidence/extract-v2/score-20260907-{1321-c54e198,1546-ce7d4cf}-a558523/actual/` and the keys,
flags and corrections under `golden/extract-v2/`. Before scoring I recomputed the SHA-256 of every answer file
and checked it against the `sha256` the committed report records for that document: **40 of 40 match**, so the
report and my run scored the same bytes.

| Kind | Published | Mine | Agreement |
|---|---|---|---|
| Header | 445 / 447 = 99.55 % | **445 / 447 = 99.5526 %** | 411 of the 447 cells recomputed directly; see below |
| Lines | 342 / 347 = 98.56 % | **342 / 347 = 98.5591 %** | exact, cell for cell |
| Movements | 1349 / 1419 = 95.07 % | **1349 / 1419 = 95.0669 %** | exact, cell for cell, same 70 misses |

**Movements.** Reproduced with no adjustment: pairing on `(posting_date ↔ movDate, amount, currency)` per
statement with the S-11 tie-break, S-7's cell set, `running_balance` only where `printsRunningBalance`, both
unpaired sides missed on every cell, the sixteen description corrections applied, statement cells against the
flags file. Per statement: BCP 325/325, BPI-CC 63/63, BPI-CRF 10/38, BPI-GOT 591/633, Revolut 360/360. Per
field identical to the report, including `running_balance` 182/192 and the six fields at 192/202. The
denominator decomposes as 187 × 7 + 10 × 6 (the card, `printsRunningBalance: false`) + 5 unpaired answer rows
× 7 + 15 statement cells = **1419**. Two things I checked deliberately because they changed this round: I
scored `movement_count` against the answer's own `statement.movements_printed_count` (R2-23) and it is still
5/5, and I implemented S-11's **new** bound (3 tokens and 25 % of characters, R2-8) and the total is unchanged.
Neither round-two change moved the published figure.

**Lines.** 342/347 exactly, with the same five misses (all I26070005), after implementing §2.6's `Split` — step
A, branches 0–3, the parafiscal fold, `VAT := VAT ?? 0` — and S-3's signs with the printed-sign exceptions.
One caveat, which is **R3-10**: implementing step C exactly as §2.6 writes it (`item.Net += other`) gives
340/347; the code also folds the surcharge into the item line's `Total`, which the spec does not say.

**Header.** I recomputed the 318 non-derived cells and got **317/318** — the published `line_count` miss on
I26070005 and nothing else — with every per-field tally identical (`document_id` 35/35, `date_doc` 35/35,
`date_due` 35/35 under S-4, `date_pay` 6/6 on the six `Invoice-Receipt` rows, `currency` 35/35, `net`/`vat`/
`total` 35/35 after `Split`, `fiscal_no` 32/32 through my own transcription of `SameTaxId`). I then recomputed
three of the five derived fields independently: `company` **29/29** and `origin_class` **29/29** by
reimplementing the two-sided gate over the six companies the keys name (GOTT, ITOO, CONF, FMAT, VIGA, SILA),
and `document_type` **34/35** with the single miss on I26020038 (`Invoice` vs `Invoice-Receipt`) — exactly the
report's other header miss. That is **411 of 447 cells verified directly, 445 matched, 2 missed**. The
remaining 36 (`account_period` 24, `service_period` 12) need `AccountPeriodService.Decide` and
`ExtractionDerivations.PeriodFor`, which I did not reimplement; I verified their denominators from the keys'
`accountPeriodRule` (12 R1 + 12 R3 = 24; the 12 R1 rows again for `service_period`), and R3-1 is the question
of whether 12 of them should be scored at all, not whether they match.

**Cross-checks that also reproduce:** 24 corrections applied and 24/2213 = 1.0845 % → "1.08 %"; 447 + 347 +
1419 = 2213; the blind reserve report's 289-cell matrix gives header 76/80 = 95.0000 %, lines 27/51 =
52.9412 %, movements 158/158, which is exactly what its new banner prints.

---

## What I verified as sound

- **Every round-two reproduction reverses.** `1348/1419` now prints FAILED and `1349/1419` PASSED (the decimal
  arithmetic is exact: `0.95m × 1419 = 1348.05m`); `GR…` matches `EL…` and `XI…` matches `GB…` through both
  `SameTaxId` and `Gate`, while `PT…`/`ES…` still do not; a "Verified genuine repeat" survives a later
  duplicate ruling with its clause, evidence, user and `declared_on` intact, and the sync-owned row is
  untouched by construction; a ruling on `printedOpeningBalance` moves 48 cells and **all 48** carry it; the
  R2-8 bound admits the sixteen real slices — I re-ran the whole statements set under it.
- **The fingerprint.** `git archive 3b15b5c | tar -x` then the committed script:
  `54d550b089ede3f69e2b4c31e94008a8b7a719577a8f1729c8c69821be19971d over 652 files` — character-for-character
  what `RED-green.txt:2348` claims, and **the same at HEAD `81ab4fc`** (the only file `81ab4fc` touches lives
  under `evidence/`, which the script excludes). The transcript's other new fingerprint also reproduces:
  `7b96422ed2aa1e5df698c47623e2ed3798db838b05d432d79d59b79059d49086 over 651 files` at `24e4a52`. `docs/` is a
  root now, so the specification is inside the fingerprint for the first time in the slice.
- **The recovered fiscal record is genuine, not fabricated.** Diffed against `git show a50e5d5:` — the same 35
  documents, the same 60 attempts, the same 60 evidence elements with the same `inputSha256`, `promptSha256`,
  `rawStdoutSha256`, `numTurns` and `durationMs` (5,050,898 ms both sides), and the true
  2026-09-07T13:21:11 → 14:45:22 window. Only `status`, `jobId` and `modelUsage` differ (R3-6).
- **The banner on the blind run is accurate**, including "passes, by nothing" for 76/80 — under the new exact
  comparison `76 >= 0.95 × 80` is true by zero margin.
- **The suites.** `Sibyla.Tests.Argus` **231/231**; `Sibyla.Tests.Platform` on the documented filter
  **250/250**; on the transcript's wider baseline filter **334/334**; on my own wider filter (adding
  `CompanyMatcher`, `WorkerStartup`, `QueueWorkerGate`, `IngestionServiceHold`) **341/341**. Both counts
  `397afb9` quotes reproduce. Build lock taken and released on every run.
- **The new oracles can fail.** `TheGateComparesTheUnroundedRateSoOneCellUnderNinetyFiveFails` runs the exact
  1348/1419 shape; `AKindWithNoCellBesideAPerfectlyScoredKindIsStillNotAPass` is the discriminator R2-16 asked
  for; `AReleaseWithoutTheTupleOrDisagreeingWithThePinnedPackageIsRefused` pins the publish script by
  expressions that exist only inside its enforcement block and checks the repository's own worker
  `appsettings.json` against the pinned golden; `AHostOverrideOfTheReleaseTupleStopsTheWorker` is a real
  three-case theory. The round-two RED transcript records genuine failing values, not compile errors, and says
  which three REDs had to be produced by reverting a fix in place.
- **`score` still touches nothing it should not.** The verb re-validates every answer through
  `ExtractionContractV2.Validate` before scoring, records a SHA-256 per file, refuses a path-shaped `--out`,
  refuses to overwrite without `--force`, refuses an incomplete set without `--partial`, and makes a mixed
  configuration tuple a gate reason rather than a note.

## What I could not verify, and why

- **The 36 remaining derived header cells** (`account_period` 24, `service_period` 12). I verified their
  denominators and read the code; I did not reimplement `AccountPeriodService.Decide` or
  `ExtractionDerivations.PeriodFor`.
- **The blind reserve run's answers.** Destroyed and never in git; I recomputed its rates from the surviving
  289-cell matrix and nothing else.
- **The round-two RED observations themselves.** They are the author's own account. Where I could check one
  against history (`NormalizeTaxId` with no call site at `21097cc^`; `HasMaxLength(20)` on the entry code) it
  held. I did not re-run them.
- **§4.5's hostile-document run.** Known blocked on the owner; out of scope.
- **Anything needing the production database, a deployment or the CLI.** `PreviewParityTests` is red pending
  three v2 migrations on Main (known, not my finding). R3-2's blast radius and R2-18's operational effect on
  the live host are arguments from the code; I read no host file and queried no database.
- **The other session's `Discarded` feature** carried whole by `21097cc`. Out of scope by instruction.

## Restrictions honoured

No secrets path, key ring, worker CLI home or production database was read.
`D:\fileStorage\repos\invoice-skill-build` was not touched. The Claude CLI was never called and nothing was
deployed. The build lock `D:/fileStorage/tmp/apollo-extraction/build.lock` was taken and released around every
build and test run. The only file I wrote inside the repository is this report; the scorer, the probe project
and the clean exports live under `D:\fileStorage\tmp\rev3`. One `python -` heredoc was started by mistake,
detected immediately and killed by pid; no orphan survives.
