| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Reproduced the defect: reintroducing the exact historical citation bug
(pointing class-1's citation back at LT.txt:12459, the value a prior
fix round corrected away from) made the tool report "147 citations
checked, 0 look wrong". The mechanism was that the corrective comment
documenting the old bug quotes the wrong historical value, and the
checker pooled every quoted phrase from the whole surrounding comment
block, so citing the wrong line matched the comment explaining why it
was wrong.
Four changes:
1. The word pool for a citation is now scoped to the entry(ies) it is
attached to only -- never to quoted text elsewhere in the comment.
This is the direct fix for the self-poisoning bug.
2. A citation whose pool has fewer than two distinctive words (Latin
liturgical headings are short and stopword-heavy) cannot
discriminate the right line from a wrong nearby one. Such a
citation is now reported CANNOT VERIFY and fails the target,
instead of silently passing.
3. The blanket +-2-line tolerance is gone. A bare "LT.txt:N" is
checked at line N only; a heading that genuinely wraps must say so
explicitly as "LT.txt:N-M". The allowance moves into the data,
where it is visible.
4. The tool gets its own test suite, tools/test_check_citations.py,
with a synthetic fixture covering: a correct citation, off-by-one
and off-by-three mismatches, an explicit wrap range, a degenerate
pool, a PATTERN-marked entry with no citation, and a dedicated
regression test for the self-poisoning case itself. Wired into
`dune test` via a new (rule (alias runtest) ...) in tools/dune (a
plain (test ...) stanza cannot run a Python script), so it runs
with the rest of the suite, not only as a `make` target.
Added a --file/--lt-file override to check_citations.py so the tool
(and its own tests) can point at a fixture without touching the real
lang/la.ini or docs/research/LT.txt. Confirmed the "SKIPPED, exit 0"
behaviour for a missing docs/research/LT.txt is unchanged.
tools/__pycache__/ (a stray artefact of this script, previously
untracked and ungitignored) is now in .gitignore.
Measured against the current lang/la.ini (another task is still
landing its sanctoral entries on this branch): 15 of 275 citations now
look wrong and 42 more cannot be verified, both far above the 0 the
unhardened tool reported. Not fixed here -- the data pass is separate,
once the sanctoral entries land.
|
|
|
Two of la.ini's LT.txt:<n> citations pointed at the wrong line -- the
Latin itself was right, only the pinned line was wrong:
- advent cited LT.txt:8631 ("Tempus Nativitatis"); the real "Tempus
Adventus" heading is at 8609.
- ef-christ-the-king and [rank]'s own citation both pointed near
"Dominica ultima Octobris" (12459) when the text they actually quote,
"D.NI NOSTRI JESU CHRISTI REGIS" and "I classis", sits two and three
lines further down, at 12461 and 12462.
ef-christmas-sunday-0 was marked PATTERN but LT.txt:8644 is the identical
string verbatim -- relabelled as a direct citation, not constructed.
Added tools/check_citations.py and `make check-citations`: for every
LT.txt:<n> citation outside a PATTERN block, confirms a +-2-line window
around line n actually contains the Latin text the citation claims,
rather than trusting each of the 38 citations by hand. Follows
check-schema/check-templates' own precedent -- docs/ is gitignored, so
the target prints SKIPPED loudly and exits 0 when docs/research/LT.txt
is absent, never a silent pass.
The checker's own teeth are proven three ways: replayed against the
pre-fix file it independently re-derives both corrections above; a fresh
mutation (redirecting one citation to an unrelated line) is caught and
reverted; the fixed file passes clean, 147 citations checked, 0 wrong.
|