1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
|
#!/usr/bin/env python3
"""check_citations.py -- verify every "LT.txt:<n>" citation in lang/la.ini
actually resolves to the Latin text it claims, in docs/research/LT.txt.
Run via `make check-citations` (or directly: `python3 tools/check_citations.py
[--file LA_INI] [--lt-file LT_TXT]`). Exits 2 with a report if any citation
looks WRONG, is MALFORMED, or CANNOT BE VERIFIED; exits 0 (with a summary
line) only if every citation checked out cleanly; exits 0 with a loud
"SKIPPED" line if docs/research/LT.txt is not present (it is gitignored --
see below).
THIS SCRIPT WAS ITSELF FOUND TO BE SELF-POISONING (2026-08-19) AND HARDENED
------------------------------------------------------------------------
An earlier version pooled distinctive words from the whole COMMENT BLOCK
around a citation, including any double-quoted phrase the comment happened
to mention -- and comments routinely quote a WRONG historical value while
explaining a past fix (e.g. "CORRECTED: previously cited LT.txt:12459,
which is 'Dominica ultima Octobris', not this heading"). That quote landed
in the pool, so re-introducing the exact bug being documented -- citing
LT.txt:12459 again -- matched the very quote correcting it, and the script
reported "0 look wrong". A verification tool whose own documentation of a
fix defeats the check for that fix is worse than no tool: it manufactures
false confidence. Proven with a reproduction: reintroducing that one wrong
citation into a real copy of lang/la.ini produced zero findings on the
pre-hardening script. Four changes closed this, in order of how directly
each one addresses the reproduction:
1. THE POOL IS SCOPED TO THE CITATION'S OWN ENTRY, never to the
surrounding comment's quoted text. A citation is verified against the
name it claims, not against anything quoted nearby -- see
`distinctive_words` / the per-citation `pool_items` construction
below. This alone closes the reproduced bug (see `test_check_citations.py`'s
own `test_self_poisoning_quote_does_not_pass`).
2. A POOL TOO THIN TO VERIFY FAILS CLOSED (round 1's shape; round 2
replaced the mechanism -- see below -- but kept the discipline: no
match is ever reported as a silent PASS just because nothing better
was checked).
3. NO MORE BLANKET +-2-LINE TOLERANCE. A citation is checked at the EXACT
line it names. A heading that genuinely spans more than one physical
line in the source must say so explicitly, "LT.txt:8609-8610" -- the
allowance moves into the data, where a reader (and a future citation
added nearby) can see it, instead of silently forgiving ANY citation
within two lines of the truth. A bare "LT.txt:N" is checked at line N
only.
4. THIS SCRIPT NOW HAS ITS OWN TEST SUITE (test_check_citations.py,
wired into `dune test` via tools/dune's own runtest rule) -- the
single most important change. The reproduced bug survived as long as
it did specifically BECAUSE nothing exercised this script's own
logic against a known-wrong citation. The self-poisoning case above
is now a permanent regression test.
ROUND 2 (2026-08-19): A REVIEW DEFEATED ROUND 1 AGAIN -- FIVE MORE FIXES
------------------------------------------------------------------------
Round 1's own hardening had three further holes, all found live against
this exact file, plus two smaller defects. Fixed in this order (matching
the order the defeats were found in, most dangerous first):
1. "; PATTERN" SILENCED A WHOLE BLOCK, NOT JUST ITS OWN ENTRY. `check()`
used to test `"PATTERN" in block_comment` -- one substring search over
every comment in the block -- and `continue` past the ENTIRE block if
it matched anywhere. In a block that mixes a PATTERN-marked entry with
an ordinary entry carrying its own, genuinely wrong citation (exactly
[season]'s own trailing-comment style: several entries back-to-back
with no blank line between them), the wrong citation was never even
looked at. Fixed by scoping PATTERN the same way citation pools are
already scoped: a PASS over the block first collects which entries a
PATTERN-bearing comment actually covers (the one preceding entry for a
trailing comment, every entry in the block for a leading one -- the
identical leading/trailing rule `check()` already uses for citation
pooling), and only THOSE entries are excluded from checking. See
`test_pattern_does_not_silence_a_different_entry_in_the_same_block`.
2. EXPLICIT RANGES HAD NO UPPER BOUND. Round 1 replaced the old blanket
+-2-line tolerance with an explicit per-citation range
("LT.txt:8609-8610") specifically so a wrap allowance would be
visible in the data instead of invisible in the checker -- but set no
cap on how WIDE that range could be. "LT.txt:8600-8650" passed if the
claimed text appeared anywhere across fifty lines: the same
"tolerance forgives a wrong line" defect the +-2 removal was meant to
close, reintroduced at 25x the radius, now with the data's own
apparent blessing. Fixed with `MAX_RANGE_WIDTH` (see below): a range
wider than a genuine heading wrap (1-2 extra physical lines) is
rejected as MALFORMED, naming the entry and the width, rather than
silently accepted. See `test_range_wider_than_cap_is_malformed`.
3. THE "TOO THIN TO VERIFY" GATE COUNTED WORDS, NOT RARITY. Round 1
required a pool item to have >=2 distinctive words before it could
even be checked, on the theory that a one-word match proves nothing.
Measured against the real data: this flagged 11 CORRECT citations
CANNOT VERIFY -- every one a short Latin hagionym whose Missal
heading genuinely has only one non-stopword ("S. Antonii Abb.": only
"antonii" survives the stopword filter) -- while a match on nothing
but "classis" (507 occurrences across LT.txt) passed the >=2 bar
freely whenever it shared a citation with three siblings ("I classis
/ II classis / III classis / IV classis"). Word COUNT was never the
right proxy; word RARITY is. Replaced with `build_frequency_table` +
`item_evidence`: a token seen n times in the whole corpus contributes
1/n of evidence, an item's evidence is its single RAREST matched
token (not a sum -- see `item_evidence`'s own docstring for why
summing would reopen this exact class of bug), and `EVIDENCE_THRESHOLD`
is the bar a match must clear to count as real proof rather than
coincidence. See `test_rare_single_word_match_passes` and
`test_common_word_only_match_is_cannot_verify`.
4. A COMMA TYPO SILENTLY BECAME A DIFFERENT CITATION. "LT.txt:12,459" (a
stray thousands-separator comma for the single number "12459") used
to parse as TWO independent bare citations, "12" and "459" -- either
of which might coincidentally match somewhere nearby while the
intended line was never checked at all. `parse_citation_spec` now
recognises the shape (a 1-2 digit token immediately followed by an
exactly-3-digit token -- the only way a thousands-grouped LT.txt line
number, which never exceeds 5 digits, can be split by one comma) and
rejects the whole spec as MALFORMED rather than silently reinterpreting
it. See `test_thousands_separator_typo_is_malformed`.
5. THE SELF-TEST SUITE OVERSTATED ITS OWN COVERAGE. A review ran round
1's seven fixture cases against the PRE-round-1 script: only two
(`test_off_by_one_line_fails`, `test_self_poisoning_quote_does_not_pass`)
actually failed on it -- the rest passed on both sides of round 1's
fix and regression-tested nothing despite their names (a third,
`test_wrap_range_required_not_just_first_line`, was added after the
"seven" and also turns out to be a genuine regression test, confirmed
the same way). Every test in this suite is now labelled REGRESSION
(shown, not just claimed, to fail against a named prior version) or
CHARACTERISATION (pins current behaviour; does not fail on the prior
version, usually because it exercises a data shape or API surface
that prior version did not have at all) -- see
`test_check_citations.py`'s own module docstring for the full,
per-test account and how each label was actually verified.
MALFORMED, PRECISELY (round 2's third finding class, alongside WRONG and
CANNOT VERIFY): a citation whose SYNTAX cannot be trusted even before its
content is checked -- a range wider than `MAX_RANGE_WIDTH`, a backwards
range, or the thousands-separator-typo shape above. Reported, counted, and
fails the run exactly like a WRONG citation: rejecting the syntax rather
than guessing at the author's intent is the whole point (see fix 2 and
fix 4 above) -- "the author either finds the real line or marks it PATTERN
honestly" is not achieved by the checker silently picking a reading.
WHAT THIS CHECKS, PRECISELY (a heuristic, not a proof)
-------------------------------------------------------
lang/la.ini's own comments cite a Missal heading in one of two shapes:
1. A LEADING comment block, then a group of entries it covers, e.g.
"; Ash Wednesday and the three days after it -- LT.txt:8686-8689."
followed by four `key = value` lines. The pool for every citation
found in such a comment is EVERY entry in the group (there is no
positional correspondence encoded in the data between a particular
cited line and a particular entry in the list).
2. A TRAILING comment immediately under the ONE entry it explains, e.g.
"advent = Tempus Adventus" then "; LT.txt:8609." on the next line,
with no blank line -- [season]'s own style. The pool is that one entry
alone.
A "PATTERN" marker follows the identical leading/trailing rule (see round
2 fix 1 above): it excludes only the entry (or entries) it is itself
attached to from citation-checking, never the rest of the block.
A citation is either a bare line number ("LT.txt:8609") or an explicit
range ("LT.txt:8609-8610", capped at `MAX_RANGE_WIDTH` lines) for a heading
that genuinely wraps across physical lines in the source; a comma-separated
list ("LT.txt:8618,8620,8622") is several independent citations, each
checked on its own -- unless the list itself looks like a thousands-typo
for one number (round 2 fix 4), in which case the whole spec is MALFORMED.
For a bare number the window is that one line; for a range it is the union
of every line in the range (inclusive). There is no other tolerance.
For each citation, the POOL is the set of candidate Latin phrases it could
be defending: the entry (trailing shape) or every entry in the group
(leading shape) -- nothing pulled from quoted prose elsewhere in the
comment (see the self-poisoning account above). Each pool item keeps its
OWN distinctive-word set (>=4 letters, not on the small stopword list
below, j/i and ae/oe/diacritics normalised) -- items are never flattened
into one shared bag, for the same reason quotes were removed: a citation
bundling two claims onto one line number must not pass on the strength of
an unrelated pool item's words.
A pool item is a candidate MATCH only if the cited window's text contains
ALL of its distinctive words -- otherwise it is not a match at all: it
never contributed to a PASS and is not what the citation is checked at
all. Among pool items that ARE contained in the window, the citation
PASSES if at least one clears `EVIDENCE_THRESHOLD` (see round 2 fix 3
above and `item_evidence`'s own docstring) -- i.e. contains a word rare
enough, across the whole LT.txt corpus, to be real evidence rather than
coincidence. A citation whose only contained items are all common-word-only
is CANNOT VERIFY, not a silent PASS: found, but not proven. A citation with
NO contained item at all -- the claimed words are simply not at the cited
line(s) -- is WRONG.
This is deliberately a LOOSE, word-overlap check within the (now exact and
bounded) window, not a byte-exact phrase match: la.ini spells abbreviations
out in full (Sanctissimi, not Ss.mi) and normalises j->i, and requiring a
byte-exact substring would either force every citation's prose to repeat
the raw OCR text verbatim (defeating the point of writing readable
comments) or produce false failures having nothing to do with a wrong line
number. What it proves is narrower than a byte-exact match, and is
disclosed as such: PASS means "the claimed name's distinctive words are
present, in full, at the exact line(s) cited, and at least one of them is
rare enough in the corpus to be real evidence" -- not that the citation is
the best possible line, only that it is not obviously wrong and is not
resting on a coincidence-prone common word.
"""
import argparse
import re
import sys
import unicodedata
from pathlib import Path
ROOT = Path(__file__).resolve().parent.parent
DEFAULT_LA_INI = ROOT / "lang" / "la.ini"
DEFAULT_LT_TXT = ROOT / "docs" / "research" / "LT.txt"
STOPWORDS = {
"in", "de", "et", "ad", "post", "ante", "cum", "per", "seu", "infra",
"vel", "si", "haec", "hoc", "hic", "qui", "quae", "quod", "quia",
"tempus", "dominica", "dominicam", "dominicae", "feria", "feriae",
"sabbato", "sabbatum", "die", "diebus", "eodem", "anno", "eius",
"sancti", "sancta", "sanctae", "sancto", "sanctorum", "sanctus",
"domini", "dominus", "octava", "octavam", "octavas",
"missae", "missa", "proprium", "gregorianus", "cantus", "pdf",
"forma", "longior", "brevior", "vide", "etiam", "dom", "prosper",
"sacro", "actio", "electronica", "formam", "novissimae", "variationes",
"copyright", "archivum", "liturgicum", "missale", "romanum", "index",
"www", "http", "https", "htm", "html", "com", "romanum", "text",
}
# A citation's explicit range is capped at this many lines (n .. n+2). A
# genuine heading wrap in the source is one or two extra physical lines;
# anything wider is not a wrap, it is a search over a neighbourhood wide
# enough to coincidentally contain almost any short phrase -- exactly the
# blanket +-2 tolerance round 1 removed, reintroduced at a much larger
# radius by an unbounded range. See round 2 fix 2 in the module docstring.
MAX_RANGE_WIDTH = 3
# The evidence bar a pool item's RAREST matched word must clear to count
# as real proof (see `item_evidence` below). 1/200: a word occurring up to
# ~200 times across the whole ~118,000-token LT.txt corpus can still be
# the deciding, sole distinctive word of a short Missal calendar-table
# entry (measured: "omnium", the only survivor in "Omnium Sanctorum",
# occurs 139 times across the whole document, mostly in unrelated legal
# prose -- "of all" is common Latin furniture -- yet is genuinely the
# correct, sole citable word for that one heading). "classis" (507
# occurrences), the round 1 false-pass this rule specifically targets,
# sits comfortably over 2.5x past this bar and stays CANNOT VERIFY.
# Round-2's own measurement of every affected real citation in
# lang/la.ini is in the branch report.
EVIDENCE_THRESHOLD = 1.0 / 200
def normalize_word(w: str) -> str:
w = w.lower()
w = unicodedata.normalize("NFKD", w)
w = "".join(c for c in w if not unicodedata.combining(c))
w = w.replace("æ", "ae").replace("œ", "oe")
w = re.sub(r"[^a-z]", "", w)
w = w.replace("j", "i")
return w
def distinctive_words(text: str) -> set:
out = set()
for tok in re.split(r"\s+", text):
w = normalize_word(tok)
if len(w) >= 4 and w not in STOPWORDS:
out.add(w)
return out
def build_frequency_table(lt_lines: list) -> dict:
"""Count how many times each normalised word occurs anywhere in the
whole LT.txt corpus -- built ONCE per run (not per citation) and
consulted by `word_evidence`/`item_evidence` below. This is what lets
the checker tell a token that could only ever mean one heading (occurs
once in 26,000+ lines) apart from common liturgical furniture (occurs
hundreds of times) -- see round 2 fix 3 in the module docstring."""
freq = {}
for line in lt_lines:
for tok in re.split(r"\s+", line):
w = normalize_word(tok)
if w:
freq[w] = freq.get(w, 0) + 1
return freq
def word_evidence(word: str, freq: dict) -> float:
"""A token seen n times contributes 1/n: a hapax (n=1) contributes 1.0
-- about as conclusive as a word-overlap check can be -- and a word
occurring hundreds of times contributes next to nothing. A plain
reciprocal is chosen over a logarithmic/IDF scale deliberately:
measured against this corpus, log-scaling compresses "occurs once" and
"occurs 500 times" into a difference of a few units, too close to
cleanly separate "essentially conclusive" from "unverified" with a
single threshold; a reciprocal keeps them many orders of magnitude
apart, which is the actual claim this rule makes."""
n = freq.get(word, 0)
return 1.0 / n if n > 0 else 0.0
def item_evidence(words: set, freq: dict) -> float:
"""A pool item's rarity evidence is the SINGLE RAREST word it matched
on -- not a sum over all its words. Summing would let several
merely-uncommon words add up to "enough" evidence between them, which
is exactly the shape round 2 fix 3 exists to close (a match consisting
only of common tokens must stay CANNOT VERIFY "regardless of how many
words sit beside it") and exactly the self-poisoning failure mode from
round 1 (quoted corrective prose tends to share several ordinary words
with its neighbour, never one rare one). One genuinely rare word is
real evidence; an accumulation of merely-uncommon words is not the
same thing and must not be treated as if it were."""
if not words:
return 0.0
return max(word_evidence(w, freq) for w in words)
class CitationRef:
"""One citation token: 'label' is what the data actually wrote
("8609" or "8609-8610"); 'lines' is the fully-expanded, sorted list of
line numbers the label names -- a single-element list for a bare
number, the whole inclusive range for an explicit wrap."""
__slots__ = ("label", "lines")
def __init__(self, label, lines):
self.label = label
self.lines = lines
def _looks_like_thousands_typo(tokens) -> bool:
"""'12,459' splits into ('12', '459') -- the classic shape of a human
thousands-separator typo for a single 4-5 digit LT.txt line number:
grouped from the right in chunks of 3, a 4-digit number groups as
'N,NNN' and a 5-digit number as 'NN,NNN' (LT.txt never reaches 6
digits, so a 3-digit leading group never arises from real grouping).
Checked against every real comma-list citation in lang/la.ini as of
round 2: none has a 1-2 digit token immediately followed by an
exactly-3-digit token, so this shape is unambiguous enough in practice
to reject outright as malformed rather than silently parsing as two
unrelated short citations (round 2 fix 4)."""
for a, b in zip(tokens, tokens[1:]):
if re.fullmatch(r"\d{1,2}", a) and re.fullmatch(r"\d{3}", b):
return True
return False
def parse_citation_spec(spec: str):
"""'8618,8620,8622' -> three exact-line CitationRefs; '8609-8610' -> one
CitationRef spanning both lines (an explicit wrapped heading, capped at
MAX_RANGE_WIDTH lines). Returns (refs, malformed): `refs` is the list of
successfully-parsed CitationRef objects; `malformed` is a list of
{"token", "reason"} dicts for anything that did NOT parse into a
trustworthy citation -- a backwards range, a range wider than the cap,
the thousands-typo shape above, or an unparseable token. Round 1
silently DROPPED all of these (no crash, but no report either); round 2
surfaces every one instead, because a malformed citation naming a real
entry deserves a human's attention exactly as much as a wrong one does
(see round 2 fixes 2 and 4 in the module docstring)."""
raw_tokens = [t.strip() for t in spec.split(",")]
if _looks_like_thousands_typo(raw_tokens):
return [], [
{
"token": spec,
"reason": (
"looks like a thousands-separator typo for one number "
"(a 1-2 digit token immediately followed by a 3-digit "
"one) rather than a genuine list of citations -- "
"remove the comma, or split into real citations"
),
}
]
refs = []
malformed = []
for tok in raw_tokens:
if not tok:
continue
m = re.fullmatch(r"(\d{2,6})-(\d{2,6})", tok)
if m:
a, b = int(m.group(1)), int(m.group(2))
width = b - a + 1
if a > b:
malformed.append({"token": tok, "reason": f"backwards range (LT.txt:{tok})"})
elif width > MAX_RANGE_WIDTH:
malformed.append(
{
"token": tok,
"reason": (
f"range is {width} lines wide (max {MAX_RANGE_WIDTH}) -- "
"not a genuine heading wrap; find the real line or mark PATTERN honestly"
),
}
)
else:
refs.append(CitationRef(tok, list(range(a, b + 1))))
continue
m = re.fullmatch(r"(\d{2,6})", tok)
if m:
refs.append(CitationRef(tok, [int(m.group(1))]))
continue
malformed.append({"token": tok, "reason": "unparseable citation token"})
return refs, malformed
CITATION_RE = re.compile(r"LT\.txt:\s*((?:\d{2,6}(?:-\d{2,6})?)(?:\s*,\s*\d{2,6}(?:-\d{2,6})?)*)")
def parse_blocks(la_ini_text: str):
"""Split la.ini into blocks on blank lines and [section] headers. Each
block is a list of (kind, content) where kind is 'comment' or 'entry',
content is the stripped comment text or (key, value)."""
blocks = []
cur = []
for raw in la_ini_text.split("\n"):
line = raw.rstrip("\n")
stripped = line.strip()
if stripped == "" or stripped.startswith("["):
if cur:
blocks.append(cur)
cur = []
continue
if stripped.startswith(";"):
cur.append(("comment", stripped[1:].strip()))
elif "=" in stripped:
k, _, v = stripped.partition("=")
cur.append(("entry", (k.strip(), v.strip())))
# anything else (shouldn't occur) is ignored
if cur:
blocks.append(cur)
return blocks
def window_text_for(lt_lines, lines):
lo, hi = lines[0], lines[-1]
lo_c, hi_c = max(1, lo), min(len(lt_lines), hi)
if lo_c > hi_c:
return ""
return " ".join(lt_lines[lo_c - 1 : hi_c])
def _pool_entries_for(entries_before, entries):
"""The leading/trailing pooling rule, shared identically by citation
checking AND PATTERN scoping (round 2 fix 1): a TRAILING comment (one
or more entries already seen in this block) pools against only the
most recent entry; a LEADING comment (no entry seen yet) pools against
every entry in the block. The same rule must decide both questions --
a PATTERN marker and a citation attached to the same comment always
cover the same entries, by construction."""
return [entries_before[-1]] if entries_before else entries
def _pattern_marked_entries(block, entries):
"""Which entries in this block are excluded from citation-checking by
their OWN "PATTERN" marker -- never by a PATTERN marker attached to a
DIFFERENT entry in the same block. Returns a set of `id()` of the
entry (key, value) tuples (safe: each entry tuple is a single object,
shared by reference between `block` and `entries`, never copied).
This is a first, standalone pass over the block, completed before any
citation is evaluated, so a citation's own PATTERN status never
depends on where in the block it happens to sit relative to its
entry's PATTERN comment."""
marked = set()
entries_before = []
for k, c in block:
if k == "comment":
if "PATTERN" in c:
for e in _pool_entries_for(entries_before, entries):
marked.add(id(e))
else:
entries_before.append(c)
return marked
def check(la_ini_text: str, lt_lines: list):
"""Returns a dict: checked (int), passed (int), findings (list --
wrong citations), unverifiable (list -- no matched item clears the
evidence bar), malformed (list -- citation syntax itself could not be
trusted). findings, unverifiable and malformed are all things a human
must look at; only 'passed' citations required no human attention."""
checked = 0
passed = 0
findings = []
unverifiable = []
malformed = []
freq = build_frequency_table(lt_lines)
for block in parse_blocks(la_ini_text):
entries = [c for k, c in block if k == "entry"]
if not entries:
continue
pattern_entries = _pattern_marked_entries(block, entries)
entries_before = []
for k, c in block:
if k == "comment":
for m in CITATION_RE.finditer(c):
pool_entries = _pool_entries_for(entries_before, entries)
if all(id(e) in pattern_entries for e in pool_entries):
# Every entry this citation could be defending is
# itself PATTERN-marked: this "LT.txt:N" is not a
# provenance claim for any of them (e.g. citing the
# GRAMMAR another day's heading attests), so
# checking it would produce a meaningless failure.
continue
entry_desc = ", ".join(f"{k2}={v2}" for k2, v2 in pool_entries)
pool_items = [distinctive_words(v2) for _, v2 in pool_entries]
nonempty_items = [p for p in pool_items if p]
pool_words = sorted(set().union(*pool_items)) if pool_items else []
refs, bad_tokens = parse_citation_spec(m.group(1))
for bad in bad_tokens:
checked += 1
malformed.append(
{
"label": bad["token"],
"entries": entry_desc,
"reason": bad["reason"],
}
)
for ref in refs:
checked += 1
window_text = window_text_for(lt_lines, ref.lines)
window_words = distinctive_words(window_text)
if not nonempty_items:
# Every pool item is fully stopwords -- there is
# nothing to check either way. Absence of
# evidence is not evidence of correctness.
unverifiable.append(
{
"label": ref.label,
"entries": entry_desc,
"reason": "no distinctive words in the entry text to check at all",
"pool_words": pool_words,
}
)
continue
contained_items = [p for p in nonempty_items if p <= window_words]
strong_items = [
p for p in contained_items if item_evidence(p, freq) >= EVIDENCE_THRESHOLD
]
if strong_items:
passed += 1
elif contained_items:
best = max(item_evidence(p, freq) for p in contained_items)
unverifiable.append(
{
"label": ref.label,
"entries": entry_desc,
"reason": (
f"matched, but every matched word is too common to trust "
f"(best evidence {best:.4f}, need >= {EVIDENCE_THRESHOLD:.4f})"
),
"pool_words": pool_words,
}
)
else:
n = ref.lines[0]
findings.append(
{
"label": ref.label,
"entries": entry_desc,
"actual": (
lt_lines[n - 1].strip()
if 1 <= n <= len(lt_lines)
else "(out of range)"
),
"window": window_text.strip()[:160],
}
)
else:
entries_before.append(c)
return {
"checked": checked,
"passed": passed,
"findings": findings,
"unverifiable": unverifiable,
"malformed": malformed,
}
def build_arg_parser():
p = argparse.ArgumentParser(description=__doc__.splitlines()[0])
p.add_argument(
"--file",
dest="la_ini",
type=Path,
default=DEFAULT_LA_INI,
help=f"the la.ini-shaped file to check (default: {DEFAULT_LA_INI})",
)
p.add_argument(
"--lt-file",
dest="lt_txt",
type=Path,
default=DEFAULT_LT_TXT,
help=f"the LT.txt transcription to check against (default: {DEFAULT_LT_TXT})",
)
return p
def main(argv=None):
args = build_arg_parser().parse_args(argv)
if not args.lt_txt.exists():
print(
f"SKIPPED: {args.lt_txt} is absent (docs/ is gitignored -- "
"present locally only). Citations are NOT verified this run."
)
return 0
la_ini_text = args.la_ini.read_text(encoding="utf-8")
lt_lines = args.lt_txt.read_text(encoding="utf-8", errors="replace").split("\n")
result = check(la_ini_text, lt_lines)
checked = result["checked"]
findings = result["findings"]
unverifiable = result["unverifiable"]
malformed = result["malformed"]
if not findings and not unverifiable and not malformed:
print(f"check-citations: {checked} LT.txt citations checked, 0 look wrong.")
return 0
print(
f"check-citations: {len(findings)} of {checked} citations look wrong, "
f"{len(malformed)} MALFORMED, {len(unverifiable)} CANNOT VERIFY (see below):\n"
)
if findings:
print("WRONG:\n")
for f in findings:
print(f" LT.txt:{f['label']} cited for [{f['entries']}]")
print(f" actual: {f['actual']!r}")
print(f" window: {f['window']!r}\n")
if malformed:
print("MALFORMED (citation syntax itself is untrustworthy):\n")
for m in malformed:
print(f" LT.txt:{m['label']} cited for [{m['entries']}]")
print(f" {m['reason']}\n")
if unverifiable:
print("CANNOT VERIFY (a human must adjudicate these by hand):\n")
for u in unverifiable:
words = ", ".join(u["pool_words"]) if u["pool_words"] else "(none)"
print(f" LT.txt:{u['label']} cited for [{u['entries']}]")
print(f" {u['reason']}; pool words: {words}\n")
return 2
if __name__ == "__main__":
sys.exit(main())
|