aboutsummaryrefslogtreecommitdiff
path: root/tools/extract_of_calendar.py
blob: 50b570e0e686418dfba089c4c00e7389e1682435 (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
#!/usr/bin/env python3
# SPDX-License-Identifier: AGPL-3.0-or-later
"""tools/extract_of_calendar.py -- transcribes the Calendarium Romanum
Generale (General Roman Calendar) from the 2002 Missale Romanum, editio
typica tertia, into data/of/calendar-2002.sexp: a Colitur_kernel.Layer.t of
Colitur_kernel.Celebration.t entries, one per FIXED (month, day) universal
calendar entry.

SOURCE DISCIPLINE (do not weaken this): the 2002 Missale Romanum PDF is the
authority. lectio's ~/git/projects/lectio/internal/caldata/roman-calendar.ini
is read here ONLY for two cross-check purposes, never as structural source:
  (a) comparing (month, day, rank) to surface divergences, each of which is
      adjudicated by re-reading the Missal (see the divergence log this
      script prints), never by silently preferring lectio;
  (b) copying `name.en` -- lectio's OWN `name.la` field is independently
      known to be unreliable (it contains English text for at least one
      entry, verified during planning: mary-mother-of-god-octave-of-christmas
      has name.la = "Mary, Mother of God (Octave of Christmas)"), so Latin
      is NEVER taken from lectio, only English.

Run: `python3 tools/extract_of_calendar.py [pretext-file] > data/of/calendar-2002.sexp`
With no argument, runs `pdftotext -layout` on the PDF itself. A pre-extracted
text file may be given instead (for reproducibility without re-invoking
pdftotext).

WHAT THIS FILE DELIBERATELY EXCLUDES, and why (each is a real design
decision, not an oversight):

1. MOVABLE universal solemnities/feasts printed in the calendar table
   under a "Dominica ... :" / "Feria ... :" / "Sabbato ... :" heading. MOST
   of these (Baptism of the Lord, Holy Family, Trinity Sunday, Corpus
   Christi, Christ the King) have no (month, day) Date_spec.Fixed key at
   all, are ALREADY computed by lib/rites/rite_of/temporal_of.ml's own
   `named`/`holy_family`/`baptism_of_the_lord` (Phase 1, already merged),
   and are excluded here for that reason -- shipping them too would create
   two competing candidates for the same day.

   MOVABLE-HEADING AUDIT (fix round 1, 2026-08-25 -- the coordinator asked
   for a full accounting after the gap below was found, not just the two
   entries that closed it). There are EXACTLY 7 such headings in the whole
   calendar table (verified mechanically: every "Dominica/Feria/Sabbato
   ... :" line the table contains, listed by MOVABLE_HEADING_RE with no
   heuristic filtering). All 7, and their fate:
     - Baptism of the Lord (Jan)      -> Temporal_of.baptism_of_the_lord
     - Trinity Sunday (end of May)    -> Temporal_of.named (off 56)
     - Corpus Christi (end of May)    -> Temporal_of.named (off 60)
     - Sacred Heart of Jesus (Jun)    -> THIS FILE, Easter_offset 68 (below)
     - Immaculate Heart of Mary (Jun) -> THIS FILE, Easter_offset 69 (below)
     - Christ the King (Nov)          -> Temporal_of.christ_the_king
     - Holy Family (Dec)              -> Temporal_of.holy_family (its own
       heading carries a FIXED-DATE FALLBACK, "vel, ea deficiente, die 30
       decembris" -- correctly Temporal_of's job, not this file's, despite
       the prose shape; left alone on the coordinator's own confirmation)
   No 8th heading exists -- the audit found nothing else of this shape.

   RESOLVED (fix round 1): the Sacred Heart of Jesus ("Feria VI post
   dominicam secundam post Pentecosten: SACRATISSIMI CORDIS IESU
   Sollemnitas") and the Immaculate Heart of Mary ("Sabbato post dominicam
   secundam post Pentecosten: Immaculati Cordis B. Mariae Virginis
   Memoria") are printed in the SAME base-2002 calendar table as Trinity/
   Corpus Christi/Christ the King, but temporal_of.ml covered NEITHER --
   found and reported by this extractor's own first pass; researched and
   confirmed by the coordinator (extracted lines ~4219-4222, immediately
   after 30 June). They belong HERE, not in temporal_of.ml: both are
   printed in the Calendarium Romanum Generale itself, the exact table this
   file transcribes, so adding them completes the transcription rather than
   amending it. Produced as ordinary Date_spec.Easter_offset entries (the
   same variant EF's Rogation Wednesday already uses, and Task 2 will reuse
   for Mary, Mother of the Church -- no kernel change). Pentecost is
   Easter+49 (Normae n. 22-23); the Second Sunday after Pentecost is
   Easter+63; the Friday after it is Easter+68 (Sacred Heart), the Saturday
   after it Easter+69 (Immaculate Heart) -- verified against three real
   Easter dates in test_calendar_of_data.ml, not merely computed on paper.
   See MOVABLE_ENTRY_HEADINGS/matching_movable_entry_heading below for how
   the extractor recognises these two headings specifically (and only
   these two) among the other 5 it still correctly skips.

2. Three FIXED entries that ARE in the printed table but are ALSO already
   computed by temporal_of.ml's own `named`: 1 January (Mary, Mother of
   God), 6 January (Epiphany), 25 December (the Nativity). Shipping them
   here too would duplicate a candidate Precedence_of.band would then have
   to arbitrate between a temporal-origin and a data-origin copy of the
   IDENTICAL office -- pointless and risks a silent divergence between the
   two copies. Excluded explicitly (SKIP_DATES below), not merely absent
   by accident.

DATA DECISIONS made while transcribing, each cited in the emitted file's
own provenance header:

- A blank grade column means Memoria ad libitum, per the calendar's own
  footnote (`* Quando non indicatur gradus celebrationis, fit Memoria ad
  libitum.`) -- Normae/Missale Romanum 2002, Calendarium Romanum Generale,
  its own footnote marker.
- 2 November (Commemoratio omnium fidelium defunctorum, All Souls) carries
  NO grade word in the table at all -- applying the footnote naively would
  misclassify the second-most solemn day of November as an optional
  memorial. Normae n. 59's own Tabula dierum liturgicorum, entry 3, places
  it explicitly alongside "Sollemnitates ... in Calendario generali
  inscriptae" ("Sollemnitates Domini, beatae Mariae Virginis, et Sanctorum
  in Calendario generali inscriptae. Commemoratio omnium fidelium
  defunctorum.") -- same table entry, same precedence tier. Tagged
  Sollemnitas here on that citation, as an explicit, documented override,
  not the generic blank-grade default.
- Colour is derived from IGMR n. 346 (a)-(d): white for Christ's
  non-Passion celebrations / the BVM / Angels / non-martyr Saints / All
  Saints / John the Baptist's Nativity / John the Evangelist / Chair of
  Peter / Conversion of Paul (346(a), the last three named explicitly);
  red for Passion-related celebrations of the Lord, "festis nataliciis
  Apostolorum et Evangelistarum" (an Apostle's or Evangelist's own feast),
  and any celebration of Martyr Saints (346(b)); violet, chosen over the
  also-permitted black, for the Commemoration of All the Souls (346(d):
  "Assumi potest etiam in Officiis et Missis defunctorum" -- explicitly
  authorised for the dead; 346(a)'s white clause has no textual reach here,
  since All Souls is not "Sanctorum" (canonised Saints) at all). This
  DIVERGES from lectio, which colours 2 November white -- see the
  divergence log below; adjudicated colitur/violet on 346(d)'s explicit
  text, flagging lectio's white as reflecting a widespread modern PASTORAL
  custom the Missal's own words do not themselves authorise as the
  default. Three explicit exceptions to the apostle/evangelist-red rule
  (Chair of Peter, Conversion of Paul, John the Evangelist) and one to the
  martyr/apostle-red rule (three "In Dedicatione ..." church-dedication
  entries, always white) are coded as override tables, cited alongside the
  rule.

EXTRACTION ARTIFACT: some titles are printed letter-spaced by pdftotext
(e.g. "S . P e t r i D a m i a n i"). `collapse_letterspacing` mechanically
rejoins runs of >=4 single-character tokens, using punctuation (periods,
commas) and a lowercase-then-uppercase transition as word-boundary signals.
This CANNOT recover a lowercase-to-lowercase boundary (no case signal
survives the original single-spacing), so it is a first pass only; the sole
entry it fires on in this table (21 February, S. Petri Damiani) is then
hand-verified against the well-attested standard title and corrected
explicitly, with the exact reason recorded at the call site. The count of
entries needing this repair is printed to stderr.
"""

import hashlib
import re
import subprocess
import sys
import unicodedata
from datetime import date, timezone, datetime

PDF_PATH = "docs/research/of/missale-romanum-2002.pdf"
TOOL_PATH = "tools/extract_of_calendar.py"
LECTIO_INI = "../lectio/internal/caldata/roman-calendar.ini"

MONTHS = [
    "IANUARIUS", "FEBRUARIUS", "MARTIUS", "APRILIS", "MAIUS", "IUNIUS",
    "IULIUS", "AUGUSTUS", "SEPTEMBER", "OCTOBER", "NOVEMBER", "DECEMBER",
]
MONTH_NUM = {m: i + 1 for i, m in enumerate(MONTHS)}

ROW_RE = re.compile(r"^(Cal\.|Prid\.|Non\.|Idib\.|[IVXL]+)\s*(\d{1,2})\b\s*(.*)$")
MOVABLE_HEADING_RE = re.compile(r"^(Dominica|Feria\s+[IVXL]+|Sabbato)\b.*\b(post|ultima|infra)\b")
ENTRY_MARKER_RE = re.compile(r"^(S\.|Ss\.|B\.)\s")
FOOTNOTE_RE = re.compile(r"^\*\s*Quando non indicatur")

GRADE_WORDS = ["Sollemnitas", "Festum", "Memoria"]
GRADE_TO_RANK = {
    "Sollemnitas": "Sollemnitas",
    "Festum": "Festum",
    "Memoria": "Memoria_obligatoria",
    None: "Memoria_ad_libitum",
}

# Excluded because Phase 1 (lib/rites/rite_of/temporal_of.ml) already
# computes these as temporal-cycle offices. See module docstring, point 2.
SKIP_DATES = {
    (1, 1): "of-mary-mother-of-god (Temporal_of.named, m=1 dd=1)",
    (1, 6): "of-epiphany (Temporal_of.named, m=1 dd=6)",
    (12, 25): "of-nativity (Temporal_of.named, m=12 dd=25)",
}

# IGMR 346(a)'s own three explicit white exceptions to the apostle/
# evangelist-red rule.
WHITE_APOSTLE_EXCEPTIONS = {(1, 25), (2, 22), (12, 27)}

# Church-dedication feasts: always white, never coloured by the martyr/
# apostle words that happen to appear in their own titles (the apostles
# named are the church's patrons, not the day's own honouree).
DEDICATION_DATES = {(8, 5), (11, 9), (11, 18)}

# All Souls: see module docstring for the 346(d) citation.
ALL_SOULS_OVERRIDE = (11, 2)

# Explicit red override: the Exaltation of the Holy Cross is a celebration
# "Passionis Domini" (346(b)) even though its own title contains neither
# "martyr" nor "apostol"/"evangelist".
RED_OVERRIDES = {(9, 14)}

# subject = Lord: celebrations of the Lord not already excluded above,
# plus the three church-dedication feasts (a dedication is itself
# classified a "festum Domini" -- see CLAUDE.md's EF precedent, and
# lectio's own roman-calendar.ini tags all three `class = lord`).
LORD_DATES = {(2, 2), (3, 25), (8, 6), (1, 3), (9, 14)} | DEDICATION_DATES


def sha256_file(path):
    h = hashlib.sha256()
    with open(path, "rb") as f:
        h.update(f.read())
    return h.hexdigest()


def get_text(pdf_path):
    out = subprocess.run(["pdftotext", "-layout", pdf_path, "-"], capture_output=True, check=True)
    return out.stdout.decode("utf-8")


def collapse_letterspacing(line):
    """Mechanically rejoins a run of >=4 single-character tokens. Returns
    (new_line, changed). See module docstring's "EXTRACTION ARTIFACT" note
    for what this can and cannot recover."""
    tokens = line.split(" ")
    out = []
    i = 0
    changed = False
    while i < len(tokens):
        if len(tokens[i]) == 1:
            j = i
            while j < len(tokens) and len(tokens[j]) == 1:
                j += 1
            run = tokens[i:j]
            if len(run) >= 4:
                changed = True
                merged = "".join(run)
                merged = re.sub(r"(?<=[.,])(?=\S)", " ", merged)
                merged = re.sub(r"(?<=[a-zà-öø-ÿæœ])(?=[A-ZÀ-ÖØ-ÝÆŒ])", " ", merged)
                out.append(merged)
            else:
                out.extend(run)
            i = j
        else:
            out.append(tokens[i])
            i += 1
    return " ".join(out), changed


def bound_table(text):
    lines = text.split("\n")
    # First IANUARIUS before line 5000, immediately preceded (allowing for a
    # blank line) by "CALENDARIUM ROMANUM GENERALE" -- rules out the second,
    # Proprium Sanctorum run near line 17874.
    start = None
    for i, l in enumerate(lines[:5000]):
        if l.strip("\x0c").strip() == "IANUARIUS":
            # confirm CALENDARIUM ROMANUM GENERALE appears in the preceding
            # few lines
            window = "\n".join(lines[max(0, i - 4):i])
            if "CALENDARIUM ROMANUM GENERALE" in window:
                start = i
                break
    if start is None:
        raise SystemExit("could not find the start of the calendar table (IANUARIUS after "
                          "CALENDARIUM ROMANUM GENERALE, before line 5000)")
    end = None
    for i in range(start, min(len(lines), 5000)):
        if lines[i].strip() == "TABELLA TEMPORARIA":
            end = i
            break
    if end is None:
        raise SystemExit("could not find the end of the calendar table (TABELLA TEMPORARIA)")
    return start, end, lines[start:end]


def is_month_header(line):
    return line.strip("\x0c").strip() in MONTH_NUM


def extract_grade(text):
    """Splits a trailing grade word off `text`. Shared by fixed-date rows
    and movable-heading entries (parse_rows below) so both go through the
    identical rule: a printed grade word (Sollemnitas/Festum/Memoria) is
    stripped; its absence, or Jan 3's own trailing footnote asterisk, both
    fall through to the calendar's blank-grade footnote (Memoria ad
    libitum)."""
    grade = None
    for g in GRADE_WORDS:
        if text == g or text.endswith(" " + g):
            grade = g
            text = text[: -len(g)].strip()
            break
    if text.endswith("*"):
        text = text[:-1].strip()
    return text, grade


# Fix round 1 (coordinator, 2026-08-25): two movable universal celebrations
# are printed as prose HEADINGS between fixed-date rows, not as (kalends,
# day, title, grade) rows -- exactly the shape parse_rows' generic
# MOVABLE_HEADING_RE skip-and-discard branch was built for, which is why the
# original extraction silently dropped them (found and reported by the
# extractor's own author; researched and confirmed by the coordinator).
# Both belong HERE, not in temporal_of.ml: they are printed in the
# Calendarium Romanum Generale itself, the exact table this file
# transcribes, so adding them completes the transcription rather than
# amending it (contrast Holy Family, whose heading is likewise prose but
# whose FIXED-DATE FALLBACK, "vel, ea deficiente, die 30 decembris", makes
# it correctly Temporal_of's job -- see the re-audit note in the module
# docstring's "MOVABLE-HEADING AUDIT" section for the full accounting of
# all 7 such headings in the table).
#
#   Feria VI post dominicam secundam post Pentecosten:
#   SACRATISSIMI CORDIS IESU                      Sollemnitas
#   Sabbato post dominicam secundam post Pentecosten:
#   Immaculati Cordis B. Mariae Virginis           Memoria
#
# (extracted lines ~4219-4222, immediately after 30 June). Pentecost is
# Easter+49 (Normae n. 22-23, already temporal_of.ml's own citation); "the
# Second Sunday after Pentecost" is therefore Easter+63, the Friday after it
# Easter+68, the Saturday after it Easter+69 -- independently verified
# against three real Easter dates (2026-04-05, 2027-03-28, 2035-03-25) in
# test_calendar_of_data.ml, not merely computed on paper here.
#
# Date_spec.Easter_offset (built for EF's Rogation Wednesday) needs no
# kernel change and is the same mechanism Task 2 will use for Mary, Mother
# of the Church -- so these become ordinary Easter_offset entries in this
# file, produced by the extractor like every other entry, not hand-appended:
# a heading this specific is no less mechanical to recognise than a
# (kalends, day) row, and hand-appending would be the one entry pair in this
# file with no SHA-256-verifiable path back to the PDF text.
MOVABLE_ENTRY_HEADINGS = [
    # (heading-substring, Easter offset, weekday name for the test/sanity check)
    ("Feria VI post dominicam secundam post Pentecosten", 68, "Fri"),
    ("Sabbato post dominicam secundam post Pentecosten", 69, "Sat"),
]


def matching_movable_entry_heading(line):
    for substring, offset, weekday in MOVABLE_ENTRY_HEADINGS:
        if line.startswith(substring):
            return offset, weekday
    return None


def parse_rows(table_lines):
    """Returns (entries, movable_entries, letterspace_count, movable_audit)
    where entries is a list of dicts with month, day, latin, grade;
    movable_entries is the same shape but with easter_offset instead of
    month/day; movable_audit lists EVERY "Dominica/Feria/Sabbato ... :"
    heading found, with its own fate (produced here, or the reason it is
    skipped), for the re-audit this fix round asked for."""
    # Pass 1: strip form feeds and collapse letter-spacing line by line.
    letterspace_count = 0
    lines = []
    for raw in table_lines:
        l = raw.replace("\x0c", "")
        l2, changed = collapse_letterspacing(l)
        if changed:
            letterspace_count += 1
        lines.append(l2)

    entries = []
    movable_entries = []
    movable_audit = []
    current_month = None
    i = 0
    n = len(lines)
    while i < n:
        line = lines[i].strip()
        if line == "":
            i += 1
            continue
        if is_month_header(line):
            current_month = MONTH_NUM[line]
            i += 1
            continue
        if FOOTNOTE_RE.match(line):
            i += 1
            continue
        if MOVABLE_HEADING_RE.match(line):
            heading = line
            i += 1
            block = []
            while i < n:
                nxt = lines[i].strip()
                if (nxt == "" or is_month_header(nxt) or ROW_RE.match(nxt)
                        or MOVABLE_HEADING_RE.match(nxt) or FOOTNOTE_RE.match(nxt)):
                    break
                block.append(nxt)
                i += 1
            match = matching_movable_entry_heading(heading)
            if match is None:
                movable_audit.append({"heading": heading, "content": block, "produced": False})
                continue
            offset, weekday = match
            text = " ".join(block).strip()
            text, grade = extract_grade(text)
            movable_entries.append({"easter_offset": offset, "weekday": weekday, "latin": text, "grade": grade})
            movable_audit.append({"heading": heading, "content": block, "produced": True,
                                   "easter_offset": offset})
            continue
        rm = ROW_RE.match(line)
        if rm:
            if current_month is None:
                raise SystemExit(f"row before any month header: {line!r}")
            _kalend, day_s, rest = rm.groups()
            day = int(day_s)
            blocks = [[]]
            if rest.strip():
                blocks[0].append(rest.strip())
            i += 1
            while i < n:
                nxt_raw = lines[i]
                nxt = nxt_raw.strip()
                if (nxt == "" or is_month_header(nxt) or ROW_RE.match(nxt)
                        or MOVABLE_HEADING_RE.match(nxt) or FOOTNOTE_RE.match(nxt)):
                    break
                if ENTRY_MARKER_RE.match(nxt) and blocks[-1]:
                    blocks.append([nxt])
                else:
                    blocks[-1].append(nxt)
                i += 1
            for cblock in blocks:
                if not cblock:
                    continue
                text = " ".join(cblock).strip()
                text, grade = extract_grade(text)
                if text:
                    entries.append({"month": current_month, "day": day, "latin": text, "grade": grade})
            continue
        raise SystemExit(f"unrecognised line in calendar table (month={current_month}): {line!r}")

    return entries, movable_entries, letterspace_count, movable_audit


LATIN_TO_ASCII = str.maketrans({
    # Æ (uppercase ligature) is verified, across every occurrence in the
    # extracted calendar table, to appear ONLY inside an ALL-CAPS heading
    # ("SANCTÆ", "PRÆSENTATIONE", "ASSUMPTIONE BEATÆ") -- never as the
    # initial letter of a Title Case word -- so it maps to "AE" (both
    # capitals), not "Ae", to stay consistent with its own surrounding case.
    "æ": "ae", "Æ": "AE", "œ": "oe", "Œ": "OE", "ø": "o", "Ø": "O",
})


def normalize_latin(text):
    """æ -> ae, œ -> oe, matching the existing la-name convention already
    used by lib/rites/rite_of/temporal_of.ml's own holy_family_names/
    baptism_names (e.g. "Sanctae Familiae Iesu, Mariae et Ioseph"), not the
    ligature glyphs the PDF itself prints. Apostrophes and other punctuation
    are left untouched here -- they are real printed characters in the `la`
    field; only slugify()/title_slug() (which call this first) strip them
    further, for the slug alone."""
    return text.translate(LATIN_TO_ASCII)


def slugify(latin):
    s = normalize_latin(latin)
    s = unicodedata.normalize("NFKD", s)
    s = "".join(c for c in s if not unicodedata.combining(c))
    s = s.lower()
    s = re.sub(r"[^a-z0-9]+", "-", s)
    s = re.sub(r"-+", "-", s).strip("-")
    return s


# Words to drop when deriving a slug from a title, so slugs read as names
# rather than descriptions (kept short and stable, matching data/ef/
# sanctoral.sexp's own style, e.g. "agnes", "all-saints").
DROP_WORDS = {
    "s", "ss", "b", "in", "et", "de", "ad", "a", "the", "of",
}


def title_slug(latin, month, day):
    words = re.split(r"[\s,]+", normalize_latin(latin))
    kept = []
    for w in words:
        base = re.sub(r"\.$", "", w)
        if base.lower() in DROP_WORDS or base == "":
            continue
        kept.append(base)
        if len(kept) >= 6:
            break
    if not kept:
        kept = [f"of-{month:02d}-{day:02d}"]
    return slugify(" ".join(kept))


def classify_colour(month, day, latin):
    key = (month, day)
    if key == ALL_SOULS_OVERRIDE:
        return "Violet"
    if key in RED_OVERRIDES:
        return "Red"
    if key in DEDICATION_DATES:
        return "White"
    if "martyr" in latin.lower():
        return "Red"
    if key in WHITE_APOSTLE_EXCEPTIONS:
        return "White"
    lower = latin.lower()
    if "apostol" in lower or "evangelist" in lower:
        return "Red"
    return "White"


def classify_subject(month, day, latin):
    key = (month, day)
    if key in LORD_DATES:
        return "Lord"
    lower = normalize_latin(latin).lower()
    if key != (3, 19) and ("maria" in lower or "b.m.v" in lower or "bmv" in lower):
        return "Bvm"
    return "Saint"


def parse_lectio_ini(path):
    entries = []
    cur = None
    try:
        with open(path, encoding="utf-8") as f:
            raw = f.read()
    except FileNotFoundError:
        return entries
    for line in raw.split("\n"):
        line = line.rstrip("\n")
        if line.startswith("[") and line.endswith("]"):
            if cur is not None:
                entries.append(cur)
            cur = {"slug": line[1:-1]}
        elif "=" in line and cur is not None:
            k, _, v = line.partition("=")
            cur[k.strip()] = v.strip()
    if cur is not None:
        entries.append(cur)
    return [e for e in entries if "date" in e]


LECTIO_RANK_TO_VOCAB = {
    "solemnity": "Sollemnitas",
    "feast": "Festum",
    "memorial": "Memoria_obligatoria",
    "optional": "Memoria_ad_libitum",
}


def significant_words(text):
    words = re.split(r"[\s,.;]+", normalize_latin(text).lower())
    return {w for w in words if w and w not in DROP_WORDS and len(w) > 2}


def cross_check(entries, lectio_entries):
    """Returns (report_lines, name_en_by_id) where name_en_by_id maps
    id(entry dict) -> matched lectio name.en, and report_lines is the
    divergence log for the provenance header."""
    by_date = {}
    for le in lectio_entries:
        d = le.get("date", "")
        m = re.match(r"^(\d\d)-(\d\d)$", d)
        if not m:
            continue
        key = (int(m.group(1)), int(m.group(2)))
        by_date.setdefault(key, []).append(le)

    by_date_missal = {}
    for e in entries:
        by_date_missal.setdefault((e["month"], e["day"]), []).append(e)

    report = []
    name_en = {}
    lectio_slug = {}
    agree_rank = 0
    total_matched = 0
    matched_lectio_ids = set()

    for key, missal_list in sorted(by_date_missal.items()):
        lectio_list = by_date.get(key, [])
        if len(missal_list) == 1 and len(lectio_list) == 1:
            pairs = [(missal_list[0], lectio_list[0])]
        else:
            pairs = []
            remaining = list(lectio_list)
            for me in missal_list:
                if not remaining:
                    break
                mwords = significant_words(me["latin"])
                best = max(remaining, key=lambda le: len(mwords & significant_words(le.get("name.la", ""))))
                pairs.append((me, best))
                remaining.remove(best)
        for me, le in pairs:
            total_matched += 1
            matched_lectio_ids.add(id(le))
            name_en[id(me)] = le.get("name.en")
            lectio_slug[id(me)] = le.get("slug")
            missal_rank = GRADE_TO_RANK[me["grade"]]
            lectio_rank = LECTIO_RANK_TO_VOCAB.get(le.get("rank", ""), "?")
            if missal_rank == lectio_rank:
                agree_rank += 1
            else:
                report.append(
                    f"  {key[0]:02d}-{key[1]:02d} RANK  colitur={missal_rank} lectio={lectio_rank} "
                    f"[{me['latin']!r} vs {le.get('name.la','')!r}]"
                )
            missal_colour = classify_colour(key[0], key[1], me["latin"])
            lectio_colour = (le.get("colour", "") or "").capitalize()
            if missal_colour != lectio_colour:
                report.append(
                    f"  {key[0]:02d}-{key[1]:02d} COLOUR colitur={missal_colour} lectio={lectio_colour} "
                    f"[{me['latin']!r}]"
                )
        if not lectio_list:
            for me in missal_list:
                report.append(f"  {key[0]:02d}-{key[1]:02d} MISSAL-ONLY {me['latin']!r} (grade={me['grade']})")
        elif len(lectio_list) > len(missal_list):
            # A shared date where lectio carries MORE entries than the
            # Missal transcription found -- the pairing loop above only
            # consumes len(missal_list) of them, so the rest would
            # otherwise vanish silently. Surfaced explicitly.
            for le in lectio_list:
                if id(le) not in matched_lectio_ids:
                    report.append(
                        f"  {key[0]:02d}-{key[1]:02d} LECTIO-EXTRA (same date, uncovered by Missal pairing) "
                        f"{le.get('name.la','')!r} rank={le.get('rank')} slug={le.get('slug')}"
                    )

    missal_dates = set(by_date_missal)
    for key, lectio_list in sorted(by_date.items()):
        if key in SKIP_DATES:
            continue
        for le in lectio_list:
            if id(le) not in matched_lectio_ids and key not in missal_dates:
                report.append(f"  {key[0]:02d}-{key[1]:02d} LECTIO-ONLY {le.get('name.la','')!r} "
                               f"rank={le.get('rank')} slug={le.get('slug')}")

    summary = [
        f"lectio cross-check: {total_matched} dates paired; rank agrees on {agree_rank}/{total_matched}.",
    ]
    return summary + report, name_en, lectio_slug


def render_sexp(entries, name_en, letterspace_count, skipped, meta):
    lines = []
    lines.append(f"; data/of/calendar-2002.sexp -- OF (post-1970) General Roman Calendar,")
    lines.append(f"; transcribed from the Missale Romanum, editio typica tertia (2002). The")
    lines.append(f"; 2002 Missal is the AUTHORITY; this file is never edited afterwards --")
    lines.append(f"; spec sec4.1. lectio's roman-calendar.ini is a cross-check only, never a")
    lines.append(f"; source (see tools/extract_of_calendar.py's own header for the full")
    lines.append(f"; source-discipline statement and every classification citation).")
    lines.append(f";")
    lines.append(f"; Generator: {TOOL_PATH} -- do not hand-edit; re-run against the same")
    lines.append(f"; PDF (this file is pinned, so a re-run should reproduce it byte-for-byte)")
    lines.append(f"; and diff before committing.")
    lines.append(f";")
    lines.append(f"; Source: {PDF_PATH}")
    lines.append(f"; SHA-256: {meta['pdf_sha256']}")
    lines.append(f"; Extracted lines (pdftotext -layout, 0-based): {meta['start']}-{meta['end']}")
    lines.append(f"; Extraction date (UTC): {meta['extraction_date']}")
    lines.append(f";")
    lines.append(f"; {len(entries)} entries. {letterspace_count} entry needed letter-spacing repair")
    lines.append(f"; (collapse_letterspacing's own mechanical pass, then hand-verified -- see")
    lines.append(f"; the module docstring's EXTRACTION ARTIFACT note).")
    lines.append(f";")
    lines.append(f"; Excluded, deliberately (see module docstring for the full reasoning):")
    for (m, d), why in sorted(skipped.items()):
        lines.append(f";   {m:02d}-{d:02d}: {why}")
    lines.append(f";   Movable universal solemnities STILL excluded, correctly (Baptism of")
    lines.append(f";   the Lord, Holy Family, Trinity, Corpus Christi, Christ the King) --")
    lines.append(f";   Temporal_of code already covers each; see the module docstring's")
    lines.append(f";   MOVABLE-HEADING AUDIT for the full 7-heading accounting.")
    lines.append(f";")
    lines.append(f"; Movable entries produced from a prose heading, not a (kalends, day) row")
    lines.append(f"; (fix round 1, 2026-08-25 -- see MOVABLE_ENTRY_HEADINGS):")
    for me in meta.get("movable_entries", []):
        lines.append(f";   Easter_offset {me['easter_offset']} ({me['weekday']}): {me['latin']!r}")
    lines.append(f";")
    for l in meta["cross_check_report"]:
        lines.append(f"; {l}")
    lines.append("")

    def esc(s):
        return s.replace("\\", "\\\\").replace('"', '\\"')

    lines.append("((id of-universal) (name \"OF (2002) General Roman Calendar\")")
    lines.append("  (entries")
    lines.append("    (")
    for idx, e in enumerate(entries):
        la = normalize_latin(e["latin"])
        en = name_en.get(id(e))
        names_parts = []
        if en:
            names_parts.append(f'(en "{esc(en)}")')
        names_parts.append(f'(la "{esc(la)}")')
        names_str = " ".join(names_parts)
        rank = GRADE_TO_RANK[e["grade"]]
        if "colour" in e:
            colour = e["colour"]
        else:
            colour = classify_colour(e["month"], e["day"], e["latin"])
        if "subject" in e:
            subject = e["subject"]
        else:
            subject = classify_subject(e["month"], e["day"], e["latin"])
        slug = e["slug"]
        prefix = "      " if idx > 0 else "      "
        if "easter_offset" in e:
            date_sexp = f"(Easter_offset {e['easter_offset']})"
        else:
            date_sexp = f"(Fixed (month {e['month']}) (day {e['day']}))"
        lines.append(f"{prefix}((date {date_sexp})")
        lines.append(f"        (cel")
        lines.append(f"          ((slug {slug})")
        lines.append(f"            (names ({names_str}))")
        lines.append(f"            (rank {rank}) (status Feast) (colour {colour})")
        lines.append(f"            (subject {subject}) (citations ()) (layer of-universal))))")
    lines.append("    )))")
    return "\n".join(lines) + "\n"


# Fix round 1: hand-assigned, clean English-style slugs for the two
# movable entries, keyed by Easter offset. Not run through title_slug's
# Latin-mechanical fallback (which would produce "sacratissimi-cordis-iesu")
# because lectio -- the only source this file ever takes a slug FROM when
# one exists (see the "Slugs:" comment in main() below) -- carries neither
# entry at all (its own header states both are "computed by internal/
# calendar", so its roman-calendar.ini has no bracket slug to reuse). This
# is the one place in the file a slug is authored rather than derived --
# recorded here rather than left implicit.
MOVABLE_SLUG_OVERRIDE = {68: "sacred-heart-of-jesus", 69: "immaculate-heart-of-mary"}

# Fix round 1: subject/colour for the two movable entries, set directly
# rather than through classify_colour/classify_subject (which key off
# (month, day) -- meaningless for an Easter_offset entry). Both white
# (IGMR 346(a): "celebrationibus Domini quae non sint de eius Passione" for
# the Sacred Heart; the same clause's "beatae Mariae Virginis" for the
# Immaculate Heart -- neither is a Passion or martyr celebration). Subject
# Lord/Bvm respectively, read directly off each title's own referent.
MOVABLE_SUBJECT_OVERRIDE = {68: "Lord", 69: "Bvm"}


def main():
    if len(sys.argv) > 1:
        with open(sys.argv[1], encoding="utf-8") as f:
            text = f.read()
    else:
        text = get_text(PDF_PATH)

    start, end, table_lines = bound_table(text)
    entries, movable_entries, letterspace_count, movable_audit = parse_rows(table_lines)

    # Hand-verified correction for the one letter-spaced entry the
    # mechanical pass cannot fully resolve (see module docstring).
    for e in entries:
        if (e["month"], e["day"]) == (2, 21) and "episcopiet" in e["latin"]:
            e["latin"] = "S. Petri Damiani, episcopi et Ecclesiae doctoris"

    skipped = {}
    kept = []
    for e in entries:
        key = (e["month"], e["day"])
        if key in SKIP_DATES:
            skipped[key] = SKIP_DATES[key]
            continue
        kept.append(e)
    entries = kept

    # All Souls override (see module docstring).
    for e in entries:
        if (e["month"], e["day"]) == ALL_SOULS_OVERRIDE:
            e["grade"] = "Sollemnitas"

    for e in movable_entries:
        e["slug"] = MOVABLE_SLUG_OVERRIDE[e["easter_offset"]]
        e["colour"] = "White"
        e["subject"] = MOVABLE_SUBJECT_OVERRIDE[e["easter_offset"]]

    lectio_entries = parse_lectio_ini(LECTIO_INI)
    cross_check_report, name_en, lectio_slug = cross_check(entries, lectio_entries)

    # Slugs: prefer lectio's OWN identifier when a confident match exists --
    # it is already a clean, conventional, human-authored identifier
    # ("raymond-of-penyafort-priest") matching data/ef/sanctoral.sexp's own
    # style, and reusing it is not a source-discipline violation: a slug is
    # a stable ENGINEERING KEY, not liturgical content -- every substantive
    # field (date, rank, colour, the `la` name) still comes from the Missal
    # alone. Falls back to a mechanical slug built from the Latin title for
    # any entry lectio has no match for (none, as it happens, in this run --
    # every one of the 206 entries paired -- but kept for robustness against
    # a future lectio update dropping an entry). Collisions suffixed either
    # way.
    seen = {}
    for e in entries:
        candidate = lectio_slug.get(id(e))
        base = slugify(candidate) if candidate else title_slug(e["latin"], e["month"], e["day"])
        if not base:
            base = title_slug(e["latin"], e["month"], e["day"])
        if base in seen:
            seen[base] += 1
            e["slug"] = f"{base}-{seen[base]}"
        else:
            seen[base] = 1
            e["slug"] = base

    # Movable entries carry their own hand-assigned slugs already (see
    # MOVABLE_SLUG_OVERRIDE); still registered in `seen` and defensively
    # collision-checked like everything else, rather than assumed safe.
    for e in movable_entries:
        base = e["slug"]
        if base in seen:
            raise SystemExit(f"movable entry slug {base!r} collides with a fixed-date entry")
        seen[base] = 1

    entries = entries + movable_entries
    entries.sort(key=lambda e: e["slug"])

    meta = {
        "pdf_sha256": sha256_file(PDF_PATH),
        "start": start,
        "end": end,
        "extraction_date": datetime.now(timezone.utc).strftime("%Y-%m-%d"),
        "cross_check_report": cross_check_report,
        "movable_entries": movable_entries,
    }

    sys.stdout.write(render_sexp(entries, name_en, letterspace_count, skipped, meta))

    print(f"[extract_of_calendar] {len(entries)} entries emitted ({len(movable_entries)} of them "
          f"movable, Easter_offset), {letterspace_count} needed letter-spacing repair, "
          f"{len(skipped)} dates excluded (temporal_of.ml coverage)",
          file=sys.stderr)
    print(f"[extract_of_calendar] MOVABLE-HEADING AUDIT: {len(movable_audit)} \"Dominica/Feria/"
          f"Sabbato ... :\" headings found in the table:", file=sys.stderr)
    for ma in movable_audit:
        fate = (f"PRODUCED as Easter_offset {ma['easter_offset']}" if ma["produced"]
                else "skipped -- covered by Temporal_of code")
        print(f"[extract_of_calendar]   {ma['heading']!r} -> {fate}", file=sys.stderr)
    for l in cross_check_report:
        print(f"[extract_of_calendar] {l}", file=sys.stderr)


if __name__ == "__main__":
    main()