#!/usr/bin/env python3 # SPDX-License-Identifier: AGPL-3.0-or-later """tools/extract_of_calendar.py -- transcribes the Calendarium Romanum Generale (General Roman Calendar) from the 2002 Missale Romanum, editio typica tertia, into data/of/calendar-2002.sexp: a Colitur_kernel.Layer.t of Colitur_kernel.Celebration.t entries, one per FIXED (month, day) universal calendar entry. SOURCE DISCIPLINE (do not weaken this). The 2002 Missale Romanum PDF is the SOLE authority for every SUBSTANTIVE field: which day exists at all, its (month, day), its rank/grade, its colour, its subject, and every character of its `la` name. None of those is ever taken from, adjusted by, or defaulted to lectio. WHAT ACTUALLY COMES FROM LECTIO, stated exactly rather than summarised (fix round 2, coordinator, 2026-08-25: a prior version of this note said lectio is read "ONLY for two cross-check purposes", which undercounted -- slugs were a third, unlisted use, and are the file's own PRIMARY KEY, not a minor detail): (a) `en` names -- lectio's OWN `name.la` field is independently known to be unreliable (it contains English text for at least one entry, verified during planning: mary-mother-of-god-octave-of-christmas has name.la = "Mary, Mother of God (Octave of Christmas)"), so Latin is NEVER taken from lectio, only English, and only when a confident per-date match exists (see cross_check's own pairing logic); (b) SLUGS -- reused verbatim from lectio's own bracket identifier when a confident match exists (main()'s own "Slugs:" comment explains why this is not an authority violation: a slug is an engineering key, not liturgical content); falls back to a mechanical slug built from the Latin title otherwise; (c) the (month, day, rank) COMPARISON itself, whose divergences are listed in this file's own provenance header, each adjudicated by RE-READING THE MISSAL, never by preferring lectio's answer. Both (a) and (b) are ENGINEERING/PRESENTATION conveniences layered on top of Missal-sourced content, not competing sources for that content -- but they are real dependencies on an external, unpinned-by-Git sibling repo, which is why lectio's own SHA-256 is now pinned in the emitted provenance header (below) and why a missing lectio file is now a hard error (see parse_lectio_ini) rather than a silent zero-entries degradation: running without lectio present used to exit 0 with a DIFFERENT SHA-256, no `en` names, and 206 mechanical Latin slugs, with nothing printed to say so. Run: `python3 tools/extract_of_calendar.py [pretext-file] > data/of/calendar-2002.sexp` With no argument, runs `pdftotext -layout` on the PDF itself. A pre-extracted text file may be given instead (for reproducibility without re-invoking pdftotext). WHAT THIS FILE DELIBERATELY EXCLUDES, and why (each is a real design decision, not an oversight): 1. MOVABLE universal solemnities/feasts printed in the calendar table under a "Dominica ... :" / "Feria ... :" / "Sabbato ... :" heading. MOST of these (Baptism of the Lord, Holy Family, Trinity Sunday, Corpus Christi, Christ the King) have no (month, day) Date_spec.Fixed key at all, are ALREADY computed by lib/rites/rite_of/temporal_of.ml's own `named`/`holy_family`/`baptism_of_the_lord` (Phase 1, already merged), and are excluded here for that reason -- shipping them too would create two competing candidates for the same day. MOVABLE-HEADING AUDIT (fix round 1, 2026-08-25 -- the coordinator asked for a full accounting after the gap below was found, not just the two entries that closed it). There are EXACTLY 7 such headings in the whole calendar table (verified mechanically: every "Dominica/Feria/Sabbato ... :" line the table contains, listed by MOVABLE_HEADING_RE with no heuristic filtering). All 7, and their fate: - Baptism of the Lord (Jan) -> Temporal_of.baptism_of_the_lord - Trinity Sunday (end of May) -> Temporal_of.named (off 56) - Corpus Christi (end of May) -> Temporal_of.named (off 60) - Sacred Heart of Jesus (Jun) -> THIS FILE, Easter_offset 68 (below) - Immaculate Heart of Mary (Jun) -> THIS FILE, Easter_offset 69 (below) - Christ the King (Nov) -> Temporal_of.christ_the_king - Holy Family (Dec) -> Temporal_of.holy_family (its own heading carries a FIXED-DATE FALLBACK, "vel, ea deficiente, die 30 decembris" -- correctly Temporal_of's job, not this file's, despite the prose shape; left alone on the coordinator's own confirmation) No 8th heading exists -- the audit found nothing else of this shape. RESOLVED (fix round 1): the Sacred Heart of Jesus ("Feria VI post dominicam secundam post Pentecosten: SACRATISSIMI CORDIS IESU Sollemnitas") and the Immaculate Heart of Mary ("Sabbato post dominicam secundam post Pentecosten: Immaculati Cordis B. Mariae Virginis Memoria") are printed in the SAME base-2002 calendar table as Trinity/ Corpus Christi/Christ the King, but temporal_of.ml covered NEITHER -- found and reported by this extractor's own first pass; researched and confirmed by the coordinator (extracted lines ~4219-4222, immediately after 30 June). They belong HERE, not in temporal_of.ml: both are printed in the Calendarium Romanum Generale itself, the exact table this file transcribes, so adding them completes the transcription rather than amending it. Produced as ordinary Date_spec.Easter_offset entries (the same variant EF's Rogation Wednesday already uses, and Task 2 will reuse for Mary, Mother of the Church -- no kernel change). Pentecost is Easter+49 (Normae n. 22-23); the Second Sunday after Pentecost is Easter+63; the Friday after it is Easter+68 (Sacred Heart), the Saturday after it Easter+69 (Immaculate Heart) -- verified against three real Easter dates in test_calendar_of_data.ml, not merely computed on paper. See MOVABLE_ENTRY_HEADINGS/matching_movable_entry_heading below for how the extractor recognises these two headings specifically (and only these two) among the other 5 it still correctly skips. 2. Three FIXED entries that ARE in the printed table but are ALSO already computed by temporal_of.ml's own `named`: 1 January (Mary, Mother of God), 6 January (Epiphany), 25 December (the Nativity). Shipping them here too would duplicate a candidate Precedence_of.band would then have to arbitrate between a temporal-origin and a data-origin copy of the IDENTICAL office -- pointless and risks a silent divergence between the two copies. Excluded explicitly (SKIP_DATES below), not merely absent by accident. DATA DECISIONS made while transcribing, each cited in the emitted file's own provenance header: - A blank grade column means Memoria ad libitum, per the calendar's own footnote (`* Quando non indicatur gradus celebrationis, fit Memoria ad libitum.`) -- Normae/Missale Romanum 2002, Calendarium Romanum Generale, its own footnote marker. - 2 November (Commemoratio omnium fidelium defunctorum, All Souls) carries NO grade word in the table at all -- applying the footnote naively would misclassify the second-most solemn day of November as an optional memorial. Normae n. 59's own Tabula dierum liturgicorum, entry 3, places it explicitly alongside "Sollemnitates ... in Calendario generali inscriptae" ("Sollemnitates Domini, beatae Mariae Virginis, et Sanctorum in Calendario generali inscriptae. Commemoratio omnium fidelium defunctorum.") -- same table entry, same precedence tier. Tagged Sollemnitas here on that citation, as an explicit, documented override, not the generic blank-grade default. - Colour is derived from IGMR n. 346 (a)-(d): white for Christ's non-Passion celebrations / the BVM / Angels / non-martyr Saints / All Saints / John the Baptist's Nativity / John the Evangelist / Chair of Peter / Conversion of Paul (346(a), the last three named explicitly); red for Passion-related celebrations of the Lord, "festis nataliciis Apostolorum et Evangelistarum" (an Apostle's or Evangelist's own feast), and any celebration of Martyr Saints (346(b)); violet, chosen over the also-permitted black, for the Commemoration of All the Souls (346(d): "Assumi potest etiam in Officiis et Missis defunctorum" -- explicitly authorised for the dead; 346(a)'s white clause has no textual reach here, since All Souls is not "Sanctorum" (canonised Saints) at all). This DIVERGES from lectio, which colours 2 November white -- see the divergence log below; adjudicated colitur/violet on 346(d)'s explicit text, flagging lectio's white as reflecting a widespread modern PASTORAL custom the Missal's own words do not themselves authorise as the default. Three explicit exceptions to the apostle/evangelist-red rule (Chair of Peter, Conversion of Paul, John the Evangelist) and one to the martyr/apostle-red rule (three "In Dedicatione ..." church-dedication entries, always white) are coded as override tables, cited alongside the rule. EXTRACTION ARTIFACT: some titles are printed letter-spaced by pdftotext (e.g. "S . P e t r i D a m i a n i"). `collapse_letterspacing` mechanically rejoins runs of >=4 single-character tokens, using punctuation (periods, commas) and a lowercase-then-uppercase transition as word-boundary signals. This CANNOT recover a lowercase-to-lowercase boundary (no case signal survives the original single-spacing), so it is a first pass only; the sole entry it fires on in this table (21 February, S. Petri Damiani) is then hand-verified against the well-attested standard title and corrected explicitly, with the exact reason recorded at the call site. The count of entries needing this repair is printed to stderr. """ import hashlib import os import re import subprocess import sys import unicodedata from datetime import date, timezone, datetime PDF_PATH = "docs/research/of/missale-romanum-2002.pdf" TOOL_PATH = "tools/extract_of_calendar.py" LECTIO_INI = "../lectio/internal/caldata/roman-calendar.ini" MONTHS = [ "IANUARIUS", "FEBRUARIUS", "MARTIUS", "APRILIS", "MAIUS", "IUNIUS", "IULIUS", "AUGUSTUS", "SEPTEMBER", "OCTOBER", "NOVEMBER", "DECEMBER", ] MONTH_NUM = {m: i + 1 for i, m in enumerate(MONTHS)} ROW_RE = re.compile(r"^(Cal\.|Prid\.|Non\.|Idib\.|[IVXL]+)\s*(\d{1,2})\b\s*(.*)$") MOVABLE_HEADING_RE = re.compile(r"^(Dominica|Feria\s+[IVXL]+|Sabbato)\b.*\b(post|ultima|infra)\b") ENTRY_MARKER_RE = re.compile(r"^(S\.|Ss\.|B\.)\s") FOOTNOTE_RE = re.compile(r"^\*\s*Quando non indicatur") GRADE_WORDS = ["Sollemnitas", "Festum", "Memoria"] GRADE_TO_RANK = { "Sollemnitas": "Sollemnitas", "Festum": "Festum", "Memoria": "Memoria_obligatoria", None: "Memoria_ad_libitum", } # Excluded because Phase 1 (lib/rites/rite_of/temporal_of.ml) already # computes these as temporal-cycle offices. See module docstring, point 2. SKIP_DATES = { (1, 1): "of-mary-mother-of-god (Temporal_of.named, m=1 dd=1)", (1, 6): "of-epiphany (Temporal_of.named, m=1 dd=6)", (12, 25): "of-nativity (Temporal_of.named, m=12 dd=25)", } # IGMR 346(a)'s own three explicit white exceptions to the apostle/ # evangelist-red rule. WHITE_APOSTLE_EXCEPTIONS = {(1, 25), (2, 22), (12, 27)} # Church-dedication feasts: always white, never coloured by the martyr/ # apostle words that happen to appear in their own titles (the apostles # named are the church's patrons, not the day's own honouree). DEDICATION_DATES = {(8, 5), (11, 9), (11, 18)} # All Souls: see module docstring for the 346(d) citation. ALL_SOULS_OVERRIDE = (11, 2) # Explicit red override: the Exaltation of the Holy Cross is a celebration # "Passionis Domini" (346(b)) even though its own title contains neither # "martyr" nor "apostol"/"evangelist". RED_OVERRIDES = {(9, 14)} # subject = Lord: celebrations of the Lord not already excluded above, # plus the three church-dedication feasts. The dedication classification # is THIS FILE'S OWN READING, not lectio's: CLAUDE.md records a primary- # source argument for the EF side (the Common of the Dedication of a # Church is itself titled "Festum Dedicationis Ecclesiae est festum # Domini" in the Missal) that applies with equal force here -- a # dedication feast commemorates Christ's own house, not a saint. Fix # round 2 (coordinator, 2026-08-25): a prior version of this comment cited # lectio's own `class = lord` tag on these three dates as corroboration, # which is the wrong way round for a file whose SOLE authority is the # Missal -- lectio agreeing is worth recording as an observation, never as # the reason. (It happens to: roman-calendar.ini tags 08-05/11-09/11-18 # `class = lord`, checked, not merely asserted.) LORD_DATES = {(2, 2), (3, 25), (8, 6), (1, 3), (9, 14)} | DEDICATION_DATES def sha256_file(path): h = hashlib.sha256() with open(path, "rb") as f: h.update(f.read()) return h.hexdigest() def get_text(pdf_path): out = subprocess.run(["pdftotext", "-layout", pdf_path, "-"], capture_output=True, check=True) return out.stdout.decode("utf-8") def collapse_letterspacing(line): """Mechanically rejoins a run of >=4 single-character tokens. Returns (new_line, changed). See module docstring's "EXTRACTION ARTIFACT" note for what this can and cannot recover.""" tokens = line.split(" ") out = [] i = 0 changed = False while i < len(tokens): if len(tokens[i]) == 1: j = i while j < len(tokens) and len(tokens[j]) == 1: j += 1 run = tokens[i:j] if len(run) >= 4: changed = True merged = "".join(run) merged = re.sub(r"(?<=[.,])(?=\S)", " ", merged) merged = re.sub(r"(?<=[a-zà-öø-ÿæœ])(?=[A-ZÀ-ÖØ-ÝÆŒ])", " ", merged) out.append(merged) else: out.extend(run) i = j else: out.append(tokens[i]) i += 1 return " ".join(out), changed def bound_table(text): lines = text.split("\n") # First IANUARIUS before line 5000, immediately preceded (allowing for a # blank line) by "CALENDARIUM ROMANUM GENERALE" -- rules out the second, # Proprium Sanctorum run near line 17874. start = None for i, l in enumerate(lines[:5000]): if l.strip("\x0c").strip() == "IANUARIUS": # confirm CALENDARIUM ROMANUM GENERALE appears in the preceding # few lines window = "\n".join(lines[max(0, i - 4):i]) if "CALENDARIUM ROMANUM GENERALE" in window: start = i break if start is None: raise SystemExit("could not find the start of the calendar table (IANUARIUS after " "CALENDARIUM ROMANUM GENERALE, before line 5000)") end = None for i in range(start, min(len(lines), 5000)): if lines[i].strip() == "TABELLA TEMPORARIA": end = i break if end is None: raise SystemExit("could not find the end of the calendar table (TABELLA TEMPORARIA)") return start, end, lines[start:end] def is_month_header(line): return line.strip("\x0c").strip() in MONTH_NUM def extract_grade(text): """Splits a trailing grade word off `text`. Shared by fixed-date rows and movable-heading entries (parse_rows below) so both go through the identical rule: a printed grade word (Sollemnitas/Festum/Memoria) is stripped; its absence, or Jan 3's own trailing footnote asterisk, both fall through to the calendar's blank-grade footnote (Memoria ad libitum).""" grade = None for g in GRADE_WORDS: if text == g or text.endswith(" " + g): grade = g text = text[: -len(g)].strip() break if text.endswith("*"): text = text[:-1].strip() return text, grade # Fix round 1 (coordinator, 2026-08-25): two movable universal celebrations # are printed as prose HEADINGS between fixed-date rows, not as (kalends, # day, title, grade) rows -- exactly the shape parse_rows' generic # MOVABLE_HEADING_RE skip-and-discard branch was built for, which is why the # original extraction silently dropped them (found and reported by the # extractor's own author; researched and confirmed by the coordinator). # Both belong HERE, not in temporal_of.ml: they are printed in the # Calendarium Romanum Generale itself, the exact table this file # transcribes, so adding them completes the transcription rather than # amending it (contrast Holy Family, whose heading is likewise prose but # whose FIXED-DATE FALLBACK, "vel, ea deficiente, die 30 decembris", makes # it correctly Temporal_of's job -- see the re-audit note in the module # docstring's "MOVABLE-HEADING AUDIT" section for the full accounting of # all 7 such headings in the table). # # Feria VI post dominicam secundam post Pentecosten: # SACRATISSIMI CORDIS IESU Sollemnitas # Sabbato post dominicam secundam post Pentecosten: # Immaculati Cordis B. Mariae Virginis Memoria # # (extracted lines ~4219-4222, immediately after 30 June). Pentecost is # Easter+49 (Normae n. 22-23, already temporal_of.ml's own citation); "the # Second Sunday after Pentecost" is therefore Easter+63, the Friday after it # Easter+68, the Saturday after it Easter+69 -- independently verified # against three real Easter dates (2026-04-05, 2027-03-28, 2035-03-25) in # test_calendar_of_data.ml, not merely computed on paper here. # # Date_spec.Easter_offset (built for EF's Rogation Wednesday) needs no # kernel change and is the same mechanism Task 2 will use for Mary, Mother # of the Church -- so these become ordinary Easter_offset entries in this # file, produced by the extractor like every other entry, not hand-appended: # a heading this specific is no less mechanical to recognise than a # (kalends, day) row, and hand-appending would be the one entry pair in this # file with no SHA-256-verifiable path back to the PDF text. MOVABLE_ENTRY_HEADINGS = [ # (heading-substring, Easter offset, weekday name for the test/sanity check) ("Feria VI post dominicam secundam post Pentecosten", 68, "Fri"), ("Sabbato post dominicam secundam post Pentecosten", 69, "Sat"), ] def matching_movable_entry_heading(line): for substring, offset, weekday in MOVABLE_ENTRY_HEADINGS: if line.startswith(substring): return offset, weekday return None def parse_rows(table_lines): """Returns (entries, movable_entries, letterspace_count, movable_audit) where entries is a list of dicts with month, day, latin, grade; movable_entries is the same shape but with easter_offset instead of month/day; movable_audit lists EVERY "Dominica/Feria/Sabbato ... :" heading found, with its own fate (produced here, or the reason it is skipped), for the re-audit this fix round asked for.""" # Pass 1: strip form feeds and collapse letter-spacing line by line. letterspace_count = 0 lines = [] for raw in table_lines: l = raw.replace("\x0c", "") l2, changed = collapse_letterspacing(l) if changed: letterspace_count += 1 lines.append(l2) entries = [] movable_entries = [] movable_audit = [] current_month = None i = 0 n = len(lines) while i < n: line = lines[i].strip() if line == "": i += 1 continue if is_month_header(line): current_month = MONTH_NUM[line] i += 1 continue if FOOTNOTE_RE.match(line): i += 1 continue if MOVABLE_HEADING_RE.match(line): heading = line i += 1 block = [] while i < n: nxt = lines[i].strip() if (nxt == "" or is_month_header(nxt) or ROW_RE.match(nxt) or MOVABLE_HEADING_RE.match(nxt) or FOOTNOTE_RE.match(nxt)): break block.append(nxt) i += 1 match = matching_movable_entry_heading(heading) if match is None: movable_audit.append({"heading": heading, "content": block, "produced": False}) continue offset, weekday = match text = " ".join(block).strip() text, grade = extract_grade(text) movable_entries.append({"easter_offset": offset, "weekday": weekday, "latin": text, "grade": grade}) movable_audit.append({"heading": heading, "content": block, "produced": True, "easter_offset": offset}) continue rm = ROW_RE.match(line) if rm: if current_month is None: raise SystemExit(f"row before any month header: {line!r}") _kalend, day_s, rest = rm.groups() day = int(day_s) blocks = [[]] if rest.strip(): blocks[0].append(rest.strip()) i += 1 while i < n: nxt_raw = lines[i] nxt = nxt_raw.strip() if (nxt == "" or is_month_header(nxt) or ROW_RE.match(nxt) or MOVABLE_HEADING_RE.match(nxt) or FOOTNOTE_RE.match(nxt)): break if ENTRY_MARKER_RE.match(nxt) and blocks[-1]: blocks.append([nxt]) else: blocks[-1].append(nxt) i += 1 for cblock in blocks: if not cblock: continue text = " ".join(cblock).strip() text, grade = extract_grade(text) if text: entries.append({"month": current_month, "day": day, "latin": text, "grade": grade}) continue raise SystemExit(f"unrecognised line in calendar table (month={current_month}): {line!r}") return entries, movable_entries, letterspace_count, movable_audit LATIN_TO_ASCII = str.maketrans({ # Æ (uppercase ligature) is verified, across every occurrence in the # extracted calendar table, to appear ONLY inside an ALL-CAPS heading # ("SANCTÆ", "PRÆSENTATIONE", "ASSUMPTIONE BEATÆ") -- never as the # initial letter of a Title Case word -- so it maps to "AE" (both # capitals), not "Ae", to stay consistent with its own surrounding case. "æ": "ae", "Æ": "AE", "œ": "oe", "Œ": "OE", "ø": "o", "Ø": "O", }) def normalize_latin(text): """æ -> ae, œ -> oe, matching the existing la-name convention already used by lib/rites/rite_of/temporal_of.ml's own holy_family_names/ baptism_names (e.g. "Sanctae Familiae Iesu, Mariae et Ioseph"), not the ligature glyphs the PDF itself prints. Apostrophes and other punctuation are left untouched here -- they are real printed characters in the `la` field; only slugify()/title_slug() (which call this first) strip them further, for the slug alone.""" return text.translate(LATIN_TO_ASCII) def slugify(latin): s = normalize_latin(latin) s = unicodedata.normalize("NFKD", s) s = "".join(c for c in s if not unicodedata.combining(c)) s = s.lower() s = re.sub(r"[^a-z0-9]+", "-", s) s = re.sub(r"-+", "-", s).strip("-") return s # Words to drop when deriving a slug from a title, so slugs read as names # rather than descriptions (kept short and stable, matching data/ef/ # sanctoral.sexp's own style, e.g. "agnes", "all-saints"). DROP_WORDS = { "s", "ss", "b", "in", "et", "de", "ad", "a", "the", "of", } def title_slug(latin, month, day): words = re.split(r"[\s,]+", normalize_latin(latin)) kept = [] for w in words: base = re.sub(r"\.$", "", w) if base.lower() in DROP_WORDS or base == "": continue kept.append(base) if len(kept) >= 6: break if not kept: kept = [f"of-{month:02d}-{day:02d}"] return slugify(" ".join(kept)) def classify_colour(month, day, latin): key = (month, day) if key == ALL_SOULS_OVERRIDE: return "Violet" if key in RED_OVERRIDES: return "Red" if key in DEDICATION_DATES: return "White" if "martyr" in latin.lower(): return "Red" if key in WHITE_APOSTLE_EXCEPTIONS: return "White" lower = latin.lower() if "apostol" in lower or "evangelist" in lower: return "Red" return "White" # Fix round 2 (coordinator, 2026-08-25): the previous rule ("maria"/"b.m.v" # substring anywhere in the Latin title) over-matched -- 12 of the 25 # resulting Bvm entries were saints who happen to carry "Maria" as part of # their OWN given name (Maximilian MARY Kolbe, John MARY Vianney, Alphonsus # MARIA de Liguori, MARY Magdalene, MARIA Goretti, Margaret MARY Alacoque, # Anthony MARY Claret, Anthony MARIA Zaccaria, Louis Grignion de Montfort's # own "MARIAE", Mary Magdalene de Pazzi) or belong to a religious order whose # NAME contains "B.M.V." (the Seven Holy Founders of the Servite Order), or # are Mary's own parents commemorated as themselves (Joachim and Anne) -- # none of these celebrations is OF Mary; the subject of a celebration is # whom it commemorates, not who is merely named in its title. Replaced with # an explicit table of the 12 genuinely Marian FIXED-date entries (read off # the Missal page by hand, each a feast/memorial of Mary herself), keyed by # (month, day) rather than slug -- consistent with every other override # table in this file and independent of lectio's own slug text, which this # file does not treat as authoritative for anything but the identifier # itself (see the module docstring's "What actually comes from lectio"). # The Immaculate Heart of Mary (the one movable Marian entry) is tagged # separately via MOVABLE_SUBJECT_OVERRIDE, not this table. BVM_DATES = { (2, 11): "our-lady-of-lourdes", (5, 13): "our-lady-of-fatima", (5, 31): "visitation-of-the-blessed-virgin-mary", (7, 16): "our-lady-of-mount-carmel", (8, 15): "assumption-of-the-blessed-virgin-mary", (8, 22): "queenship-of-blessed-virgin-mary", (9, 8): "birth-of-the-blessed-virgin-mary", (9, 12): "holy-name-of-the-blessed-virgin-mary", (9, 15): "our-lady-of-sorrows", (10, 7): "our-lady-of-the-rosary", (11, 21): "presentation-of-the-blessed-virgin-mary", (12, 8): "immaculate-conception-of-the-blessed-virgin-mary", } def classify_subject(month, day, latin): key = (month, day) if key in LORD_DATES: return "Lord" if key in BVM_DATES: return "Bvm" return "Saint" def parse_lectio_ini(path): # Fix round 2: this used to catch FileNotFoundError and return [], # letting main() run to completion with zero lectio-derived data (no # en names, mechanical Latin slugs) and exit 0 -- SILENTLY, with a # DIFFERENT SHA-256 than the shipped file, and nothing printed to # explain why. lectio is a real, load-bearing dependency for slugs and # en names (see the module docstring's "WHAT ACTUALLY COMES FROM # LECTIO"), so its absence is now fatal, not degraded-and-silent. if not os.path.exists(path): raise SystemExit( f"extract_of_calendar: lectio's cross-check file is required and was not found: " f"{path!r}. Running without it would silently change the output (no `en` names, " f"mechanical Latin slugs instead of lectio's, a different SHA-256) with no warning " f"-- exactly the failure mode fix round 2 closed. Clone/update " f"~/git/projects/lectio and re-run." ) with open(path, encoding="utf-8") as f: raw = f.read() entries = [] cur = None for line in raw.split("\n"): line = line.rstrip("\n") if line.startswith("[") and line.endswith("]"): if cur is not None: entries.append(cur) cur = {"slug": line[1:-1]} elif "=" in line and cur is not None: k, _, v = line.partition("=") cur[k.strip()] = v.strip() if cur is not None: entries.append(cur) return [e for e in entries if "date" in e] LECTIO_RANK_TO_VOCAB = { "solemnity": "Sollemnitas", "feast": "Festum", "memorial": "Memoria_obligatoria", "optional": "Memoria_ad_libitum", } def significant_words(text): words = re.split(r"[\s,.;]+", normalize_latin(text).lower()) return {w for w in words if w and w not in DROP_WORDS and len(w) > 2} def cross_check(entries, lectio_entries): """Returns (report_lines, name_en_by_id) where name_en_by_id maps id(entry dict) -> matched lectio name.en, and report_lines is the divergence log for the provenance header.""" by_date = {} for le in lectio_entries: d = le.get("date", "") m = re.match(r"^(\d\d)-(\d\d)$", d) if not m: continue key = (int(m.group(1)), int(m.group(2))) by_date.setdefault(key, []).append(le) by_date_missal = {} for e in entries: by_date_missal.setdefault((e["month"], e["day"]), []).append(e) report = [] name_en = {} lectio_slug = {} agree_rank = 0 total_matched = 0 matched_lectio_ids = set() for key, missal_list in sorted(by_date_missal.items()): lectio_list = by_date.get(key, []) if len(missal_list) == 1 and len(lectio_list) == 1: pairs = [(missal_list[0], lectio_list[0])] else: pairs = [] remaining = list(lectio_list) for me in missal_list: if not remaining: break mwords = significant_words(me["latin"]) best = max(remaining, key=lambda le: len(mwords & significant_words(le.get("name.la", "")))) pairs.append((me, best)) remaining.remove(best) for me, le in pairs: total_matched += 1 matched_lectio_ids.add(id(le)) name_en[id(me)] = le.get("name.en") lectio_slug[id(me)] = le.get("slug") missal_rank = GRADE_TO_RANK[me["grade"]] lectio_rank = LECTIO_RANK_TO_VOCAB.get(le.get("rank", ""), "?") if missal_rank == lectio_rank: agree_rank += 1 else: report.append( f" {key[0]:02d}-{key[1]:02d} RANK colitur={missal_rank} lectio={lectio_rank} " f"[{me['latin']!r} vs {le.get('name.la','')!r}]" ) missal_colour = classify_colour(key[0], key[1], me["latin"]) lectio_colour = (le.get("colour", "") or "").capitalize() if missal_colour != lectio_colour: report.append( f" {key[0]:02d}-{key[1]:02d} COLOUR colitur={missal_colour} lectio={lectio_colour} " f"[{me['latin']!r}]" ) if not lectio_list: for me in missal_list: report.append(f" {key[0]:02d}-{key[1]:02d} MISSAL-ONLY {me['latin']!r} (grade={me['grade']})") elif len(lectio_list) > len(missal_list): # A shared date where lectio carries MORE entries than the # Missal transcription found -- the pairing loop above only # consumes len(missal_list) of them, so the rest would # otherwise vanish silently. Surfaced explicitly. for le in lectio_list: if id(le) not in matched_lectio_ids: report.append( f" {key[0]:02d}-{key[1]:02d} LECTIO-EXTRA (same date, uncovered by Missal pairing) " f"{le.get('name.la','')!r} rank={le.get('rank')} slug={le.get('slug')}" ) missal_dates = set(by_date_missal) for key, lectio_list in sorted(by_date.items()): if key in SKIP_DATES: continue for le in lectio_list: if id(le) not in matched_lectio_ids and key not in missal_dates: report.append(f" {key[0]:02d}-{key[1]:02d} LECTIO-ONLY {le.get('name.la','')!r} " f"rank={le.get('rank')} slug={le.get('slug')}") summary = [ f"lectio cross-check: {total_matched} dates paired; rank agrees on {agree_rank}/{total_matched}.", ] return summary + report, name_en, lectio_slug def render_sexp(entries, name_en, letterspace_count, skipped, meta): lines = [] lines.append(f"; data/of/calendar-2002.sexp -- OF (post-1970) General Roman Calendar,") lines.append(f"; transcribed from the Missale Romanum, editio typica tertia (2002). The") lines.append(f"; 2002 Missal is the SOLE authority for every substantive field (date, rank,") lines.append(f"; colour, subject, every character of `la`); this file is never edited") lines.append(f"; afterwards -- spec sec4.1. lectio's roman-calendar.ini supplies `en` names") lines.append(f"; and, when a confident match exists, the slug (an engineering key, not") lines.append(f"; content -- see tools/extract_of_calendar.py's own docstring, \"WHAT") lines.append(f"; ACTUALLY COMES FROM LECTIO\", for exactly what and why) -- corrected fix") lines.append(f"; round 2, 2026-08-25: a prior header here said lectio was \"never a") lines.append(f"; source\", which omitted slugs and was not quite true.") lines.append(f";") lines.append(f"; Generator: {TOOL_PATH} -- do not hand-edit; re-run against the same") lines.append(f"; PDF AND the same lectio snapshot (its SHA-256 is pinned below; a missing") lines.append(f"; or different lectio file now fails loudly rather than silently changing") lines.append(f"; slugs/en-names/SHA-256 -- fix round 2) to reproduce this file byte-for-") lines.append(f"; byte, and diff before committing.") lines.append(f";") lines.append(f"; Source (authority): {PDF_PATH}") lines.append(f"; SHA-256: {meta['pdf_sha256']}") lines.append(f"; Extracted lines (pdftotext -layout, 0-based): {meta['start']}-{meta['end']}") lines.append(f"; Source (slugs + en names only, never content -- see above): {LECTIO_INI}") lines.append(f"; SHA-256: {meta['lectio_sha256']}") lines.append(f"; Extraction date (UTC): {meta['extraction_date']}") lines.append(f";") lines.append(f"; {len(entries)} entries. {letterspace_count} entry needed letter-spacing repair") lines.append(f"; (collapse_letterspacing's own mechanical pass, then hand-verified -- see") lines.append(f"; the module docstring's EXTRACTION ARTIFACT note); {meta['diacritic_repair_count']}") lines.append(f"; more carry a SEPARATE pdftotext artifact (a combining diacritic rendered") lines.append(f"; as a stray spacing character), also hand-repaired -- fix round 2, see") lines.append(f"; main()'s own diacritic_repairs table.") lines.append(f";") lines.append(f"; Excluded, deliberately (see module docstring for the full reasoning):") for (m, d), why in sorted(skipped.items()): lines.append(f"; {m:02d}-{d:02d}: {why}") lines.append(f"; Movable universal solemnities STILL excluded, correctly (Baptism of") lines.append(f"; the Lord, Holy Family, Trinity, Corpus Christi, Christ the King) --") lines.append(f"; Temporal_of code already covers each; see the module docstring's") lines.append(f"; MOVABLE-HEADING AUDIT for the full 7-heading accounting.") lines.append(f";") lines.append(f"; Movable entries produced from a prose heading, not a (kalends, day) row") lines.append(f"; (fix round 1, 2026-08-25 -- see MOVABLE_ENTRY_HEADINGS):") for me in meta.get("movable_entries", []): lines.append(f"; Easter_offset {me['easter_offset']} ({me['weekday']}): {me['latin']!r}") lines.append(f";") for l in meta["cross_check_report"]: lines.append(f"; {l}") lines.append("") def esc(s): return s.replace("\\", "\\\\").replace('"', '\\"') lines.append("((id of-universal) (name \"OF (2002) General Roman Calendar\")") lines.append(" (entries") lines.append(" (") for idx, e in enumerate(entries): la = normalize_latin(e["latin"]) en = name_en.get(id(e)) names_parts = [] if en: names_parts.append(f'(en "{esc(en)}")') names_parts.append(f'(la "{esc(la)}")') names_str = " ".join(names_parts) rank = GRADE_TO_RANK[e["grade"]] if "colour" in e: colour = e["colour"] else: colour = classify_colour(e["month"], e["day"], e["latin"]) if "subject" in e: subject = e["subject"] else: subject = classify_subject(e["month"], e["day"], e["latin"]) slug = e["slug"] prefix = " " if idx > 0 else " " if "easter_offset" in e: date_sexp = f"(Easter_offset {e['easter_offset']})" else: date_sexp = f"(Fixed (month {e['month']}) (day {e['day']}))" lines.append(f"{prefix}((date {date_sexp})") lines.append(f" (cel") lines.append(f" ((slug {slug})") lines.append(f" (names ({names_str}))") lines.append(f" (rank {rank}) (status Feast) (colour {colour})") lines.append(f" (subject {subject}) (citations ()) (layer of-universal))))") lines.append(" )))") return "\n".join(lines) + "\n" # Fix round 1: hand-assigned, clean English-style slugs for the two # movable entries, keyed by Easter offset. Not run through title_slug's # Latin-mechanical fallback (which would produce "sacratissimi-cordis-iesu") # because lectio -- the only source this file ever takes a slug FROM when # one exists (see the "Slugs:" comment in main() below) -- carries neither # entry at all (its own header states both are "computed by internal/ # calendar", so its roman-calendar.ini has no bracket slug to reuse). This # is the one place in the file a slug is authored rather than derived -- # recorded here rather than left implicit. MOVABLE_SLUG_OVERRIDE = {68: "sacred-heart-of-jesus", 69: "immaculate-heart-of-mary"} # Fix round 1: subject/colour for the two movable entries, set directly # rather than through classify_colour/classify_subject (which key off # (month, day) -- meaningless for an Easter_offset entry). Both white # (IGMR 346(a): "celebrationibus Domini quae non sint de eius Passione" for # the Sacred Heart; the same clause's "beatae Mariae Virginis" for the # Immaculate Heart -- neither is a Passion or martyr celebration). Subject # Lord/Bvm respectively, read directly off each title's own referent. MOVABLE_SUBJECT_OVERRIDE = {68: "Lord", 69: "Bvm"} def main(): if len(sys.argv) > 1: with open(sys.argv[1], encoding="utf-8") as f: text = f.read() else: text = get_text(PDF_PATH) start, end, table_lines = bound_table(text) entries, movable_entries, letterspace_count, movable_audit = parse_rows(table_lines) # Hand-verified correction for the one letter-spaced entry the # mechanical pass cannot fully resolve (see module docstring). for e in entries: if (e["month"], e["day"]) == (2, 21) and "episcopiet" in e["latin"]: e["latin"] = "S. Petri Damiani, episcopi et Ecclesiae doctoris" # Fix round 2 (coordinator, 2026-08-25): a SECOND, different pdftotext # artifact -- distinct from the letter-spacing above -- affects four # more entries. Each is a proper noun carrying a COMBINING diacritic # (macron/breve/tilde/ogonek/dot-below) the font's cmap maps to a # SEPARATE spacing character positioned after its base letter with a # stray space, instead of combining onto it (e.g. "Makhlu‾ f" where # ‾ is U+203E OVERLINE standing in for a macron over "u"). Not # mechanically recoverable the way collapse_letterspacing's own runs # are (there is no generic rule for "which glyph substitutes for which # combining mark"), so each is hand-repaired here, exactly like Peter # Damian above -- and each is corroborated by the IDENTICAL corruption # recurring a second time elsewhere in the same extracted text (the # Missal's own alphabetical index of saints, lines ~33857/34079: the # stray mark pairs with the same base letter both times, not assumed # from outside knowledge alone), then checked against the standard # Latin liturgical/romanised spelling for each canonised saint. diacritic_repairs = [ (7, 24, "Makhlu‾ f", "Makhlūf"), # OVERLINE -> u-macron: Sarbelii Makhlūf (9, 20, "Tae-go¡ n", "Tae-gŏn"), # INV. EXCL. MARK -> o-breve: Kim Taegŏn (9, 20, "Cho¡ ng", "Chŏng"), # same artifact, same entry: Chŏng Ha-sang (11, 24, "Du˜ ng La.c", "Dũng Lạc"), # TILDE -> u-tilde; "."-> a-dot-below: Dũng Lạc (12, 23, "Ke¸ty", "Kęty"), # CEDILLA -> e-ogonek: Ioannis de Kęty ] diacritic_repair_count = 0 for e in entries: for m, d, bad, good in diacritic_repairs: if (e["month"], e["day"]) == (m, d) and bad in e["latin"]: e["latin"] = e["latin"].replace(bad, good) diacritic_repair_count += 1 skipped = {} kept = [] for e in entries: key = (e["month"], e["day"]) if key in SKIP_DATES: skipped[key] = SKIP_DATES[key] continue kept.append(e) entries = kept # All Souls override (see module docstring). for e in entries: if (e["month"], e["day"]) == ALL_SOULS_OVERRIDE: e["grade"] = "Sollemnitas" for e in movable_entries: e["slug"] = MOVABLE_SLUG_OVERRIDE[e["easter_offset"]] e["colour"] = "White" e["subject"] = MOVABLE_SUBJECT_OVERRIDE[e["easter_offset"]] lectio_entries = parse_lectio_ini(LECTIO_INI) lectio_sha256 = sha256_file(LECTIO_INI) cross_check_report, name_en, lectio_slug = cross_check(entries, lectio_entries) # Slugs: prefer lectio's OWN identifier when a confident match exists -- # it is already a clean, conventional, human-authored identifier # ("raymond-of-penyafort-priest") matching data/ef/sanctoral.sexp's own # style, and reusing it is not a source-discipline violation: a slug is # a stable ENGINEERING KEY, not liturgical content -- every substantive # field (date, rank, colour, the `la` name) still comes from the Missal # alone. Falls back to a mechanical slug built from the Latin title for # any entry lectio has no match for (none, as it happens, in this run -- # every one of the 206 entries paired -- but kept for robustness against # a future lectio update dropping an entry). Collisions suffixed either # way. seen = {} for e in entries: candidate = lectio_slug.get(id(e)) base = slugify(candidate) if candidate else title_slug(e["latin"], e["month"], e["day"]) if not base: base = title_slug(e["latin"], e["month"], e["day"]) if base in seen: seen[base] += 1 e["slug"] = f"{base}-{seen[base]}" else: seen[base] = 1 e["slug"] = base # Movable entries carry their own hand-assigned slugs already (see # MOVABLE_SLUG_OVERRIDE); still registered in `seen` and defensively # collision-checked like everything else, rather than assumed safe. for e in movable_entries: base = e["slug"] if base in seen: raise SystemExit(f"movable entry slug {base!r} collides with a fixed-date entry") seen[base] = 1 entries = entries + movable_entries entries.sort(key=lambda e: e["slug"]) meta = { "pdf_sha256": sha256_file(PDF_PATH), "lectio_sha256": lectio_sha256, "start": start, "end": end, "extraction_date": datetime.now(timezone.utc).strftime("%Y-%m-%d"), "cross_check_report": cross_check_report, "movable_entries": movable_entries, "diacritic_repair_count": diacritic_repair_count, } sys.stdout.write(render_sexp(entries, name_en, letterspace_count, skipped, meta)) print(f"[extract_of_calendar] {len(entries)} entries emitted ({len(movable_entries)} of them " f"movable, Easter_offset), {letterspace_count} needed letter-spacing repair, " f"{diacritic_repair_count} needed diacritic repair, " f"{len(skipped)} dates excluded (temporal_of.ml coverage)", file=sys.stderr) print(f"[extract_of_calendar] lectio SHA-256: {lectio_sha256}", file=sys.stderr) print(f"[extract_of_calendar] MOVABLE-HEADING AUDIT: {len(movable_audit)} \"Dominica/Feria/" f"Sabbato ... :\" headings found in the table:", file=sys.stderr) for ma in movable_audit: fate = (f"PRODUCED as Easter_offset {ma['easter_offset']}" if ma["produced"] else "skipped -- covered by Temporal_of code") print(f"[extract_of_calendar] {ma['heading']!r} -> {fate}", file=sys.stderr) for l in cross_check_report: print(f"[extract_of_calendar] {l}", file=sys.stderr) if __name__ == "__main__": main()