aboutsummaryrefslogtreecommitdiff
path: root/tools/extract_fiuv_ordo.ml
diff options
context:
space:
mode:
Diffstat (limited to 'tools/extract_fiuv_ordo.ml')
-rw-r--r--tools/extract_fiuv_ordo.ml855
1 files changed, 855 insertions, 0 deletions
diff --git a/tools/extract_fiuv_ordo.ml b/tools/extract_fiuv_ordo.ml
new file mode 100644
index 0000000..26f5b16
--- /dev/null
+++ b/tools/extract_fiuv_ordo.ml
@@ -0,0 +1,855 @@
+(* Witnesses task (2026-08-22-colitur-celebrant-rubrics-phase1): turns
+ pdftotext's -layout dump of the FIUV (Foederatio Internationalis Una
+ Voce) universal Ordo into test/fixtures/fiuv-ordo-2025-2026.sexp, one
+ row per day, 2025-11-27..2026-12-31.
+
+ A DIFFERENT format from tools/extract_lms_ordo.ml's own LMS PDFs in
+ every way that matters (see docs/research/ordo/PROVENANCE-ordo-corpus.md):
+ UNIVERSAL (no diocesan variants to exclude at all -- one block per day,
+ not "the first of several"), IN LATIN, using the rubrics' own
+ vocabulary ("Missa pr., Gloria, sine Credo, praef. comm."), and BOTH
+ directions of Gloria/Credo are stated explicitly in the source's own
+ words ("Gloria"/"sine Gloria", "Credo"/"sine Credo") rather than by a
+ printed flag letter. Hand-rolled, no Str/regex (frozen deps) -- the
+ same String.sub/index/split_on_char discipline extract_lms_ordo.ml
+ already uses for a different publisher's layout.
+
+ THE KEY STRUCTURAL FACTS this parser leans on, established during
+ characterisation (see the fixture's own provenance header for the
+ full record):
+
+ 1. A day-start line's own FIRST token is always "<1-2 digit day>."
+ (e.g. "27.", "3."), immediately followed by a colour code (not
+ anchored on below -- colour is captured as raw prose, never parsed
+ into a closed set, because a few real days use a COMPOUND colour
+ ("Viol. in Off., Alb. in Missa." on Holy Thursday) or drop the
+ token to a blank cell entirely (13 December 2026, a rendering
+ quirk) -- neither is needed for anything this task compares).
+
+ 2. Month headers/running-footers name a month EITHER in English with a
+ trailing ROMAN-NUMERAL year ("November MMXXV", the book's own first
+ two months only) OR bare in Latin with no year at all
+ ("Januarius", "Februarius", ... every month after). The English
+ Nov/Dec spellings happen to be IDENTICAL to their Latin
+ equivalents, so one 12-entry table covers both vocabularies.
+ CHARACTERISATION FINDING: the SECOND header ("December MMXV") is a
+ publisher-side TYPO -- MMXV is 2015, ten years off, contradicted by
+ every neighbouring date and by month-order inference alone. This
+ parser therefore reads a Roman-numeral year ONLY off the very
+ FIRST header encountered (the anchor) and tracks every subsequent
+ month purely by wraparound inference (increment the year exactly
+ once, at the Dec->Jan transition) -- the typo is never read at all,
+ not merely tolerated.
+
+ 3. Every real Mass rubric line starts with the literal token "Missa"
+ (401 of ~400 real day-blocks, confirmed structurally). A SECOND,
+ alternate Mass option, when the day offers one, is introduced by
+ the CAPITALISED two-token sequence "Vel Missa" -- never bare lower-
+ case "vel" alone, which is the ordinary Latin conjunction "or" and
+ appears constantly inside ordinary prose (e.g. "praef. comm. vel de
+ Martyribus", part of ONE preface's own text, not a second Mass
+ option) -- so only "Vel Missa", both tokens, anchors a real
+ alternate-option boundary. This parser reads Gloria/Credo/praef.
+ from the FIRST Mass option only, bounded above by whichever comes
+ first: "Vel Missa" or "VESPERAE"/"VESPERA" -- the identical
+ "day's own PRIMARY office, not a menu entry" discipline
+ extract_lms_ordo.ml already uses for the LMS diocesan-variant
+ colon boundary.
+
+ 4. SUBSTRING TRAP, found and defended against: "Gloria" is also the
+ first word of "Gloria Patri" (the psalm doxology, "Glory be to the
+ Father"), a DIFFERENT liturgical unit that has nothing to do with
+ whether the Mass's own Gloria in excelsis is said -- and it can
+ appear INSIDE the very Mass clause being scanned (Good Friday:
+ "Missa pr., (omittuntur ps. Iudica me et Gloria Patri), Gloria,
+ sine Credo, praef. comm." -- the real, standalone "Gloria," follows
+ immediately after). [find_word] below rejects any "Gloria" hit
+ whose very next token is "Patri", the same whole-token discipline
+ extract_lms_ordo.ml's own header describes for "V Mass of BVM"
+ being a substring of "IV Mass of BVM".
+
+ 5. A WHOLE-SECOND-OFFICE TRAP, found and defended against
+ (celebrant-rubrics-phase1 Phase 2 correction, 2026-08-22, closing
+ data/ef/expected-divergences-fiuv.sexp's own F3): six real days
+ print BOTH a strict-1962 reading (a privileged Lenten/Passiontide
+ feria, saint reduced to a commemoration, "Off. feriale") AND, after
+ a literal separator token, a SECOND, complete alternative office
+ (the same saint kept outright under the 2020 decree Cum
+ Sanctissima, "Off. ordinarium", with its OWN "Ad Mat. ... Te Deum"
+ lessons). The separator is an en-dash, "VEL" (Latin "or"), another
+ en-dash, glued by pdftotext into ONE token with no internal spaces
+ -- "\xe2\x80\x93VEL\xe2\x80\x93" ("VEL" wrapped in en-dashes) --
+ confirmed by direct byte inspection of the raw dump, NOT the same
+ shape as [extract_missa_fields]'s own "Vel"+"Missa" two-token
+ boundary above (a narrower "choice of Mass formulary within the
+ SAME office" marker, e.g. a Common vs a Proper Mass for one saint).
+ [extract_missa_fields]'s own Gloria/Credo/praef span happened to
+ stay correctly bounded to the PRIMARY office on all six real
+ instances even before this fix, because a Vespers line ("VESPER..."
+ -- already one of its own boundary tokens) always intervenes before
+ the separator in this corpus; [extract_te_deum] had NO such
+ protection at all -- it scans the WHOLE block's own text for "Ad
+ Mat." unbounded, so on these six days (whose PRIMARY office states
+ no "Ad Mat." of its own at all -- an ordinary privileged feria's
+ Matins needs no special note) it read the SECOND office's own "Ad
+ Mat. ... Te Deum." instead, manufacturing a same-day contradiction
+ that F3 first recorded as "colitur's own possible gap". [split_at_alt_office]
+ below cuts every block's own token array at this separator BEFORE
+ any field extraction runs, so every extractor now reads the PRIMARY
+ office only; the alternative is captured, not discarded, in its own
+ [alt_te_deum] field (the row type's own comment has the full
+ account of why only this one field, not a parallel alt_gloria/
+ alt_credo/alt_class -- those three were never wrong, so widening
+ them now would be undirected scope creep, not a fix).
+
+ Usage:
+ pdftotext -layout docs/research/ordo/fiuv-ordo-2025-2026.pdf /tmp/fiuv.txt
+ dune exec tools/extract_fiuv_ordo.exe -- /tmp/fiuv.txt \
+ docs/research/ordo/fiuv-ordo-2025-2026.pdf \
+ test/fixtures/fiuv-ordo-2025-2026.sexp *)
+
+open Sexplib0.Sexp_conv
+module Date = Colitur_kernel.Date
+
+let die fmt = Printf.ksprintf (fun s -> prerr_endline ("extract_fiuv_ordo: " ^ s); exit 1) fmt
+
+(* Mirrored, not shared, by test/test_fiuv_ordo.ml -- tools/ and test/
+ have no common .mli either could hang a shared type from, the same
+ reasoning extract_lms_ordo.ml's own header already gives. *)
+type row = {
+ date : string; (** ISO-8601 *)
+ class_ : string option; (** raw, e.g. "III cl.", "III cl. (Priv.)", "I cl." -- [None] only if genuinely unparseable, read off the PRIMARY office alone *)
+ title : string; (** the day-start line's own text before the class marker, PRIMARY office alone *)
+ te_deum : bool option; (** [None] only if neither "Te Deum" nor "non dicitur Te Deum"/"sine Te Deum" is found WITHIN THE PRIMARY OFFICE (see [split_at_alt_office]) *)
+ gloria : bool option; (** PRIMARY office *)
+ credo : bool option; (** PRIMARY office *)
+ praef : string option; (** raw trailing text after "praef." within the primary Mass clause -- CAPTURED, NOT VALIDATED *)
+ alt_te_deum : bool option;
+ (** the SECOND, Cum-Sanctissima-alternative office's own Te Deum
+ status, when this day prints one (see [split_at_alt_office]'s
+ own citation) -- [None] both when the day has no such
+ alternative at all (394 of 400 real days) AND when it has one
+ but that alternative's own text doesn't state Te Deum either
+ (never observed live: on all six real instances the alternative
+ states it unnegated). Captured rather than discarded because a
+ FUTURE Cum-Sanctissima overlay (CLAUDE.md's own binding
+ decision 2 -- an addition, never core) would have exactly this
+ shape to compare against; not compared by anything in this
+ project yet. Deliberately NOT joined by an [alt_gloria]/
+ [alt_credo]/[alt_class] -- see [split_at_alt_office]'s own
+ citation for why only this one field was ever wrong. *)
+}
+[@@deriving sexp]
+
+(* The literal separator between a day's PRIMARY (strict 1962) office and
+ a printed SECOND, Cum-Sanctissima-2020 alternative -- see the module
+ header's own point 5 for the full account of the bug this closes
+ (data/ef/expected-divergences-fiuv.sexp's own F3) and why it is a
+ DIFFERENT token from [extract_missa_fields]'s own "Vel"+"Missa"
+ two-token boundary. Exactly 6 occurrences in the whole corpus,
+ confirmed by direct grep of the raw pdftotext dump -- all 6 on the
+ exact dates F3 named (2026-03-06/07/09/12/21/24). *)
+let alt_office_marker = "\xe2\x80\x93VEL\xe2\x80\x93"
+
+let split_at_alt_office tokens =
+ let n = Array.length tokens in
+ let rec go i = if i >= n then None else if tokens.(i) = alt_office_marker then Some i else go (i + 1) in
+ match go 0 with
+ | None -> (tokens, None)
+ | Some i -> (Array.sub tokens 0 i, Some (Array.sub tokens (i + 1) (n - i - 1)))
+
+(* --- tiny hand-rolled helpers, no Str/regex ------------------------------ *)
+
+let split_ws s =
+ String.split_on_char ' ' (String.map (fun c -> if c = '\t' then ' ' else c) s)
+ |> List.filter (fun t -> t <> "")
+
+let is_all_digits s = s <> "" && String.for_all (fun c -> c >= '0' && c <= '9') s
+
+let has_prefix ~prefix s =
+ String.length s >= String.length prefix && String.sub s 0 (String.length prefix) = prefix
+
+(* Strips ASCII trailing punctuation only (.,();) -- tokens in this source
+ never carry other trailing marks that matter to comparisons below. *)
+let rec strip_trailing_punct s =
+ let n = String.length s in
+ if n = 0 then s
+ else
+ match s.[n - 1] with
+ | '.' | ',' | ')' | ';' | ':' -> strip_trailing_punct (String.sub s 0 (n - 1))
+ | _ -> s
+
+let contains s ~sub =
+ let ls = String.length s and lu = String.length sub in
+ if lu = 0 then true
+ else
+ let rec go i = if i + lu > ls then false else if String.sub s i lu = sub then true else go (i + 1) in
+ go 0
+
+let index_of s ~sub =
+ let ls = String.length s and lu = String.length sub in
+ if lu = 0 then Some 0
+ else
+ let rec go i = if i + lu > ls then None else if String.sub s i lu = sub then Some i else go (i + 1) in
+ go 0
+
+let months =
+ [| "Januarius"; "Februarius"; "Martius"; "Aprilis"; "Maius"; "Junius"; "Julius"; "Augustus"; "September";
+ "October"; "November"; "December" |]
+
+let month_index name =
+ let rec go i = if i >= Array.length months then None else if months.(i) = name then Some (i + 1) else go (i + 1) in
+ go 0
+
+(* Roman numerals, subtractive form -- only ever called once, on the
+ FIRST month header's own year token (the anchor); see the module
+ header on why every LATER header's own year, when present at all, is
+ deliberately never read. *)
+let roman_value = function
+ | 'I' -> Some 1 | 'V' -> Some 5 | 'X' -> Some 10 | 'L' -> Some 50 | 'C' -> Some 100 | 'D' -> Some 500
+ | 'M' -> Some 1000 | _ -> None
+
+let parse_roman s =
+ let n = String.length s in
+ let rec go i acc =
+ if i >= n then Some acc
+ else
+ match roman_value s.[i] with
+ | None -> None
+ | Some v -> (
+ let next = if i + 1 < n then roman_value s.[i + 1] else None in
+ match next with
+ | Some nv when nv > v -> go (i + 2) (acc + nv - v)
+ | _ -> go (i + 1) (acc + v))
+ in
+ if n = 0 then None else go 0 0
+
+(* --- day-block scanning --------------------------------------------------- *)
+
+type block = { b_day : int; b_month : int; b_year : int; b_lines : string list }
+
+let is_day_start_line toks =
+ match toks with
+ | t :: _ :: _ when String.length t >= 2 && t.[String.length t - 1] = '.' ->
+ let core = String.sub t 0 (String.length t - 1) in
+ if is_all_digits core then
+ let d = int_of_string core in
+ if d >= 1 && d <= 31 then Some d else None
+ else None
+ | _ -> None
+
+(* A month running-header/footer line: first token is a recognised month
+ name (English Nov/Dec spelled identically to Latin), optionally
+ followed by a Roman-numeral year and/or a page number -- ALWAYS pure
+ noise for the day-block scanner (never a continuation line's own
+ content), but the FIRST time a given month index is seen it also
+ drives the [cur_month]/[cur_year] state transition. *)
+let is_month_marker_line toks = match toks with m :: _ -> month_index m <> None | [] -> false
+
+(* Left-margin indulgence-notation codes ("Ind." = an indulgence is
+ available, "Plen." = a plenary one, "DFP" unexplained but structurally
+ identical -- all three print in the SAME left-margin column pdftotext
+ -layout preserves, glued onto the following content line as literal
+ leading tokens, the identical trap extract_lms_ordo.ml's own
+ [strip_left_column] already documents for the LMS PDFs' "Pl"/"Ind"
+ column. Found by tracing a real parse failure (8 December 2025's own
+ class marker split across two physical lines by "DFP" sitting between
+ "I" and "cl."), not assumed in advance: catalogued via a frequency
+ scan of every short capitalised token opening a continuation line, and
+ deliberately does NOT include "RM" (11 occurrences, ALSO short and
+ capitalised) -- checked directly, "RM" is a real citation abbreviation
+ ("Rubricae Missalis", e.g. "vide RM 440"), not a margin code, so
+ stripping it would corrupt real text for no parsing benefit. *)
+let strip_left_column tokens =
+ let is_marker t = t = "Ind." || t = "Plen." || t = "DFP" in
+ let rec go = function t :: rest when is_marker t -> go rest | ts -> ts in
+ go tokens
+
+(* The malformed trailing entry's own distinctive two-token prefix,
+ "1st Jan" -- NOT a generic "first token is a bare 4-digit number"
+ check (extract_lms_ordo.ml's own approach): that generic shape false-
+ positived here, on real body content deep in October ("...20 augusti
+ / 1885, 26 augusti 1886..." -- an indulgence-decree date citation
+ whose line-wrap happens to put a 4-digit year token first on its own
+ physical line). This source's OWN tail artefact is reliably
+ distinguished only by its literal, unique "1st Jan" opening -- found
+ by tracing the false stop, not assumed in advance. *)
+let is_stop_tail_line toks = match toks with "1st" :: "Jan" :: _ -> true | _ -> false
+
+let read_lines path =
+ let ic = open_in path in
+ let rec loop acc = match input_line ic with l -> loop (l :: acc) | exception End_of_file -> List.rev acc in
+ let ls = loop [] in
+ close_in ic;
+ ls
+
+(* --- per-block field extraction ------------------------------------------- *)
+
+(* First WHOLE-TOKEN hit of [word] in [tokens] starting at-or-after
+ [from], REJECTING a hit whose immediately-following token (stripped)
+ is [reject_next] -- the "Gloria Patri" trap, see the module header.
+ Returns the hit's own index and whether the token immediately BEFORE
+ it (stripped) is "sine" (a negation). *)
+let find_word tokens ~from ~word ~reject_next =
+ let n = Array.length tokens in
+ let rec go i =
+ if i >= n then None
+ else if strip_trailing_punct tokens.(i) = word then
+ let next_is_rejected = i + 1 < n && strip_trailing_punct tokens.(i + 1) = reject_next in
+ if next_is_rejected then go (i + 1)
+ else
+ let negated = i > 0 && strip_trailing_punct tokens.(i - 1) = "sine" in
+ Some (i, negated)
+ else go (i + 1)
+ in
+ go from
+
+(* [class_]: the roman-numeral token immediately followed by a token
+ starting "cl" (covers "cl.", "cl.,"), whole-token matched -- never a
+ substring search (the same discipline the LMS extractor's own BVM
+ numeral anchor uses, for the identical reason: "I" is a substring of
+ "II"/"III"/"IV"). Also detects the trailing "(Priv.)" two tokens later. *)
+let extract_class tokens =
+ let n = Array.length tokens in
+ let is_class_numeral t = t = "I" || t = "II" || t = "III" || t = "IV" in
+ let rec go i =
+ if i + 1 >= n then None
+ else if is_class_numeral tokens.(i) && has_prefix ~prefix:"cl" tokens.(i + 1) then
+ let base = tokens.(i) ^ " cl." in
+ if i + 2 < n && strip_trailing_punct tokens.(i + 2) = "(Priv" then Some (base ^ " (Priv.)") else Some base
+ else go (i + 1)
+ in
+ go 0
+
+let extract_title tokens =
+ let n = Array.length tokens in
+ let is_class_numeral t = t = "I" || t = "II" || t = "III" || t = "IV" in
+ let rec find i = if i + 1 >= n then n else if is_class_numeral tokens.(i) && has_prefix ~prefix:"cl" tokens.(i + 1) then i else find (i + 1) in
+ let stop = find 0 in
+ String.concat " " (Array.to_list (Array.sub tokens 0 stop))
+
+(* Te Deum: bounded to the "Ad Mat." .. "Ad Laudes" span specifically
+ (both casings occur: "Ad MAT."/"Ad Mat."/"ad Mat.", "Ad LAUDES"/
+ "Ad Laudes") -- NOT the whole block, because a day can separately
+ mention "Te Deum" in an unrelated INDULGENCE note (31 December: "Hodie
+ ad solemnem recitationem hymni Te Deum, indulgentia plenaria lucrari
+ potest" -- about gaining an indulgence for the New Year's Eve Te Deum
+ of thanksgiving, not about whether it is sung at that day's own
+ Matins) that a whole-block search would wrongly read as a positive hit. *)
+let extract_te_deum full_text =
+ let find_ci needles start =
+ List.fold_left
+ (fun acc needle ->
+ match acc with
+ | Some _ -> acc
+ | None -> (
+ match index_of (String.sub full_text start (String.length full_text - start)) ~sub:needle with
+ | Some i -> Some (start + i)
+ | None -> None))
+ None needles
+ in
+ match find_ci [ "Ad Mat."; "Ad MAT."; "ad Mat." ] 0 with
+ | None -> None
+ | Some mat_start -> (
+ let laudes_start =
+ match find_ci [ "Ad Laudes"; "Ad LAUDES"; "ad Laudes" ] mat_start with
+ | Some i -> i
+ | None -> String.length full_text
+ in
+ let span = String.sub full_text mat_start (laudes_start - mat_start) in
+ (* TWO negative phrasings the source actually uses, found live
+ (celebrant-rubrics-phase1 Phase 2, 2026-08-22): "non dicitur Te
+ Deum" (e.g. ordinary Time-after-Pentecost ferias, 27/28 November)
+ AND, separately, "sine Te Deum" (e.g. every Sunday, every Ember
+ day, every privileged Lenten/Passiontide feria carrying a
+ commemoration -- Advent I, 30 November: "...3a de homilia (cum
+ suo R), sine Te Deum."). The ORIGINAL version of this function
+ checked only the first phrasing, so "sine Te Deum" fell through
+ to the bare "Te Deum" substring test and was wrongly read as a
+ POSITIVE hit -- confirmed by grepping the raw pdftotext dump
+ directly against a first, uncorrected run's own output: EVERY
+ date this bug affected showed "sine Te Deum" in the source and
+ [Some true] in the fixture, a 100% correlation, not a handful of
+ coincidences. This is the SAME "does the source negate the
+ hymn's own name with a DIFFERENT word than the one this parser
+ already checks for" shape [find_word]'s own "Gloria Patri"
+ substring trap already documents for Gloria -- this trap simply
+ went unnoticed until Phase 2 actually compared the extracted
+ values against colitur's own output and against the raw text by
+ hand, rather than only checking coverage counts. *)
+ if contains span ~sub:"non dicitur Te Deum" then Some false
+ else if contains span ~sub:"sine Te Deum" then Some false
+ else if contains span ~sub:"Te Deum" then Some true
+ else None)
+
+(* Gloria/Credo/praef.: bounded to the FIRST Mass option only -- from the
+ token "Missa" to whichever comes first: "Vel" immediately followed by
+ "Missa" (a real second option), or a token starting "VESPER" (Vespers
+ info -- a PREFIX match, not one or two hardcoded exact spellings,
+ found necessary by tracing a real over-capture: this source uses at
+ least four surface forms depending on grammatical case and whether
+ pdftotext renders the Æ ligature as one glyph or splits it --
+ "VESPERÆ", "VESPERAS", "VESPERA" (Ad VESPERA, rare) and "VESPERÆ"
+ with the ligature reproduced as a single non-ASCII codepoint an exact-
+ string match against the two ASCII spellings alone could never catch,
+ which let a whole day's own trailing Vespers/Compline prose leak into
+ [praef] uncaught until checked against the task brief's own worked
+ example day, 19 September), or the end of the block. *)
+let extract_missa_fields tokens =
+ let n = Array.length tokens in
+ (* [Missa] (singular, 549 of the corpus's own occurrences of any "Miss-"
+ word) anchors 398 of the 400 real days; the two genuine exceptions
+ are Good Friday and Holy Saturday, which have no Mass at all in the
+ 1955-restored Holy Week (matching extract_lms_ordo.ml's own
+ identically-shaped finding for the LMS corpus). CHRISTMAS DAY is a
+ THIRD, found by tracing a real false-None result: its own rubric
+ reads "Hodie celebrantur tres Missae pr., Gloria, Credo..." ("today
+ three Masses are celebrated..."), using the PLURAL "Missae"/"Missæ"
+ (both spellings occur across the fixture's two Christmas Days) --
+ [tokens.(i) = "Missa"] alone never matches it, silently returning
+ [None] for Gloria/Credo/praef on the single most doctrinally
+ unambiguous day in the whole calendar. Widened to accept "Missae"/
+ "Missæ" too, guarded against the ONE real collision risk: "Hodie
+ prohibentur omnes Missae defunctorum..." (a Requiem-Mass-prohibition
+ notice, common at the end of many day-blocks, unrelated to the
+ day's own Mass) also contains "Missae" -- rejected here by checking
+ the immediately FOLLOWING token is not "defunctorum". *)
+ let missa_idx =
+ let rec go i =
+ if i >= n then None
+ else if (tokens.(i) = "Missa" || tokens.(i) = "Missae" || tokens.(i) = "Missæ")
+ && not (i + 1 < n && strip_trailing_punct tokens.(i + 1) = "defunctorum")
+ then Some i
+ else go (i + 1)
+ in
+ go 0
+ in
+ match missa_idx with
+ | None -> (None, None, None)
+ | Some m ->
+ let end_idx =
+ let rec go i =
+ if i >= n then n
+ else if tokens.(i) = "Vel" && i + 1 < n && tokens.(i + 1) = "Missa" then i
+ (* A THIRD, narrower alternate-Mass marker, found while auditing
+ for the "-VEL-" trap above (celebrant-rubrics-phase1 Phase 2
+ correction, 2026-08-22): "Vel (sec. decretum <<Cum sanct.>>)
+ Missa ..." offers an alternate MASS FORMULARY under the 2020
+ decree Cum Sanctissima WITHIN the same single printed office
+ (a narrower thing than [split_at_alt_office]'s own full
+ second-office split above -- there is no second "Ad Mat." set
+ of lessons here, just a second Mass choice). Checked against
+ the raw dump directly: 20 occurrences, ALL 20 of the
+ capitalised "Vel" + open-paren shape, zero false positives
+ (the ordinary lower-case "vel" conjunction, e.g. "praef.
+ comm. vel de Martyribus", never starts with an open-paren
+ token, so this check never fires on it). Every one of the 20
+ real days already restates Gloria/Credo explicitly BEFORE
+ reaching this marker, so this was NEVER a live divergence
+ (confirmed by a direct diff of every field this fixture
+ compares, before and after this hardening: zero rows
+ changed) -- added defensively, so a FUTURE year whose
+ primary Mass clause happens not to restate Gloria/Credo
+ before this marker cannot silently read the Cum-Sanctissima
+ alternative's own value instead, the identical failure shape
+ as the "-VEL-" bug above, just not yet observed live here. *)
+ else if tokens.(i) = "Vel" && i + 1 < n && has_prefix ~prefix:"(" tokens.(i + 1) then i
+ else if has_prefix ~prefix:"VESPER" tokens.(i) then i
+ else go (i + 1)
+ in
+ go (m + 1)
+ in
+ let span = Array.sub tokens m (end_idx - m) in
+ let gloria =
+ match find_word span ~from:0 ~word:"Gloria" ~reject_next:"Patri" with
+ | Some (_, negated) -> Some (not negated)
+ | None -> None
+ in
+ let credo =
+ match find_word span ~from:0 ~word:"Credo" ~reject_next:"__never__" with
+ | Some (_, negated) -> Some (not negated)
+ | None -> None
+ in
+ let praef =
+ let sn = Array.length span in
+ let rec go i = if i >= sn then None else if strip_trailing_punct span.(i) = "praef" then Some i else go (i + 1) in
+ match go 0 with
+ | None -> None
+ | Some i ->
+ let rest = Array.to_list (Array.sub span (i + 1) (sn - i - 1)) in
+ if rest = [] then None else Some (String.concat " " rest)
+ in
+ (gloria, credo, praef)
+
+(* --- top level -------------------------------------------------------------- *)
+
+let () =
+ if Array.length Sys.argv <> 4 then
+ die "usage: extract_fiuv_ordo <pdftotext-layout.txt> <source.pdf> <dest.sexp>";
+ let txt_path = Sys.argv.(1) in
+ let pdf_path = Sys.argv.(2) in
+ let dest = Sys.argv.(3) in
+ let all_lines = read_lines txt_path in
+ let arr = Array.of_list all_lines in
+ let n = Array.length arr in
+ let trimmed i = String.trim arr.(i) in
+ (* Anchor: the FIRST month-marker line carrying a parseable Roman year
+ -- established during characterisation to be line 47, "November
+ MMXXV". Search generically rather than hardcoding the line number,
+ so a re-extraction against a re-flowed pdftotext dump still finds it. *)
+ let anchor_idx, anchor_month, anchor_year =
+ let rec go i =
+ if i >= n then die "no month header with a parseable Roman-numeral year found (expected the book's own first header, e.g. \"November MMXXV\")"
+ else
+ let toks = split_ws (trimmed i) in
+ match toks with
+ | m :: y :: _ -> (
+ match (month_index m, parse_roman y) with
+ | Some mi, Some yr when yr > 1000 -> (i, mi, yr)
+ | _ -> go (i + 1))
+ | _ -> go (i + 1)
+ in
+ go 0
+ in
+ let stop_idx =
+ let rec go i = if i >= n then n else if is_stop_tail_line (split_ws (trimmed i)) then i else go (i + 1) in
+ go (anchor_idx + 1)
+ in
+ Printf.printf "extract_fiuv_ordo: Ordo body lines %d..%d (of %d total), anchor %d/%d\n" anchor_idx stop_idx n
+ anchor_month anchor_year;
+ let cur_month = ref anchor_month and cur_year = ref anchor_year in
+ let blocks = ref [] in
+ let cur = ref None in
+ let flush () = match !cur with Some b -> blocks := b :: !blocks; cur := None | None -> () in
+ for i = anchor_idx + 1 to stop_idx - 1 do
+ let raw = trimmed i in
+ if raw = "" then ()
+ else
+ let toks = split_ws raw in
+ if is_month_marker_line toks then (
+ match toks with
+ | m :: _ -> (
+ match month_index m with
+ | Some mi ->
+ if mi <> !cur_month then begin
+ if mi < !cur_month then incr cur_year;
+ cur_month := mi
+ end
+ | None -> ())
+ | [] -> ())
+ else
+ match is_day_start_line toks with
+ | Some day ->
+ flush ();
+ cur := Some { b_day = day; b_month = !cur_month; b_year = !cur_year; b_lines = toks }
+ | None -> (
+ match !cur with
+ | None -> ()
+ | Some b -> cur := Some { b with b_lines = b.b_lines @ strip_left_column toks })
+ done;
+ flush ();
+ let blocks = List.rev !blocks in
+ Printf.printf "extract_fiuv_ordo: %d day-blocks parsed\n" (List.length blocks);
+ let rows =
+ List.map
+ (fun b ->
+ let date =
+ match Date.make ~year:b.b_year ~month:b.b_month ~day:b.b_day with
+ | Ok d -> d
+ | Error e -> die "%d/%d/%d: %s" b.b_year b.b_month b.b_day e
+ in
+ let tokens = Array.of_list b.b_lines in
+ (* [b_lines] tokens 0 is the day's own leading "<n>." marker --
+ dropped before field extraction, kept for nothing (the day
+ number is already known from [b_day]). *)
+ let content = if Array.length tokens > 0 then Array.sub tokens 1 (Array.length tokens - 1) else tokens in
+ (* Cut at the "–VEL–" alternative-office boundary, if this day has
+ one, BEFORE any field extraction -- see [split_at_alt_office]'s
+ own citation. Every field below now reads the PRIMARY office
+ only; [alt_te_deum] alone also reads the second, when present. *)
+ let primary, alt = split_at_alt_office content in
+ let class_ = extract_class primary in
+ let title = extract_title primary in
+ let full_text = String.concat " " (Array.to_list primary) in
+ let te_deum = extract_te_deum full_text in
+ let alt_te_deum =
+ match alt with None -> None | Some alt_tokens -> extract_te_deum (String.concat " " (Array.to_list alt_tokens))
+ in
+ let gloria, credo, praef = extract_missa_fields primary in
+ { date = Date.to_iso8601 date; class_; title; te_deum; gloria; credo; praef; alt_te_deum })
+ blocks
+ in
+ let missing_class = List.filter (fun r -> r.class_ = None) rows in
+ let missing_gloria = List.filter (fun r -> r.gloria = None) rows in
+ let missing_credo = List.filter (fun r -> r.credo = None) rows in
+ let missing_te_deum = List.filter (fun r -> r.te_deum = None) rows in
+ if missing_class <> [] then
+ Printf.printf "extract_fiuv_ordo: WARNING %d rows with no class found: %s\n" (List.length missing_class)
+ (String.concat ", " (List.map (fun r -> r.date) missing_class));
+ Printf.printf
+ "extract_fiuv_ordo: %d rows; %d with no Gloria found; %d with no Credo found; %d with no Te Deum found\n"
+ (List.length rows) (List.length missing_gloria) (List.length missing_credo) (List.length missing_te_deum);
+ (* CHARACTERISATION, Step C of the task brief, non-negotiable: probed
+ BEFORE trusting a single row of this fixture, and RE-VERIFIED here
+ against this run's own [rows] -- a failed probe is a hard [die], the
+ same self-verifying discipline extract_lms_ordo.ml already uses,
+ never a printed claim resting on a one-off manual check that could
+ silently rot on a re-extraction. Every date/expectation pair below
+ was confirmed directly against the raw pdftotext dump before being
+ encoded here. *)
+ let creed_of_date d = List.find_map (fun r -> if String.equal r.date d then r.credo else None) rows in
+ let probes =
+ [ ("2025-11-30", true, "Advent Sunday, I class, RG 475(a)");
+ ("2025-12-01", false, "ordinary Advent feria, III class, RG 476(b)/(d)");
+ ("2025-12-26", true, "St Stephen, II class, inside the Nativity octave, RG 475(d)");
+ ("2026-08-10", false, "St Laurence, II class martyr, no Lord/BVM/apostle clause applies, RG 476(b)");
+ ("2026-08-15", true, "the Assumption, I class, RG 475(b)") ]
+ in
+ List.iter
+ (fun (d, expected, desc) ->
+ match creed_of_date d with
+ | Some got when Bool.equal got expected -> ()
+ | Some got -> die "characterisation probe FAILED: %s (%s) -- expected creed=%b, extracted creed=%b" d desc expected got
+ | None -> die "characterisation probe FAILED: %s (%s) -- no row extracted for this date at all" d desc)
+ probes;
+ Printf.printf "extract_fiuv_ordo: all %d characterisation probes confirmed\n" (List.length probes);
+ (* Second characterisation probe, same discipline: the "–VEL–"
+ alternative-office split (module header point 5) must fire on
+ EXACTLY these six dates, no more, no fewer -- a hard [die], not a
+ printed claim, so a future re-extraction against a re-flowed dump
+ silently gaining or losing an instance fails loudly rather than
+ quietly changing F3's own population. *)
+ let alt_office_dates = List.filter_map (fun r -> if r.alt_te_deum <> None then Some r.date else None) rows in
+ let expected_alt_office_dates =
+ [ "2026-03-06"; "2026-03-07"; "2026-03-09"; "2026-03-12"; "2026-03-21"; "2026-03-24" ]
+ in
+ if List.sort String.compare alt_office_dates <> expected_alt_office_dates then
+ die "alt-office-split characterisation probe FAILED: expected exactly %s, got %s"
+ (String.concat ", " expected_alt_office_dates)
+ (String.concat ", " (List.sort String.compare alt_office_dates));
+ Printf.printf "extract_fiuv_ordo: alt-office-split characterisation probe confirmed (%d dates)\n"
+ (List.length expected_alt_office_dates);
+ let probe_lines =
+ String.concat "\n"
+ (List.map
+ (fun (d, expected, desc) -> Printf.sprintf "; %s (%s) -> %s" d desc (if expected then "TRUE" else "FALSE"))
+ probes)
+ in
+ let sha256_of_file path =
+ let cmd = Printf.sprintf "sha256sum %s" (Filename.quote path) in
+ let ic = Unix.open_process_in cmd in
+ let line = try input_line ic with End_of_file -> die "sha256sum produced no output for %s" path in
+ (match Unix.close_process_in ic with Unix.WEXITED 0 -> () | _ -> die "sha256sum failed for %s" path);
+ match String.index_opt line ' ' with Some i -> String.sub line 0 i | None -> die "unexpected sha256sum output: %S" line
+ in
+ let today () =
+ let tm = Unix.gmtime (Unix.time ()) in
+ Printf.sprintf "%04d-%02d-%02d" (tm.Unix.tm_year + 1900) (tm.Unix.tm_mon + 1) tm.Unix.tm_mday
+ in
+ let pdf_sha = sha256_of_file pdf_path in
+ let pdf_base = Filename.basename pdf_path in
+ let gloria_true = List.length (List.filter (fun r -> r.gloria = Some true) rows) in
+ let gloria_false = List.length (List.filter (fun r -> r.gloria = Some false) rows) in
+ let credo_true = List.length (List.filter (fun r -> r.credo = Some true) rows) in
+ let credo_false = List.length (List.filter (fun r -> r.credo = Some false) rows) in
+ let te_deum_true = List.length (List.filter (fun r -> r.te_deum = Some true) rows) in
+ let te_deum_false = List.length (List.filter (fun r -> r.te_deum = Some false) rows) in
+ let alt_te_deum_true = List.length (List.filter (fun r -> r.alt_te_deum = Some true) rows) in
+ let alt_te_deum_false = List.length (List.filter (fun r -> r.alt_te_deum = Some false) rows) in
+ let alt_te_deum_none = List.length (List.filter (fun r -> r.alt_te_deum = None) rows) in
+ let header =
+ Printf.sprintf
+ {|; %s -- Witnesses task (2026-08-22-colitur-
+; celebrant-rubrics-phase1)'s SEVENTH validation layer, and the first
+; UNIVERSAL (non-diocesan), SECOND-COMPILER witness: the FIUV (Foederatio
+; Internationalis Una Voce) Ordo, compiled independently of the Latin Mass
+; Society's own three editions (test/test_lms_ordo.ml) -- see
+; docs/research/ordo/PROVENANCE-ordo-corpus.md for the full four-reasons
+; account of why this source is the more valuable of the two newly-
+; acquired ones.
+;
+; Generator: tools/extract_fiuv_ordo.ml -- do not hand-edit; re-run
+; against a fresh pdftotext dump and commit the diff instead.
+;
+; Source: "Ordo Divini Officii recitandi sacrique peragendi secundum
+; antiquam Ritus Romani formam pro anno Domini 2026", Foederatio
+; Internationalis Una Voce, compiled by Joseph Shaw (Praefatio signed
+; "Joseph Shaw, President", Feast of the Nativity of Our Lady [8 Sept]
+; 2025) -- title page and Praefatio, verified directly against the PDF's
+; own extracted text. 2025 is stated as the Federation's own 60th
+; anniversary year.
+; PDF metadata: Creator "TeX", Producer "pdfTeX-1.40.22", CreationDate
+; 2025-11-08, 128 pages.
+; URL: identified via web search as
+; https://lms.org.uk/sites/default/files/u5374/fiuv_ordo_2025-2026_1.1_a5_format.pdf
+; (hosted by the Latin Mass Society's own site alongside its own three
+; editions -- Joseph Shaw is both FIUV President and LMS Chairman, which
+; plausibly also explains why this PDF's own trailing malformed entry
+; (see COVERAGE below) is structurally identical to the LMS PDFs' own,
+; despite the two being compiled by different people); NOT independently
+; re-downloaded and byte-compared in this session -- recorded honestly
+; rather than presented as verified, the same discipline the LMS
+; fixtures' own headers use.
+; The local PDF (docs/research/ordo/%s) is gitignored (CLAUDE.md: "docs/
+; is gitignored"), so this header, not git, is this fixture's only record
+; of where it came from.
+; SHA-256 of docs/research/ordo/%s: %s
+; Extracted (UTC): %s
+; Exact commands:
+; pdftotext -layout docs/research/ordo/%s /tmp/fiuv.txt
+; dune exec tools/extract_fiuv_ordo.exe -- /tmp/fiuv.txt \
+; docs/research/ordo/%s %s
+;
+; COVERAGE: %d day-rows, %s through %s. Starts at the book's own first
+; real content line ("November MMXXV", the FIRST month header found with
+; a parseable Roman-numeral year -- used as this parser's sole date
+; anchor, see the tool's own header on why no LATER header's year is ever
+; read, the second one being a publisher typo, "December MMXV" for 2025).
+; Stops before a malformed trailing entry, structurally the SAME artefact
+; found in all three LMS editions (a stray, mislabelled duplicate of the
+; fixture's own already-extracted 1 January row, headed "1st Jan / 2025."
+; -- see extract_lms_ordo.ml's own COVERAGE section for the fuller
+; account of this shared publisher-side quirk).
+;
+; CHARACTERISATION FINDINGS (non-negotiable per the task brief -- the
+; extraordinaryform.org Ordo silently omitted St Lawrence's vigil on
+; EVERY date it covered, which nearly produced a false corroboration
+; during the RG 33 work; this Ordo was probed with the same discipline
+; before a single divergence was adjudicated):
+;
+; 1. UNIVERSAL, not diocesan: unlike the three LMS editions (which had to
+; exclude diocesan variants from every comparison, quantified in each
+; of their own fixture headers), this Ordo carries ONE block per day,
+; always. There is no [has_diocesan_variant] field in this fixture's
+; own [row] at all -- there is nothing to exclude.
+;
+; 2. OPPOSITE-PREDICTION PROBES against Rite_ef.Rubrics_ef.creed (RG
+; 475-476), RE-CHECKED BY THIS TOOL RUN against its own extracted
+; [rows] before this header was even written -- a failed probe is a
+; hard [die], never a printed claim:
+%s
+; All %d predictions confirmed, both directions -- this source is
+; discriminating, not a constant, on the Creed.
+;
+; 3. THE SUBSTRING TRAP this parser defends against, found live in this
+; corpus: "Gloria" is also the first word of "Gloria Patri" (the psalm
+; doxology), which can appear INSIDE the very Mass clause being
+; scanned (Good Friday, 3 April 2026: "Missa pr., (omittuntur ps.
+; Iudica me et Gloria Patri), Gloria, sine Credo, praef. comm." -- the
+; real, standalone Gloria mention follows immediately after). See
+; [find_word]'s own citation.
+;
+; 4. RAW SUBSTRING COUNTS measured against this session's own pdftotext
+; dump, for comparison against the task brief's own figures (328
+; Gloria / 132 sine Gloria / 119 Credo / 260 sine Credo, 262 Te Deum):
+; "sine Gloria" 132 (EXACT match), "sine Credo" 260 (EXACT match), raw
+; "Te Deum" substring 262 (EXACT match, includes negated occurrences
+; as substring hits -- 42 of them are "non dicitur Te Deum"). Raw
+; "Gloria"/"Credo" substring totals (494/411) do NOT match the brief's
+; implied positive totals (328/119) even after subtracting the 25
+; "Gloria Patri" occurrences (469, still 9 over) -- not chased further
+; here; this extractor's own token-scoped, Missa-clause-bounded counts
+; below are what this fixture actually asserts, not a forced match to
+; the brief's approximate figures.
+;
+; 5. THE ALTERNATIVE-OFFICE TRAP (celebrant-rubrics-phase1 Phase 2
+; correction, 2026-08-22, closing F3): six real days print BOTH a
+; strict-1962 reading and a complete SECOND office (the same saint
+; kept outright under the 2020 decree Cum Sanctissima), separated by
+; a literal "-VEL-" token. [extract_te_deum] used to scan the WHOLE
+; block unbounded, so on these six days -- whose PRIMARY office states
+; no "Ad Mat." of its own -- it silently read the SECOND office's own
+; Te Deum instead, manufacturing a same-day contradiction with colitur
+; that F3 first recorded as "colitur's own possible gap". Every field
+; now reads the PRIMARY office only (module header point 5 has the
+; full account); the alternative is captured, not discarded, in its
+; own [alt_te_deum] field. Confirmed to fire on EXACTLY six dates by a
+; hard characterisation probe (this run's own output, above), not
+; merely asserted here.
+;
+; MEASURED DISTRIBUTION (this extraction, not the brief's figures): Gloria
+; true=%d false=%d unresolved=%d; Credo true=%d false=%d unresolved=%d;
+; Te Deum true=%d false=%d unresolved=%d; alt Te Deum (the six "-VEL-"
+; Cum Sanctissima alternatives, point 5 above) true=%d false=%d none=%d
+; (a day contributes to "none" here either because it has no "-VEL-"
+; alternative at all -- 394 of 400 -- or, for a day that does, because
+; that alternative's own text doesn't state Te Deum either; not
+; distinguished by this one count).
+;
+; FIX HISTORY (fixed text, HISTORICAL counts from the run that found each
+; bug, not re-verified live by every later run -- this section exists so
+; regenerating this fixture from a fresh pdftotext dump, which rewrites
+; this whole header from the template, does not silently erase either
+; finding; earlier this header carried the first entry below as a
+; hand-appended note that a regeneration between then and now had
+; already dropped once):
+;
+; 1. RE-EXTRACTED (celebrant-rubrics-phase1 Phase 2, 2026-08-22): the FIRST
+; extraction's own Te Deum parser recognised only ONE of the source's two
+; negative phrasings ("non dicitur Te Deum") -- "sine Te Deum" (how the
+; source actually negates a SUNDAY's own Te Deum, among others) fell
+; through to a bare "Te Deum" substring match and was wrongly read
+; [true]. Found by hand, comparing the raw pdftotext dump against this
+; fixture's own values directly, not by any coverage check (both counted
+; the same 127 unresolved rows before and after). Fixed in
+; tools/extract_fiuv_ordo.ml's own [extract_te_deum] (see its own
+; citation); this fixture re-extracted from the SAME pdftotext dump
+; (SHA-256 of the source PDF unchanged, above). 24 of 400 rows changed,
+; every one Te Deum true->false, every one independently confirmed
+; against the raw text: true 231->207, false 42->66, unresolved 127
+; unchanged. See lib/rites/rite_ef/rubrics_ef.ml's own [te_deum] header
+; (237(a)'s comment) for the full account of what this corrected, and
+; the .superpowers task report for the complete before/after date list.
+;
+; 2. RE-EXTRACTED (celebrant-rubrics-phase1 Phase 2 correction, 2026-08-22):
+; module header point 5 and CHARACTERISATION point 5 above have the
+; full account of the "-VEL-" alternative-office bug this closes
+; (data/ef/expected-divergences-fiuv.sexp's own F3). This fixture
+; re-extracted from the SAME pdftotext dump (SHA-256 of the source PDF
+; unchanged, above). 6 of 400 rows changed, every one Te Deum
+; true->unresolved (2026-03-06/07/09/12/21/24), every one independently
+; confirmed against the raw text: Te Deum true 207->201 (this run's own
+; unresolved count went 127->133, consistent with those same 6 rows).
+; Gloria/Credo/class_/title/praef are UNCHANGED on all 400 rows --
+; those fields were already correctly bounded to the primary office
+; before this fix (a Vespers-line token always intervened before the
+; "-VEL-" separator on all six real instances), confirmed by a direct
+; diff of the fixture before and after, not merely by argument.
+;
+; 3. RE-EXTRACTED (celebrant-rubrics-phase1 same correction, defensive
+; hardening, 2026-08-22): auditing for the SAME trap elsewhere (the
+; task brief's own item 5) found a THIRD, narrower alternate-Mass
+; marker, "Vel (sec. decretum <<Cum sanct.>>) Missa ...", offering a
+; Cum-Sanctissima Mass FORMULARY choice within a single office
+; (unlike point 5 above, no second "Ad Mat." set of lessons -- a
+; narrower thing than a full alternate office). 20 occurrences in the
+; whole corpus, confirmed by direct grep; on all 20, Gloria/Credo were
+; ALREADY stated before reaching this marker, so this was NEVER a
+; live divergence for those two fields (confirmed by this same diff:
+; 0 Gloria/Credo/class_/title/alt_te_deum rows changed). [praef]
+; alone changed on 20 of 400 rows (its own "rest of the primary
+; clause" capture no longer runs past this marker into the
+; alternative's own text) -- harmless in itself ([praef] is
+; documented CAPTURED, NOT VALIDATED, compared by nothing), but added
+; as a defensive boundary in [extract_missa_fields] so a FUTURE year
+; whose primary clause does not restate Gloria/Credo before this
+; marker cannot silently read the alternative's own value instead --
+; the identical failure shape as point 2 above, closed before it was
+; ever observed live rather than after.
+|}
+ dest pdf_base pdf_base pdf_sha (today ()) pdf_base pdf_base dest (List.length rows)
+ (match rows with r :: _ -> r.date | [] -> "?")
+ (match List.rev rows with r :: _ -> r.date | [] -> "?")
+ probe_lines (List.length probes) gloria_true gloria_false (List.length missing_gloria) credo_true credo_false
+ (List.length missing_credo) te_deum_true te_deum_false (List.length missing_te_deum) alt_te_deum_true
+ alt_te_deum_false alt_te_deum_none
+ in
+ let body = Sexplib.Sexp.to_string_hum ~indent:2 (sexp_of_list sexp_of_row rows) in
+ let oc = open_out dest in
+ Fun.protect
+ ~finally:(fun () -> close_out_noerr oc)
+ (fun () ->
+ output_string oc header;
+ output_string oc body;
+ output_string oc "\n");
+ Printf.printf "extract_fiuv_ordo: wrote %d rows to %s\n" (List.length rows) dest