<feed xmlns='http://www.w3.org/2005/Atom'>
<title>colitur.git/lib/citation/book.ml, branch v1.1.0</title>
<subtitle>deterministic OCaml engine to compute and validate liturgical calendars for multiple rites, template-driven output to year 9999</subtitle>
<id>https://git.labunix.xyz/colitur.git/atom?h=v1.1.0</id>
<link rel='self' href='https://git.labunix.xyz/colitur.git/atom?h=v1.1.0'/>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/'/>
<updated>2026-08-26T11:30:02Z</updated>
<entry>
<title>fix(citation): recognise English-canonical OF books, verse sub-letters</title>
<updated>2026-08-26T11:30:02Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-08-26T11:30:02Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/commit/?id=9cc763c169a6ce4ebbdbed506c5eb2dad1f9f3c1'/>
<id>urn:sha1:9cc763c169a6ce4ebbdbed506c5eb2dad1f9f3c1</id>
<content type='text'>
35% of OF citation fields (259/730 on colitur readings --rite of 2026)
printed unconverted -- 1 John renders on 2 January but not 3 January.
Two independent causes, both in the citation/siglum path, neither in
the OF data itself:

1. Book.table only ever surveyed the three EF citation-bearing files,
   so 345 references citing a book no EF file happens to use (Job,
   Ruth, Judges, 1/2 Samuel, 1/2 Chronicles, 1/2 Maccabees, Baruch,
   Ecclesiastes, Habakkuk, Haggai, Nahum, Zechariah, Zephaniah,
   Deuteronomy, Amos, Micah, Lamentations, Ezra, Joshua, 2/3 John,
   Jude, Philemon, plus "Isaiah"/"Jeremiah"/"Ezekiel"/"Malachi"/"Mat"/
   "The Acts"/"Tobit"/"Song of Solomon" spelling variants of books EF
   already knows) failed as "unknown book". Added as a new, separate
   of_lectionary_table rather than folded into the EF-surveyed table:
   none of these 25 new books is attested in the EF's own scans, and
   lang/la.ini's own header refuses to fabricate an uncited Latin
   title, so they are resolvable (parse + render, falling back to
   their own English spelling) but deliberately excluded from Book.all
   -- test_lang_coverage.ml's la.ini-completeness promise is preserved
   exactly for the ids it already covered, not silently weakened.

2. Parse's grammar could not read a verse number carrying a lectionary
   sub-verse letter ("11a", "1bcde") at all -- the dominant remaining
   failure shape once (1) was fixed. verse_range now carries a
   verse_num { n; suffix } on each boundary, PRESERVED through
   rendering rather than dropped (dropping would silently lose real
   precision the source text carries). Chapter numbers are untouched
   (nothing in the data ever attaches a letter to one).

A third, subtler bug surfaced by (1): registering "jude"/"philemon"/
"2 John"/"3 John" exposed Parse's existing "leading comma-number is a
chapter" heuristic misreading a single-chapter book's bare verse list
("Jude 17,20b-25") as chapter 17 -- a wrong PARSE, worse than the
previous safe "unknown book" failure. Book.is_single_chapter now tells
Parse to skip that heuristic for the four one-chapter books and default
to chapter 1.

Residual, honestly enumerated rather than forced to zero: 41 distinct
references (of 1540) are hyphenated ranges crossing a chapter boundary
("2:29-3:6") -- a Parse.t shape verse_range/part do not represent, a
type restructuring deliberately not attempted this task. Pinned exactly
by the new test_citation_coverage_of.ml, both directions (a new failure
or one of these 41 starting to convert both go red), and disclosed in
data/of/lectionary.sexp's own regenerated provenance header (tools/
bootstrap_lectionary_of.ml now runs the same parser at generation time
and names the count and the set).

Verified EF-unaffected: git diff v1.0.0..HEAD -- lib/kernel/
lib/rites/rite_ef/ data/ef/ is empty, and `colitur day`/`readings`
output for 2027 is byte-identical against the pre-fix binary.
</content>
</entry>
<entry>
<title>fix: audit findings — parser strictness, name ambiguity, and errors</title>
<updated>2026-08-20T20:11:17Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-08-20T20:11:17Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/commit/?id=1988d350242b47aa52aa07904c495e7e2c0eba82'/>
<id>urn:sha1:1988d350242b47aa52aa07904c495e7e2c0eba82</id>
<content type='text'>
Found by auditing the shipped program rather than the diff.

The parser accepted OCaml integer-literal syntax, so "Luke 1_1:5" read as
chapter ELEVEN and "+5" as 5 -- a typo silently becoming a different
chapter, reachable through any user overlay. Numbers are now plain digits
and positive, and a descending range is rejected: 1:20-10 is always a
transcription error. No shipped citation changed.

FOUR PAIRS OF DIFFERENT BOOKS SHARED A FULL TITLE. 1 and 2 Corinthians
both rendered "Epistola ad Corinthios", as did Thessalonians, Timothy and
Peter -- 108 citations in 2027 alone that a reader cannot resolve to a
book. This is the Kings defect fixed earlier and not generalised. The
titles now carry their volume numeral, marked CONSTRUCTED, and a test
asserts no two books share a name -- while allowing the case where two
ids ARE the same book under different numbering, which a tradition
relates.

Spec section 8.5 is now delivered rather than merely recorded. Shipped
styles did not re-parse their own output: 32 of 52 Latin abbreviations
and 49 of 52 full titles failed, so a citation copied from colitur's own
output into an overlay was passed through untouched and printed in the
wrong language, silently. Every shipped name is registered as a spelling
and split_book learned multi-word titles by longest-token match. Now 0
of 52 fail beyond the same-book aliases.

Overlay errors were written for a compiler author: they named an OCaml
source file the reader does not have and buried the useful token. The
existing five-path rewriter is replaced by a generic one, applied to
every load path rather than one, so "rank: is not one of the allowed
values (at Class9)" replaces the raw Of_sexp_error dump.

Also: the new-overlay scaffold documented citations and layer without
showing them, and its comment implied the wrong nesting -- the single
easiest thing to get wrong; error messages echoed whole file lines,
copying an unrelated file's contents into stderr when a flag pointed at
one; and config --show validated partway down its table, exiting 2 after
writing five rows to stdout.
</content>
</entry>
<entry>
<title>feat(citation): default_spelling, and close two test gaps</title>
<updated>2026-08-20T13:13:37Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-08-20T13:13:37Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/commit/?id=ec70286daa72fd2a2a78a1cbab0fe2c268e5c743'/>
<id>urn:sha1:ec70286daa72fd2a2a78a1cbab0fe2c268e5c743</id>
<content type='text'>
Book.default_spelling returns the first registered spelling for an id. It
is the fallback display name, and it exists because the alternative is
worse: a language file's [bible] lookup is total and returns THE KEY on a
miss, so a book with no entry would render as "luke.abbr 5:12-14".
Falling back to the data's own spelling makes it render as "Luke
5:12-14" instead -- what colitur printed before this feature existed.
The degraded case is the old behaviour, the same principle Lang states
for its own key-returning misses.

Two test gaps closed, both found by mutation rather than by reading:

Parse's split_book scans a leading ordinal digit over '1'..'4', and no
case in the suite used an ordinal above 1. Narrowing the range to
'1'..'3' passed every test while seven real citations depend on it
("3 Kings 17:8-16", "4 Kings 5:1-15"). Parse-layer cases added; the
first attempt at this test asserted through Book.of_token, which is a
table lookup and never reaches split_book at all.

default_spelling is asserted to round-trip: every cited id's fallback
spelling must itself resolve back to that id, or Render and Parse
disagree the moment a book goes unnamed.
</content>
</entry>
<entry>
<title>fix(citation): survey all three citation-bearing files, not just the lectionary</title>
<updated>2026-08-20T12:54:01Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-08-20T12:54:01Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/commit/?id=a4eb9ed3cd93cbac493a873991dd5817906aad56'/>
<id>urn:sha1:a4eb9ed3cd93cbac493a873991dd5817906aad56</id>
<content type='text'>
The book table was built against data/ef/lectionary.sexp alone. That
undercounts: sanctoral.sexp carries more citations than the lectionary
and, together with commons.sexp, cites 21 book tokens the table was
missing entirely (62 distinct tokens across all three files, not 42),
several of them common (2 Tim appears 8+ times).

Add the missing spellings to existing ids (2 Cor, Col., Exod, Ezek,
James, Sir, Eccli) and ten new ids for books not cited before (1-2
Timothy, 2 Peter, Apocalypse, Judith, Malachi, Proverbs, Song of Songs,
Tobit, Wisdom). Sir and Rev are modern spellings sitting inside
Vulgate data, so both resolve to their Vulgate ids (ecclesiasticus,
apocalypse) rather than to the sirach/revelation tradition targets --
mapping them to a second id would double-map the same book.

Add a duplicate-spelling invariant test (List.assoc_opt would silently
prefer the first match on a collision) and a test that re-derives the
token set from all three data files at test time and asserts every
token resolves, rather than trusting a survey performed once by hand.
</content>
</entry>
<entry>
<title>feat(citation): canonical book ids and tradition mapping</title>
<updated>2026-08-20T12:40:12Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-08-20T12:40:12Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/colitur.git/commit/?id=405910d2fd245e7a11e09eecb8c6fffb68d2169c'/>
<id>urn:sha1:405910d2fd245e7a11e09eecb8c6fffb68d2169c</id>
<content type='text'>
Seven books arrive in two spellings, inherited from lectio's ini and
ultimately from Divinum Officium. Collapse them onto one id here rather
than editing generated data.

Naming and renumbering are kept apart: a tradition decides which book an
id denotes, a language file decides what it is called.
</content>
</entry>
</feed>
