aboutsummaryrefslogtreecommitdiff
path: root/internal/bible/corpora_optional
Commit message (Collapse)AuthorAgeFilesLines
* bible(wuj): strip leaked verse numbers from the deuterocanon textLukasz Kasprzak2026-07-291-3797/+3797
| | | | | | | | | | | | | | ~3,800 verses carried their own verse number at the start of the text (a biblia.info.pl harvest artifact concentrated in the deuterocanonical books -- Sirach, 1-2 Maccabees, Wisdom, Judith, Baruch, at 90-98% of each, plus 3 strays in Ezra/Proverbs). E.g. Sirach 4:4 read "4 Nie odrzucaj..." instead of "Nie odrzucaj...". It passed every check: verse counts are unaffected, and the verses aren't short. Strip a leading token only where it exactly equals the verse number (so a genuine "40 dni" in a non-matching verse is untouched), and only when text follows. 3797 verses fixed; 35,810-row parity preserved, no verse emptied, no new warnings; --corpus-check wuj still clean; make test green.
* bible(grb): restrict to the Vulgate canon; restore Hosea and ZechariahLukasz Kasprzak2026-07-291-26696/+22951
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | grb failed corpus-check outright: 15 errors, and two whole books absent from lectio at runtime with no error and no warning. Upstream grb is a standalone Septuagint reader carrying 87 books. lectio needs one uniform canon, so scripts/gen-grb-lectio.py derives the 73 books vul has rather than forking -- upstream stays intact and keeps everything. Eleven books have no Vulgate counterpart and are dropped, including 2 Esdras, which upstream's README explains IS Ezra + Nehemiah under the Septuagint's name; verified byte-identical, so nothing is lost. Five books ARE canonical under other names and are remapped, and the verse counts show which witness the Vulgate follows -- Jerome translated Theodotion, not the Old Greek: Bel and the Dragon (Theodotion) 42 = vul Daniel 14 (42) exact Bel and the Dragon (LXX) 37 no Sussana (Theodotion) 64 = vul Daniel 13 (65) Sussana (LXX) 37 no Letter of Jeremiah 73 = vul Baruch 6 (72) Wisdom of Solomon 435 = vul Wisdom (439) Upstream's plain "Daniel" is the Old Greek and is missing chapter 4 outright, so Theodotion supplies Daniel throughout: complete, and the tradition the lectionary actually cites. lectio thereby GAINS Wisdom, Daniel 4, Daniel 13-14 and Baruch 6 in Greek rather than losing them. Three upstream defects are repaired in transit: * 408 rows carried 5 fields, not 6 -- a lost tab fused the book number and chapter ("281" = book 28, chapter 1). Every one was in Hosea or Zechariah, and neither book had a single well-formed row, so bible.go skipped both entirely. Hosea (197 verses) and Zechariah (211) are back. * 5 merged verse labels ("27-28") that strconv.Atoi turns into verse 0. * A UTF-8 BOM welded to the first book name, making "Genesis" a phantom 74th book matching nothing. The SBLGNT apparatus sigla are stripped, as TestGrbNoApparatusMarkers requires; regenerating from raw upstream reintroduces ~8700 of them. grb now passes. The remaining warnings are the Septuagint being itself -- Jeremiah is LXX-numbered (grb 33:2 is the Vulgate's 26:2), Esther integrates its additions into chapters 1-10, LXX Malachi has three chapters, and 3 Kingdoms carries supplements like 10:22a that upstream stores as a duplicate verse 22. Silencing those would mean deleting real Greek text.
* bible(drb): renumber three chapters, fill eight absent versesLukasz Kasprzak2026-07-291-4618/+4626
| | | | | | | | | | | | | | | | | | | | | | Eight corpus-check warnings, of two kinds. Three chapters (Exodus 40, Genesis 49, Song of Solomon 1) held exactly as many verses as the Vulgate but numbered with a gap. The count pins the mapping, so renumbering 1..N is safe and changes no text. Five chapters were genuinely short. The missing verses came from get.bible's douayrheims -- the same public-domain source and API scripts/gen-deutero.py already uses for this corpus, so no new provenance is introduced. Scope was deliberately narrowed to the flagged chapters. A first pass over every chapter shorter than the Vulgate pulled in 938 verses, 820 of them in Psalms -- but drb.ini declares psalm_system = drb, a different numbering, so matching by Vulgate verse number there would have inserted the wrong text under the wrong numbers across the Psalter. One warning remains and is not a defect: Baruch 6 runs 1-72 with verse 37 absent in drb and in get.bible alike, a genuine Douay/Vulgate versification difference faithfully represented.
* bible(wuj): rebuild from source, repair to exact Vulgate parityLukasz Kasprzak2026-07-291-362/+314
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The Wujek corpus was missing text and mis-numbering what it had: 453 corpus-check warnings, 8 absent chapters, 84 duplicate rows, and 357 chapters holding fewer verses than the Vulgate. Most of it was never a sourcing problem. The original harvest of biblia.info.pl assumed one verse is one <p>, but many psalms print the superscription in the anchored paragraph and the psalm's opening line in an unanchored one behind a drop cap -- so 44 psalms lost their first line, and Acts 6:5 and others went the same way. Re-harvested with the anchor treated as a verse START rather than a whole verse (scripts/scrape-wujek.py), which also had to cope with four anchor shapes across books, hidden page markers opening paragraphs mid-verse, chapter ids that are simply wrong (Mark labels 87 anchors "15:*" across chapters 14-16), and Psalms heading its divisions "Psalm CXVII" where every other book says "Rozdzial N". Structure was then repaired against the Clementine Vulgate, by hand where a rule would have guessed: * The Acts 5 interleaved three streams -- Acts 4:27-37 duplicated, the real 5:1-37, and 5:38-39 mislabelled 28-29. Verified against the Latin, the duplicates dropped, the two renumbered. Acts 4:27 kept the site's own wording in place of a 1962 reading, removing a seam. * The site heads psalms by Hebrew division; its "Psalm 114" is two psalms. Split into Vulgate 113:9-26 and 114:1-9, confirmed at all four boundaries. * A corrupt anchor for 1 Chronicles 9:11 pushed 34 paragraphs of genealogy into chapter 10. Counts corroborate: 10 + 34 = 44 verses, remainder 14, both exactly the Vulgate's. * 16 chapters looked short at the end; only two verses were truly absent. The rest were merges, split at anchors read off the Latin -- including one across a chapter boundary (Colossians 4:1 sat inside 3:25) and four numbering offsets where a mid-chapter merge shifted everything after it. wuj now carries exactly the Vulgate's verse set: 35810 rows, every one of the 1334 chapters matching, no duplicates, no gaps. corpus-check: 0 warnings. The ~300 verses that could not come from biblia.info.pl are recorded in NOTICE; they are not public domain.
* bible: embed only the Vulgate by default; other corpora opt-inLukasz Kasprzak2026-07-286-0/+110763
Split the bundled corpora so the default binary carries only the public-domain Latin Vulgate. wuj/drb/grb move to corpora_optional/ and are compiled in only with `-tags fullbible`, or dropped into the user corpora dir at runtime. This keeps the distributed AGPL binary free of third-party scripture text -- bible corpora are not covered by lectio's licence. - bible: read corpus files/metadata from core + optional embed FS (embReadCorpus/embCorpusFiles); optionalFS is set only under the tag. - config: default `versions` now reflects corpora actually available in the build (Vulgate first), so a vul-only binary offers no absent versions. - Makefile: `make build` = vul only; `make build-full` embeds the rest; `make test` runs with -tags fullbible; check-corpora validates both dirs. - tests: corpusmeta asserts on vul (the always-embedded corpus); the versions-default expectation follows availableVersions().