<feed xmlns='http://www.w3.org/2005/Atom'>
<title>lectio.git/internal/bible/corpora_optional, branch v0.46.0</title>
<subtitle>offline Catholic daily readings and liturgical calendar in Go, with CLI, TUI and web clients</subtitle>
<id>https://git.labunix.xyz/lectio.git/atom?h=v0.46.0</id>
<link rel='self' href='https://git.labunix.xyz/lectio.git/atom?h=v0.46.0'/>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/'/>
<updated>2026-07-29T11:20:56Z</updated>
<entry>
<title>bible(wuj): strip leaked verse numbers from the deuterocanon text</title>
<updated>2026-07-29T11:20:56Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-07-29T11:20:56Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/commit/?id=8c15fac1c65806f97e1bfec2fd8af3e17d8d7b22'/>
<id>urn:sha1:8c15fac1c65806f97e1bfec2fd8af3e17d8d7b22</id>
<content type='text'>
~3,800 verses carried their own verse number at the start of the text (a
biblia.info.pl harvest artifact concentrated in the deuterocanonical books --
Sirach, 1-2 Maccabees, Wisdom, Judith, Baruch, at 90-98% of each, plus 3 strays
in Ezra/Proverbs). E.g. Sirach 4:4 read "4 Nie odrzucaj..." instead of "Nie
odrzucaj...". It passed every check: verse counts are unaffected, and the
verses aren't short.

Strip a leading token only where it exactly equals the verse number (so a
genuine "40 dni" in a non-matching verse is untouched), and only when text
follows. 3797 verses fixed; 35,810-row parity preserved, no verse emptied, no
new warnings; --corpus-check wuj still clean; make test green.
</content>
</entry>
<entry>
<title>bible(grb): restrict to the Vulgate canon; restore Hosea and Zechariah</title>
<updated>2026-07-29T09:13:03Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-07-29T09:13:03Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/commit/?id=ab9a02b34d88e6c85e4d494077bb474dc19adbce'/>
<id>urn:sha1:ab9a02b34d88e6c85e4d494077bb474dc19adbce</id>
<content type='text'>
grb failed corpus-check outright: 15 errors, and two whole books absent from
lectio at runtime with no error and no warning.

Upstream grb is a standalone Septuagint reader carrying 87 books. lectio needs
one uniform canon, so scripts/gen-grb-lectio.py derives the 73 books vul has
rather than forking -- upstream stays intact and keeps everything.

Eleven books have no Vulgate counterpart and are dropped, including 2 Esdras,
which upstream's README explains IS Ezra + Nehemiah under the Septuagint's
name; verified byte-identical, so nothing is lost. Five books ARE canonical
under other names and are remapped, and the verse counts show which witness
the Vulgate follows -- Jerome translated Theodotion, not the Old Greek:

    Bel and the Dragon (Theodotion)  42 = vul Daniel 14 (42)   exact
    Bel and the Dragon (LXX)         37                        no
    Sussana (Theodotion)             64 = vul Daniel 13 (65)
    Sussana (LXX)                    37                        no
    Letter of Jeremiah               73 = vul Baruch 6  (72)
    Wisdom of Solomon               435 = vul Wisdom   (439)

Upstream's plain "Daniel" is the Old Greek and is missing chapter 4 outright,
so Theodotion supplies Daniel throughout: complete, and the tradition the
lectionary actually cites. lectio thereby GAINS Wisdom, Daniel 4, Daniel 13-14
and Baruch 6 in Greek rather than losing them.

Three upstream defects are repaired in transit:

  * 408 rows carried 5 fields, not 6 -- a lost tab fused the book number and
    chapter ("281" = book 28, chapter 1). Every one was in Hosea or Zechariah,
    and neither book had a single well-formed row, so bible.go skipped both
    entirely. Hosea (197 verses) and Zechariah (211) are back.
  * 5 merged verse labels ("27-28") that strconv.Atoi turns into verse 0.
  * A UTF-8 BOM welded to the first book name, making "Genesis" a phantom
    74th book matching nothing.

The SBLGNT apparatus sigla are stripped, as TestGrbNoApparatusMarkers
requires; regenerating from raw upstream reintroduces ~8700 of them.

grb now passes. The remaining warnings are the Septuagint being itself --
Jeremiah is LXX-numbered (grb 33:2 is the Vulgate's 26:2), Esther integrates
its additions into chapters 1-10, LXX Malachi has three chapters, and 3
Kingdoms carries supplements like 10:22a that upstream stores as a duplicate
verse 22. Silencing those would mean deleting real Greek text.
</content>
</entry>
<entry>
<title>bible(drb): renumber three chapters, fill eight absent verses</title>
<updated>2026-07-29T09:12:41Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-07-29T09:12:41Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/commit/?id=fe2a08daf55daa74810cf6e17ab49b9d8b53c89e'/>
<id>urn:sha1:fe2a08daf55daa74810cf6e17ab49b9d8b53c89e</id>
<content type='text'>
Eight corpus-check warnings, of two kinds.

Three chapters (Exodus 40, Genesis 49, Song of Solomon 1) held exactly as many
verses as the Vulgate but numbered with a gap. The count pins the mapping, so
renumbering 1..N is safe and changes no text.

Five chapters were genuinely short. The missing verses came from get.bible's
douayrheims -- the same public-domain source and API scripts/gen-deutero.py
already uses for this corpus, so no new provenance is introduced.

Scope was deliberately narrowed to the flagged chapters. A first pass over
every chapter shorter than the Vulgate pulled in 938 verses, 820 of them in
Psalms -- but drb.ini declares psalm_system = drb, a different numbering, so
matching by Vulgate verse number there would have inserted the wrong text
under the wrong numbers across the Psalter.

One warning remains and is not a defect: Baruch 6 runs 1-72 with verse 37
absent in drb and in get.bible alike, a genuine Douay/Vulgate versification
difference faithfully represented.
</content>
</entry>
<entry>
<title>bible(wuj): rebuild from source, repair to exact Vulgate parity</title>
<updated>2026-07-29T09:12:24Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-07-29T09:12:24Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/commit/?id=390535ff00d61df63a15de6fba23761a2c57cead'/>
<id>urn:sha1:390535ff00d61df63a15de6fba23761a2c57cead</id>
<content type='text'>
The Wujek corpus was missing text and mis-numbering what it had: 453
corpus-check warnings, 8 absent chapters, 84 duplicate rows, and 357 chapters
holding fewer verses than the Vulgate.

Most of it was never a sourcing problem. The original harvest of
biblia.info.pl assumed one verse is one &lt;p&gt;, but many psalms print the
superscription in the anchored paragraph and the psalm's opening line in an
unanchored one behind a drop cap -- so 44 psalms lost their first line, and
Acts 6:5 and others went the same way. Re-harvested with the anchor treated as
a verse START rather than a whole verse (scripts/scrape-wujek.py), which also
had to cope with four anchor shapes across books, hidden page markers opening
paragraphs mid-verse, chapter ids that are simply wrong (Mark labels 87
anchors "15:*" across chapters 14-16), and Psalms heading its divisions
"Psalm CXVII" where every other book says "Rozdzial N".

Structure was then repaired against the Clementine Vulgate, by hand where a
rule would have guessed:

  * The Acts 5 interleaved three streams -- Acts 4:27-37 duplicated, the real
    5:1-37, and 5:38-39 mislabelled 28-29. Verified against the Latin, the
    duplicates dropped, the two renumbered. Acts 4:27 kept the site's own
    wording in place of a 1962 reading, removing a seam.
  * The site heads psalms by Hebrew division; its "Psalm 114" is two psalms.
    Split into Vulgate 113:9-26 and 114:1-9, confirmed at all four boundaries.
  * A corrupt anchor for 1 Chronicles 9:11 pushed 34 paragraphs of genealogy
    into chapter 10. Counts corroborate: 10 + 34 = 44 verses, remainder 14,
    both exactly the Vulgate's.
  * 16 chapters looked short at the end; only two verses were truly absent.
    The rest were merges, split at anchors read off the Latin -- including one
    across a chapter boundary (Colossians 4:1 sat inside 3:25) and four
    numbering offsets where a mid-chapter merge shifted everything after it.

wuj now carries exactly the Vulgate's verse set: 35810 rows, every one of the
1334 chapters matching, no duplicates, no gaps. corpus-check: 0 warnings.

The ~300 verses that could not come from biblia.info.pl are recorded in
NOTICE; they are not public domain.
</content>
</entry>
<entry>
<title>bible: embed only the Vulgate by default; other corpora opt-in</title>
<updated>2026-07-28T16:01:33Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-07-28T16:01:33Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/lectio.git/commit/?id=5e1664b8008675177a551aff1de371f54d4e20e4'/>
<id>urn:sha1:5e1664b8008675177a551aff1de371f54d4e20e4</id>
<content type='text'>
Split the bundled corpora so the default binary carries only the
public-domain Latin Vulgate. wuj/drb/grb move to corpora_optional/ and are
compiled in only with `-tags fullbible`, or dropped into the user corpora
dir at runtime. This keeps the distributed AGPL binary free of third-party
scripture text -- bible corpora are not covered by lectio's licence.

- bible: read corpus files/metadata from core + optional embed FS
  (embReadCorpus/embCorpusFiles); optionalFS is set only under the tag.
- config: default `versions` now reflects corpora actually available in the
  build (Vulgate first), so a vul-only binary offers no absent versions.
- Makefile: `make build` = vul only; `make build-full` embeds the rest;
  `make test` runs with -tags fullbible; check-corpora validates both dirs.
- tests: corpusmeta asserts on vul (the always-embedded corpus); the
  versions-default expectation follows availableVersions().
</content>
</entry>
</feed>
