| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
~3,800 verses carried their own verse number at the start of the text (a
biblia.info.pl harvest artifact concentrated in the deuterocanonical books --
Sirach, 1-2 Maccabees, Wisdom, Judith, Baruch, at 90-98% of each, plus 3 strays
in Ezra/Proverbs). E.g. Sirach 4:4 read "4 Nie odrzucaj..." instead of "Nie
odrzucaj...". It passed every check: verse counts are unaffected, and the
verses aren't short.
Strip a leading token only where it exactly equals the verse number (so a
genuine "40 dni" in a non-matching verse is untouched), and only when text
follows. 3797 verses fixed; 35,810-row parity preserved, no verse emptied, no
new warnings; --corpus-check wuj still clean; make test green.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
grb failed corpus-check outright: 15 errors, and two whole books absent from
lectio at runtime with no error and no warning.
Upstream grb is a standalone Septuagint reader carrying 87 books. lectio needs
one uniform canon, so scripts/gen-grb-lectio.py derives the 73 books vul has
rather than forking -- upstream stays intact and keeps everything.
Eleven books have no Vulgate counterpart and are dropped, including 2 Esdras,
which upstream's README explains IS Ezra + Nehemiah under the Septuagint's
name; verified byte-identical, so nothing is lost. Five books ARE canonical
under other names and are remapped, and the verse counts show which witness
the Vulgate follows -- Jerome translated Theodotion, not the Old Greek:
Bel and the Dragon (Theodotion) 42 = vul Daniel 14 (42) exact
Bel and the Dragon (LXX) 37 no
Sussana (Theodotion) 64 = vul Daniel 13 (65)
Sussana (LXX) 37 no
Letter of Jeremiah 73 = vul Baruch 6 (72)
Wisdom of Solomon 435 = vul Wisdom (439)
Upstream's plain "Daniel" is the Old Greek and is missing chapter 4 outright,
so Theodotion supplies Daniel throughout: complete, and the tradition the
lectionary actually cites. lectio thereby GAINS Wisdom, Daniel 4, Daniel 13-14
and Baruch 6 in Greek rather than losing them.
Three upstream defects are repaired in transit:
* 408 rows carried 5 fields, not 6 -- a lost tab fused the book number and
chapter ("281" = book 28, chapter 1). Every one was in Hosea or Zechariah,
and neither book had a single well-formed row, so bible.go skipped both
entirely. Hosea (197 verses) and Zechariah (211) are back.
* 5 merged verse labels ("27-28") that strconv.Atoi turns into verse 0.
* A UTF-8 BOM welded to the first book name, making "Genesis" a phantom
74th book matching nothing.
The SBLGNT apparatus sigla are stripped, as TestGrbNoApparatusMarkers
requires; regenerating from raw upstream reintroduces ~8700 of them.
grb now passes. The remaining warnings are the Septuagint being itself --
Jeremiah is LXX-numbered (grb 33:2 is the Vulgate's 26:2), Esther integrates
its additions into chapters 1-10, LXX Malachi has three chapters, and 3
Kingdoms carries supplements like 10:22a that upstream stores as a duplicate
verse 22. Silencing those would mean deleting real Greek text.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Eight corpus-check warnings, of two kinds.
Three chapters (Exodus 40, Genesis 49, Song of Solomon 1) held exactly as many
verses as the Vulgate but numbered with a gap. The count pins the mapping, so
renumbering 1..N is safe and changes no text.
Five chapters were genuinely short. The missing verses came from get.bible's
douayrheims -- the same public-domain source and API scripts/gen-deutero.py
already uses for this corpus, so no new provenance is introduced.
Scope was deliberately narrowed to the flagged chapters. A first pass over
every chapter shorter than the Vulgate pulled in 938 verses, 820 of them in
Psalms -- but drb.ini declares psalm_system = drb, a different numbering, so
matching by Vulgate verse number there would have inserted the wrong text
under the wrong numbers across the Psalter.
One warning remains and is not a defect: Baruch 6 runs 1-72 with verse 37
absent in drb and in get.bible alike, a genuine Douay/Vulgate versification
difference faithfully represented.
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
The Wujek corpus was missing text and mis-numbering what it had: 453
corpus-check warnings, 8 absent chapters, 84 duplicate rows, and 357 chapters
holding fewer verses than the Vulgate.
Most of it was never a sourcing problem. The original harvest of
biblia.info.pl assumed one verse is one <p>, but many psalms print the
superscription in the anchored paragraph and the psalm's opening line in an
unanchored one behind a drop cap -- so 44 psalms lost their first line, and
Acts 6:5 and others went the same way. Re-harvested with the anchor treated as
a verse START rather than a whole verse (scripts/scrape-wujek.py), which also
had to cope with four anchor shapes across books, hidden page markers opening
paragraphs mid-verse, chapter ids that are simply wrong (Mark labels 87
anchors "15:*" across chapters 14-16), and Psalms heading its divisions
"Psalm CXVII" where every other book says "Rozdzial N".
Structure was then repaired against the Clementine Vulgate, by hand where a
rule would have guessed:
* The Acts 5 interleaved three streams -- Acts 4:27-37 duplicated, the real
5:1-37, and 5:38-39 mislabelled 28-29. Verified against the Latin, the
duplicates dropped, the two renumbered. Acts 4:27 kept the site's own
wording in place of a 1962 reading, removing a seam.
* The site heads psalms by Hebrew division; its "Psalm 114" is two psalms.
Split into Vulgate 113:9-26 and 114:1-9, confirmed at all four boundaries.
* A corrupt anchor for 1 Chronicles 9:11 pushed 34 paragraphs of genealogy
into chapter 10. Counts corroborate: 10 + 34 = 44 verses, remainder 14,
both exactly the Vulgate's.
* 16 chapters looked short at the end; only two verses were truly absent.
The rest were merges, split at anchors read off the Latin -- including one
across a chapter boundary (Colossians 4:1 sat inside 3:25) and four
numbering offsets where a mid-chapter merge shifted everything after it.
wuj now carries exactly the Vulgate's verse set: 35810 rows, every one of the
1334 chapters matching, no duplicates, no gaps. corpus-check: 0 warnings.
The ~300 verses that could not come from biblia.info.pl are recorded in
NOTICE; they are not public domain.
|
|
|
Split the bundled corpora so the default binary carries only the
public-domain Latin Vulgate. wuj/drb/grb move to corpora_optional/ and are
compiled in only with `-tags fullbible`, or dropped into the user corpora
dir at runtime. This keeps the distributed AGPL binary free of third-party
scripture text -- bible corpora are not covered by lectio's licence.
- bible: read corpus files/metadata from core + optional embed FS
(embReadCorpus/embCorpusFiles); optionalFS is set only under the tag.
- config: default `versions` now reflects corpora actually available in the
build (Vulgate first), so a vul-only binary offers no absent versions.
- Makefile: `make build` = vul only; `make build-full` embeds the rest;
`make test` runs with -tags fullbible; check-corpora validates both dirs.
- tests: corpusmeta asserts on vul (the always-embedded corpus); the
versions-default expectation follows availableVersions().
|