aboutsummaryrefslogtreecommitdiff
path: root/test/test_citation.ml
diff options
context:
space:
mode:
authorLukasz Kasprzak <lukas@labunix.xyz>2026-08-20 22:11:17 +0200
committerLukasz Kasprzak <lukas@labunix.xyz>2026-08-20 22:11:17 +0200
commit1988d350242b47aa52aa07904c495e7e2c0eba82 (patch)
tree463254012c555da51aa976dbe350415070b80a63 /test/test_citation.ml
parent2d8942e08edcbe5430f97270bdb1233e287600ee (diff)
downloadcolitur-1988d350242b47aa52aa07904c495e7e2c0eba82.tar.gz
colitur-1988d350242b47aa52aa07904c495e7e2c0eba82.zip
fix: audit findings — parser strictness, name ambiguity, and errors
Found by auditing the shipped program rather than the diff. The parser accepted OCaml integer-literal syntax, so "Luke 1_1:5" read as chapter ELEVEN and "+5" as 5 -- a typo silently becoming a different chapter, reachable through any user overlay. Numbers are now plain digits and positive, and a descending range is rejected: 1:20-10 is always a transcription error. No shipped citation changed. FOUR PAIRS OF DIFFERENT BOOKS SHARED A FULL TITLE. 1 and 2 Corinthians both rendered "Epistola ad Corinthios", as did Thessalonians, Timothy and Peter -- 108 citations in 2027 alone that a reader cannot resolve to a book. This is the Kings defect fixed earlier and not generalised. The titles now carry their volume numeral, marked CONSTRUCTED, and a test asserts no two books share a name -- while allowing the case where two ids ARE the same book under different numbering, which a tradition relates. Spec section 8.5 is now delivered rather than merely recorded. Shipped styles did not re-parse their own output: 32 of 52 Latin abbreviations and 49 of 52 full titles failed, so a citation copied from colitur's own output into an overlay was passed through untouched and printed in the wrong language, silently. Every shipped name is registered as a spelling and split_book learned multi-word titles by longest-token match. Now 0 of 52 fail beyond the same-book aliases. Overlay errors were written for a compiler author: they named an OCaml source file the reader does not have and buried the useful token. The existing five-path rewriter is replaced by a generic one, applied to every load path rather than one, so "rank: is not one of the allowed values (at Class9)" replaces the raw Of_sexp_error dump. Also: the new-overlay scaffold documented citations and layer without showing them, and its comment implied the wrong nesting -- the single easiest thing to get wrong; error messages echoed whole file lines, copying an unrelated file's contents into stderr when a flag pointed at one; and config --show validated partway down its table, exiting 2 after writing five rows to stdout.
Diffstat (limited to 'test/test_citation.ml')
-rw-r--r--test/test_citation.ml19
1 files changed, 19 insertions, 0 deletions
diff --git a/test/test_citation.ml b/test/test_citation.ml
index 133553b..dae4610 100644
--- a/test/test_citation.ml
+++ b/test/test_citation.ml
@@ -283,6 +283,11 @@ let parses input expected () =
| Error e -> Alcotest.failf "%s did not parse: %s" input e
| Ok t -> Alcotest.(check s) input expected (show t)
+let rejects input () =
+ match P.parse input with
+ | Ok t -> Alcotest.failf "%s should not parse, got %s" input (show t)
+ | Error _ -> ()
+
let test_rejects_unknown_book () =
Alcotest.(check bool) "error" true (Result.is_error (P.parse "Nonesuch 1:1"))
@@ -321,6 +326,20 @@ let parse_suite =
("ordinal 3 parses", `Quick, parses "3 Kings 17:8-16" "kings_3|17:8-16");
("ordinal 3 dotted", `Quick, parses "3 Kgs. 19:3-8" "kings_3|19:3-8");
("ordinal 4 parses", `Quick, parses "4 Kings 5:1-15" "kings_4|5:1-15");
+ (* A citation number must be plain digits and positive. OCaml's
+ [int_of_string] also accepts its own literal syntax, so "1_1" parsed
+ as chapter ELEVEN and "+5" as 5 -- a typo silently becoming a
+ DIFFERENT chapter, which no amount of downstream care can catch.
+ Reachable through a user overlay, which supplies arbitrary strings. *)
+ ("rejects underscore in chapter", `Quick, rejects "Luke 1_1:5");
+ ("rejects underscore in verse", `Quick, rejects "Luke 1:5_0");
+ ("rejects plus in chapter", `Quick, rejects "Luke +1:5");
+ ("rejects plus in verse", `Quick, rejects "Luke 1:+5");
+ ("rejects chapter zero", `Quick, rejects "Luke 0:5");
+ ("rejects verse zero", `Quick, rejects "Luke 1:0");
+ (* A descending range is always a transcription error. *)
+ ("rejects descending range", `Quick, rejects "Luke 1:20-10");
+ ("allows a single-verse range", `Quick, parses "Luke 1:5-5" "luke|1:5-5");
("unknown book", `Quick, test_rejects_unknown_book);
("garbage", `Quick, test_rejects_garbage) ]