aboutsummaryrefslogtreecommitdiff
path: root/lib/render/escape.mli
Commit message (Collapse)AuthorAgeFilesLines
* feat(render): Typst as a seventh escaping flavourLukasz Kasprzak2026-08-201-5/+5
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Adds Escape.Typst: type, all, to_string ("typst"), of_string, of_extension (.typ), and a per-character expand function, structured identically to the existing latex escaper (one pass, no re-scan, so double-escaping stays impossible by construction). The metacharacter set was verified against the installed typst 0.14.2 binary, not assumed: a probe document escaping each of #, *, _, $, @, <, >, `, \, ~ and - was compiled and pdftotext'd back to confirm the literal character survives, and each was separately confirmed to do something else when left bare (# opens code mode, */_ toggle strong/emph, $ opens math, @ opens a reference -- a bare unresolved @word is a hard compile error, not merely mangled output -- </> can close around a bare word into label syntax that swallows it whole, ` opens raw, ~ is a non-breaking space, and a run of two or three '-' becomes an en/em dash). All ten are backslash-escapable; none needed a non-backslash workaround. '-' is escaped unconditionally rather than only inside a detected run, since this escaper has no lookahead -- a probe confirmed escaping every hyphen independently still typesets as literal hyphens for a run of any length, so the single per-character rule is sufficient. test_escape.ml's new Typst cases were written first and shown to fail against a stubbed identity apply before the real escaper landed, per this project's own regression-test discipline. Comments in both files mark which new tests are genuine regression tests (the per-character escaping, including a real shipped citation and a synthetic markdown-habit overlay name) versus characterisation (to_string/ of_string/of_extension are flat table lookups with no logic to have been wrong).
* feat(render): per-flavour escaping and RFC 5545 line foldingLukasz Kasprzak2026-08-191-0/+28
Six flavours: latex, groff, html, xml, ics, none. Markdown, AsciiDoc and plain text map to none deliberately -- their metacharacters are context-dependent and escaping them aggressively produces worse output than not escaping. An unrecognised extension returns None rather than falling back to none: guessing the flavour wrong produces malformed output that looks fine until it does not. Folding backs off to a non-continuation byte, so a fold never splits a UTF-8 sequence -- the failure mode that would corrupt Polish and Latin names in a published feed.