<feed xmlns='http://www.w3.org/2005/Atom'>
<title>krino.git/internal, branch v0.0.12</title>
<subtitle>rule-based file sorter in Go, shows the plan before touching anything, applies only the files you choose, and can undo any run, with s-expression rules and a GTK 4 window</subtitle>
<id>https://git.labunix.xyz/krino.git/atom?h=v0.0.12</id>
<link rel='self' href='https://git.labunix.xyz/krino.git/atom?h=v0.0.12'/>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/'/>
<updated>2026-09-17T12:17:54Z</updated>
<entry>
<title>a file's name is folded once, not once per name test</title>
<updated>2026-09-17T12:17:54Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T12:17:54Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=b66850cc8c584cc00bcd796eb3aa238bcf87394a'/>
<id>urn:sha1:b66850cc8c584cc00bcd796eb3aa238bcf87394a</id>
<content type='text'>
Folding is the expensive half of a name test on a name with diacritics,
and every name test of every rule folded the same name again: on Polish
names it was most of the matching work. The per-file facts memoise it,
which is where one file's work belongs - the object is per file and per
goroutine, so no lock.

  4000 Polish names, twelve rules with name tests: 0.33s -&gt; 0.13s
</content>
</entry>
<entry>
<title>max-read bounds what a file becomes, not only what is read</title>
<updated>2026-09-17T12:07:37Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T12:07:37Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=4c6fadfab5434317357ad0272ace7927d2945942'/>
<id>urn:sha1:4c6fadfab5434317357ad0272ace7927d2945942</id>
<content type='text'>
max-read gates on file size before reading, then the text was read
whole and decoded: a file that is not valid UTF-8 decodes one byte per
code point and doubles, and normalising and folding copy that again per
set of options, with GOMAXPROCS files in flight. Twelve 40 MB files
reached 3.3 GB - enough to put a laptop into the OOM killer, with no
attacker involved, just a few big .log or .csv files.

Two bounds. The decoded text is cut to max-read at a rune boundary, so
the ceiling means what a reader takes it to mean. And extraction of
files at or above 4 MiB is rationed to two at a time, since holding
several large texts at once is what multiplies the ceiling; smaller
files, which is nearly all of them, are untouched.

  twelve 40 MB files: peak RSS 3294 MB -&gt; 728 MB, wall 27s -&gt; 46s
  two thousand small files: 0.05s both ways

The wall-clock cost falls entirely on large files needing extraction,
and buys a program that finishes instead of being killed.
</content>
</entry>
<entry>
<title>{now} is the start of the run, as the spec always said</title>
<updated>2026-09-17T12:03:44Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T12:03:44Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=33d771438f1da363af7c9ff2d5f4c368c326e84d'/>
<id>urn:sha1:33d771438f1da363af7c9ff2d5f4c368c326e84d</id>
<content type='text'>
The clock was read once per directory, so a run of several directories
stamped several different times - and a review that took a few seconds
could put one run's files into two folders, or at midnight two dates.
The spec (§7.3) says "the start of the run"; krino.conf(5) documented
the behaviour rather than the intent, so the two contradicted each
other.

The session now carries the run's clock and every directory plans with
it. Engine.Plan keeps its meaning for a caller with no session of its
own; Session.Plan passes the run's own start.
</content>
</entry>
<entry>
<title>a symlink in the sorted directory no longer redirects a step</title>
<updated>2026-09-17T12:02:24Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T12:02:24Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=a65e8a0587d3e5c971bef3062dd0d8a3b5cb53b7'/>
<id>urn:sha1:a65e8a0587d3e5c971bef3062dd0d8a3b5cb53b7</id>
<content type='text'>
Placeholders were already stopped from sending a file out of the
directory a rule named. A symlink is a name too, and the directory
krino sorts is by the threat model's own premise a place the internet
writes into: a link named after a rule's destination sent moves and
copies anywhere, and under (on-conflict overwrite) trashed a file
OUTSIDE the sorted directory - while the plan showed the in-tree text
and the run reported success.

A step whose destination passes through a symlink at or below the
directory being sorted now fails. A destination the configuration names
outside it - ~/docs on another disk - is the user's own arrangement and
is followed as before; both cases have a test.

End to end, the review's scenario (Out -&gt; ~/secret, overwrite):

  before: 1 applied, the user's file replaced and trashed
  after:  0 applied 1 failed, the file untouched, nothing trashed

Two bookkeeping bugs in MkdirAllTracked went with it: a dangling
symlink read as a missing directory and was then recorded as one krino
had created - undo would have unlinked a link krino never made - and a
directory created by someone else between the check and the mkdir was
recorded the same way.
</content>
</entry>
<entry>
<title>a duplicate test in an exclude protects like one in a rule</title>
<updated>2026-09-17T11:59:52Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:59:52Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=b44222fc2b061382dc601014cde287cc41b857ca'/>
<id>urn:sha1:b44222fc2b061382dc601014cde287cc41b857ca</id>
<content type='text'>
The scopes that drive "no rule deletes a file krino found to be a
duplicate" were collected from rules only. A directory whose duplicate
tests lived in (exclude ...) forms had no scopes at all, so the
protection never engaged: krino explain said "yes duplicate" and the
next rule permanently deleted every copy. The README's promise was
false in that shape, and the spec's wording permitted it.

What matters is what krino knows, not which form taught it. The spec
and krino.conf(5) now say so too.
</content>
</entry>
<entry>
<title>a chain that deletes itself still gives back the file it displaced</title>
<updated>2026-09-17T11:58:25Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:58:25Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=9ab6686b98499c745024a474a71e3d99b6e14973'/>
<id>urn:sha1:9ab6686b98499c745024a474a71e3d99b6e14973</id>
<content type='text'>
(on-conflict overwrite) trashes the file in the way; §7.4 promises undo
restores it. Walking a file's log entries stopped dead at a permanent
delete, so the displace written earlier in the same chain was never
reached: the user's file stayed in the Trash, the refusal named only
the file they did not care about, and krino log called the run undone.

The displaced file is a different file, so it is offered as its own
entry in the undo plan, keyed by its own path - the deleted file stays
refused, since nothing of it can come back, and the copy or move that
preceded the delete stays unreversed too (undoing a copy whose original
was then deleted would destroy the last remaining copy).

The accounting matched: every reversible step of a deleted file was
subtracted, its displace included, so the run read (undone). Only what
genuinely cannot come back is subtracted now.

End to end, the scenario from the review: the only copy of a file is
displaced by an incoming one that is then permanently deleted.

  before: archive/ empty, "(undone)", nothing offered
  after:  archive/a.pdf restored, run reads partly undone
</content>
</entry>
<entry>
<title>a file trashed to make room is logged the moment it is trashed</title>
<updated>2026-09-17T11:53:01Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:53:01Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=a80d470cbad7335fd5e534f32d935e5a49aadb24'/>
<id>urn:sha1:a80d470cbad7335fd5e534f32d935e5a49aadb24</id>
<content type='text'>
The displace was carried out first and logged only when the whole step
finished - for a copy or a cross-device move, the entire data transfer
later. A process killed in that window left the user's file in the
Trash with nothing recording it: krino log said "nothing applied" and
undo offered nothing.

It is its own action with its own line (spec §9), so it is now written
the moment trash.Put returns, through a hook ChainLogged calls before
the step that needed the name begins. A displace that cannot be logged
fails the step rather than compounding an unrecorded destructive act
with a second one.

Verified by killing krino -9 mid-copy with 600 MB in flight:

  before: krino log "nothing applied", 0 displace lines
  after:  krino log "1 displaced",     1 displace line
</content>
</entry>
<entry>
<title>the suffix search continues instead of starting again at _1</title>
<updated>2026-09-17T11:50:01Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:50:01Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=78d8313791f05defc9e0a9f2bad8e9710f741a60'/>
<id>urn:sha1:78d8313791f05defc9e0a9f2bad8e9710f741a60</id>
<content type='text'>
Every file renamed onto one name probed stem_1, stem_2, ... from the
beginning, so N files cost N^2/2 Exists calls - a test here counts 1890
of them for 60 files. The search now continues from the highest suffix
already tried for that stem.

Within one plan that is the same answer: the taken set only grows while
a plan is built and the disk is not being written to, so a suffix taken
once stays taken. Proved rather than argued - with same_1 and same_3
already on disk and same_2 free, both versions put a file in the gap,
and the two plans are byte-identical.

  1500 files renamed to one name: 2.63s -&gt; 0.05s
</content>
</entry>
<entry>
<title>a content class is worked out once, not once per copy</title>
<updated>2026-09-17T11:47:35Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:47:35Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=b596085d2391ce3701fa6820a6728ba634ac453b'/>
<id>urn:sha1:b596085d2391ce3701fa6820a6728ba634ac453b</id>
<content type='text'>
Lookup walked the whole size group for every file in it, so N copies of
one file cost N walks of an N-member group, each taking the index's
lock at every step - which is why more workers made it slower rather
than faster. Every member of a class elects the same original (the
invariant identicalTo already documents and a test already pins), so
the class is memoised on the first walk and every later member is a
map lookup.

Measured over identical files, invented data, same machine:

  500 files   0.17s -&gt; 0.12s
  2000 files  1.89s -&gt; 0.58s

and the plans are byte-identical once the sandbox path is normalised.
The test counts walks: one per content class, not one per file.
</content>
</entry>
<entry>
<title>the directory lock is the kernel's, not a pid we believe</title>
<updated>2026-09-17T11:34:20Z</updated>
<author>
<name>Lukasz Kasprzak</name>
<email>lukas@labunix.xyz</email>
</author>
<published>2026-09-17T11:34:20Z</published>
<link rel='alternate' type='text/html' href='https://git.labunix.xyz/krino.git/commit/?id=6971543d4749574d4ca575c4e8acf04f9e86d6bb'/>
<id>urn:sha1:6971543d4749574d4ca575c4e8acf04f9e86d6bb</id>
<content type='text'>
flock(2) on the lock file's descriptor replaces "write my pid, and
decide whether the pid in the file is still alive". The kernel drops
the lock when the process ends, however it ends, so there is no stale
krino lock to detect and no takeover to race over.

What that fixes:

- Two runs that both judged a lock stale could remove and recreate it
  and both believe they held it. Remove-then-create cannot be made
  atomic; there is nothing to make atomic now. Pinned by a test with
  eight callers over twenty rounds.
- A pid reused after a crash made the lock live for ever, and the
  message named neither the file nor the pid, so there was nothing to
  act on. The message now names both.
- Signal(0) reads EPERM as "not running", so a lock held by another
  user was taken over. There is no such judgement left to get wrong.

A run that waits for a held lock now says so first. Waiting is what
the spec asks for, but the wait has no timeout, and in silence it is
indistinguishable from a hang - I spent two minutes on one myself
today, waiting on a lock the window was holding.

Tests that faked a held lock by writing a file now hold a real one.
</content>
</entry>
</feed>
