The orphan-id retraction ceremony
In this section
A brain's metadata index keeps an id table: one entry per row it has indexed, nouns and relations alike, mapping the row's id to the small integer every posting list and column is keyed by. When a row is deleted, the index retracts it, and the table entry goes last.
12.13.0's removeMany() skipped that retraction for every relation it removed as a side effect of deleting a noun. The relation left canonical and the graph correctly, but its id-table entry, its postings, its column values and its share of the posted-relation tally stayed behind with nothing behind them. No read returns a deleted relation because of it, since every read loads the row and drops what is not there. What it does cost:
checkHealth().entityCountreads too high by the number of leaked entries.The posted-relation coverage row over-reads.
Negation filters (
ne,not,exists: false, …) count the leaked numbers in their "every row" set, so a page can come back shorter than its limit.
Nothing that runs by itself removes them.
The id-table consistency sweep checks the table's two directions against each other, and a leaked entry has both.
A metadata rebuild re-derives postings from canonical and keeps the table, because the table is identity.
The coverage census that does retract gone rows reads nouns only.
The fix in 13.0.0 stops new leaks. This ceremony removes the old ones.
const dry = await brain.retractOrphanedIndexIds()
console.log(dry.summary)
// metadata-id-orphan-retraction finished: found 5515 orphaned index id(s) an
// apply would retract (dry run: nothing changed); ...
const done = await brain.retractOrphanedIndexIds({ apply: true })
console.log(done.census.retracted) // 5515
// later, from any process:
const receipt = await brain.orphanedIndexIdsStatus()TSWhat it guarantees
Nothing changes without your word. The default is a dry run. It performs the same walk and the same locked re-check as an apply, and reports
retractable: exactly what an apply would retract at that moment.Nothing written after it starts is examined. The run pins the id table's next integer when it begins. Integers are never reused, so every row minted afterwards sits above the pin.
Nothing real is retracted. Every candidate is re-checked in canonical and retracted inside one hold of the commit lock, so no write can land between the check and the retraction. A candidate that canonical gained since the walk is counted as
revivedand left alone.Nothing in flight is retracted. An id the table maps but the index never indexed has no committed index record. That is a relation endpoint resolved before its commit lands, or an entry older than the index's reverse records. It is counted as
unindexedSkipped, never retracted.Live rows keep their numbers. Only unmarked, still-absent ids are touched, and removing an entry never renumbers another.
It runs behind the doors. It is a maintenance-class background job,
metadata-id-orphan-retraction. It yields to foreground traffic and stands down atclose(). What it retracted before a stand-down is flushed, and the receipt records how far it got.It resumes. A run a killed process left unfinished continues on the next writer open, with the pin and the mode it started with. Retraction is idempotent: an id the table no longer maps is never a candidate again, so a crash costs at most a repeat, never a double.
A second run finds nothing.
What it reports, by class
field | meaning |
|---|---|
| rows the walk read from canonical |
| integers below the pin that the table still mapped |
| mapped integers canonical did not name at the walk |
| candidates with no committed index record — left alone |
| candidates canonical held by the time of the locked re-check |
| candidates still absent under the lock — what an apply retracts |
| retracted by this run; always |
The receipt also carries mode (dry-run or apply) and phase (marking, sweeping or done). It records highWater (the pin), cursor (the integer the sweep has reached), the start and finish times, whether the run resumed, and a one-sentence summary. It lives in the store at _cor_orphan_id_retraction/receipt.json.
What it costs
PROJECTED, not measured. None of these figures has been measured yet. The receipt from the first run on a real brain replaces them.
There are three terms:
The walk. One streaming pass over canonical ids, nouns then relations, through the native canonical reader. It is the same enumeration every rebuild and census makes. The bitmap it fills costs one bit per integer below the pin: about 35 KB at 282,000 entries.
The sweep. One id-table lookup per integer below the pin. The table's own per-entry walk is about 0.8 µs an entry (
docs/open-law.md§2b), which is about a quarter of a second at 282,000.The candidates. For each one: an index-record probe, then two canonical reads and one index retraction inside a short commit-lock hold. A few thousand candidates take seconds, not minutes. The first metadata-less retraction loads every known field index once, and later ones reuse them.
The index is flushed once per 4,096 integers swept, and only when that stretch retracted something. The receipt is written at the same points.
When to run it
Run it once on any brain that ran removeMany() on 12.13.0 against rows with relations. The dry run tells you whether there is anything to do. retractable: 0 means the brain was never touched.
It needs the native engine and a writer. A read-only brain refuses by name.