Brainy Brainy
Docs Brainy

The rerank stage

In this section
// The fixed rule: topK 40, budgetMs 300.
const rows = await brain.find({
  query: 'what did we decide about the embedding spaces',
  limit: 20
})TS

The top K rows of the candidate set are re-scored by a cross-encoder that reads the query and each row together, and returned re-ordered — each scored row carrying rerankScore beside the score it already had.

A boolean since 11.4. The always-on stage is now late interaction, which orders the whole candidate pool from token vectors stored at index time; this stage is a SECOND, dearer pass over the page that ordering returned. rerank: true asks for it, omitted and false do not, and there are no numbers to pass — topK and the 300 ms budget are the fixed rule below, and the model comes from the store's stamp. An object is REFUSED by name: a call that passed { topK: 64 } and silently got the fixed rule would be measuring a number the engine never used.

The rule: hardware never changes the answer, only the wait

The numbers a default rerank uses are a function of the request, and of nothing else:

topK     = min(candidateCount, max(32, 2 x limit)), capped at 128
budgetMs = 300

No core count, no measured throughput, no adaptive K, no "score fewer rows on a busy box". Those would all be the same defect wearing different clothes: the same store answering the same question two different ways depending on where it ran. A slow host waits longer, and if it cannot finish inside the budget it refuses by name with its numbers. It never serves a different order.

limit

topK

why

5

32

the floor — a window narrower than that cannot change an answer worth changing

20

40

2 x limit — to change the top 20, the model has to read rows 21..40

100

128

the ceiling — cost is linear in K, so the rule is bounded

20, on a 7-row page

7

narrowed to the candidate set; the stamp reports facts

2 x limit is the whole shape of it: a re-order that cannot reach past the page it is re-ordering cannot promote anything into that page.

The constants are exported — RERANK_DEFAULT_TOPK_FLOOR, RERANK_DEFAULT_TOPK_CEILING, RERANK_DEFAULT_BUDGET_MS — and so is the rule itself, defaultRerankParams(limit, candidateCount), so a caller can compute what a call will do before making it.

The switch, and the receipt that flips it

RERANK_DEFAULT_ON in src/rerank/nativeRerankStage.ts is a single build-time constant, and it currently ships false. The rule, the stamp, the cache and the refusals are all built and pinned; what is not yet true is the millisecond law. What it costs is the reason: the fp32 L12 forward pass is measured at ~24 ms per pair on 30-word rows and ~130 ms on 120-word rows, so the rule's 32-pair floor is 8x to 40x over 300 ms. Turning it on today would convert almost every find({ query }) into a RerankBudgetExceededError.

The receipt that flips it: the int8 cross-encoder path measured at a per-pair wall that meets the 300 ms budget at K=32 on the fleet's own hosts — that is, 300 / 32 >= the measured p95 ms per pair, at the row lengths the fleet's brains actually hold, on the slowest host that serves. When that measurement exists, the constant becomes true and nothing else changes.

This receipt is not the int8 path — see Precision: the int8 path, and why it is not the default: re-measured against the checkpoint this package ships today, weight-only Q8_0 is 9.6x to 14.3x SLOWER than fp32 on this CPU kernel shape, not faster, so it closes this lever rather than opening it. The lever that remains open is a kernel that quantizes the activations too, or a non-CPU device — neither is built.

Opting out

await brain.find({ query, limit: 20, rerank: false })TS

false and omitted both mean the same thing since 11.4 — no second pass — and both are honoured before anything else: no provider lookup, no stamp read, no model. What find() returns is the RANKED order (late interaction over the whole pool), which is not the same as the order before any of this existed.

Why a second model at all

The vector leg scores a query and a row separately: two forward passes that never see each other, compared by cosine distance. That separation is what makes it fast enough to walk a corpus, and it is also why the top of its list is only approximately the top of the list. A row can sit close to the query in embedding space and answer a different question.

A cross-encoder reads the pair as one sequence — [CLS] query [SEP] row [SEP] — and emits a single relevance logit. It cannot be pre-computed and it cannot be indexed, because the score does not exist until the query arrives. So it is only ever run over a short list the cheap leg already produced. That is exactly this stage: find() builds the candidate set, and its head is re-scored.

Where it runs, and why that is the point

On the engine's threads. The forward pass is a napi AsyncTask on libuv's threadpool, so find(), get() and every write keep being answered while the batch computes.

This matters because the alternative was measured. A consuming service ran this same model on its own JS thread and paid 0.7–1.4 s per recall for it, in a process that also served every other door. Moving the model into the engine buys nothing if the engine then runs it on the caller's thread anyway — so that is pinned, not asserted: src/rerank/rerankEventLoop.e2e.test.ts runs the identical batch through the blocking door and through the async one while sampling event-loop lag, and requires the async arm's p95 lag to be a small fraction of the blocking arm's. The blocking arm is the control: the suite fails if the probe does not catch it stalling the thread, because a probe that cannot detect a known stall proves nothing about the door that matters.

The model

As of 11.3, cross-encoder/ms-marco-MiniLM-L6-v2 is the DEFAULT — Apache-2.0, a BertForSequenceClassification with one label and an identity output activation, 6 layers at hidden 384. Through 11.2.5 the default was its 12-layer sibling, cross-encoder/ms-marco-MiniLM-L12-v2; see The 11.3 reversal, measured for why and the measurements behind it. L12 stays vendored in this repo's source and reachable by an explicit variant — see Loading L12 explicitly — but from 11.4.0 it no longer ships inside the @soulcraft/brainy tarball; see Models: what ships inside, what is optional below.

L6 ships inside the package: safetensors, tokenizer and config under dist/assets/models/ms-marco-MiniLM-L-6-v2, with the Apache-2.0 text and a PROVENANCE.md beside it. Nothing is downloaded at runtime and no scoring leaves your machine. The weights are memory-mapped by candle, so a box running many brainy processes holds one copy of the loaded file's pages, not one per process.

Models: what ships inside, what is optional

Three models ship inside @soulcraft/brainy (~305 MB together), always, on every platform, with zero configuration:

model

role

size

all-MiniLM-L6-v2

the embedder

~88 MB

cross-encoder/ms-marco-MiniLM-L6-v2

this page's DEFAULT rerank stage

~88 MB

answerai-colbert-small-v1

late interaction (see What is coming)

~129 MB

One model is optional, as its own npm package — @soulcraft/brainy-model-ms-marco-minilm-l12-v2, the 12-layer cross-encoder. It left the tarball at 11.4.0: with the late-interaction token encoder aboard, the package would exceed what the npm client can upload in one request (~400 MB). Install it to load L12 explicitly:

npm install @soulcraft/brainy-model-ms-marco-minilm-l12-v2BASH

Once installed, resolveCrossEncoderAssetsDir('L12') finds it — see Loading L12 explicitly. Not installed, and the in-package source tree absent too (an ordinary npm install, rather than this repo's own checkout): the refusal names the package and this exact install line, never a network fetch of its own. Every optional package (this model, the embed-service and brainy-serve binaries — see the CLI doc) is an optionalDependency, never a dependency — install none of them and lose nothing else.

Air-gapped hosts use npx brainy bundle --out <dir> instead of installing packages one at a time — it resolves every optional package this install's package.json pins (including this model, when pinned) and fetches their exact tarballs once, from a machine with registry access, for an offline install elsewhere. See the CLI doc for the full ceremony.

Both checkpoints sit under assets/models/ at the package root rather than beside the embedder's weights in engine/assets/models/, and that is deliberate: src/engine/ is the vendored reference engine, carrying its own provenance manifest and per-file digests, and its asset directory is checked for integrity by npm run vendor:verify — which exits non-zero on a model-asset problem there. A model the reference engine never had does not belong inside a record of what the reference engine shipped.

L6 and L12's SOURCE directories are identical in shape (weights, tokenizer, config, licence, provenance) and both stay in this repo's checkout — but since 11.4.0 they part ways at the files[] law: L6's directory is listed and ships inside @soulcraft/brainy's tarball; L12's is deliberately NOT listed there, and instead scripts/model-packages.mjs copies it verbatim into @soulcraft/brainy-model-ms-marco-minilm-l12-v2 — see Models: what ships inside, what is optional.

Loading L12 explicitly

There is no per-store rerank stamp on this base yet that would let the stage pick a checkpoint automatically per store (that lands with a later stage of the always-on work); today reaching L12 is a deliberate caller choice:

import {
  NativeCrossEncoder,
  resolveCrossEncoderAssetsDir,
} from '@soulcraft/brainy'

const l12Dir = await resolveCrossEncoderAssetsDir('L12')
if (l12Dir === null) {
  // Neither this repo's own source tree nor the optional package
  // (@soulcraft/brainy-model-ms-marco-minilm-l12-v2) is present — install
  // it: npm install @soulcraft/brainy-model-ms-marco-minilm-l12-v2
  throw new Error('L12 checkpoint not found')
}
const l12 = NativeCrossEncoder.forAssetsDir(l12Dir)
await l12.initialize()TS

resolveCrossEncoderAssetsDir('L12') resolves the in-package asset directory when present and, since 11.4.1, the optional model package's directory when installed instead — identical either way to the caller; see Models for the install line and docs/embed-service-cli.md's "Models" section for the resolution order in full. CROSS_ENCODER_VARIANTS.L12 names its directory and upstream model id directly, for anything that just needs the strings.

What comes back

Every scored row carries rerankScore: the model's raw logit, higher meaning more relevant. It sits beside score, which stays the hybrid rank the candidate set was built on. The logits are unbounded and are comparable only within one call — never across queries, and never across model variants.

An absent rerankScore on a reranked page is meaningful: that row was not scored, not scored zero. Three rows can be unscored — one past topK, one with no embeddable text, and one on a page shorter than topK.

The call's own stamp comes back two ways:

import { rerankStampOf } from '@soulcraft/brainy'

const stamp = rerankStampOf(rows)
// { model: { id: 'cross-encoder/ms-marco-MiniLM-L6-v2', variant: 'L6' },
//   k: 20, ms: 512, budgetMs: 2000, unscorable: 0,
//   pairsScored: 20, perPairMs: 25.6, projectedMs: 0, subBatches: 6,
//   textRule: 'shortest-faithful', summaries: 14, capped: 3, textCapTokens: 128,
//   applied: 'explicit', precision: 'fp32', cacheHits: 0, cacheMisses: 20 }TS

applied says where the numbers came from — 'default' for the fixed rule, 'explicit' for numbers you passed. precision is what the weights were loaded in. cacheHits / cacheMisses are the score cache, below.

and, when you are already using the per-call engine budget, at report.rerank:

const { results, report } = await brain.findWithReport({ query, limit: 20, rerank })
report.rerank              // the same stamp
report.phases.rerank       // { ms, calls, scored, unscorable, perPairMs,
                           //   subBatches, textRule, summaries, capped }
report.phases.rerankLoad   // the one-time model load, if this call paid itTS

The stamp on the array is the one to read under concurrent traffic: findWithReport() arms a process-wide budget and refuses to run two at once.

Ordering

Rows are ordered by rerankScore, descending, with one rule on top: scores within 1e-3 of each other are a tie, and a tie keeps the candidate set's own order. Cross-encoder logits are computed in f32 and their last bits depend on how the batch was blocked; without a tie band, that noise would silently reshuffle rows the model considers equivalent and a page's order would not be reproducible.

Rows inside the window that carry no embeddable text are not scored. They keep their relative order after the scored rows and are counted in the stamp's unscorable, so a page the model only partly ordered says so.

Rows past topK are untouched, in their original order.

Batching

The window is scored in sub-batches, ordered by text length: one pair first, then four at a time. Every pair in one forward pass is padded to the longest pair in it, and a padded position costs the same matmuls as a real one — so one long row among short ones would make the whole batch cost what the long row costs. Ordering by length puts similar lengths together; a pair's score depends only on that pair, so which batch it travelled in cannot change it. MEASURED: the largest disagreement between a pair scored alone, in one pass, and in the 1-then-4 shape is 1.9e-6 of a logit, three orders below the tie band (src/rerank/rerankSubBatch.native.test.ts).

Longest-first, deliberately: the probe pair is then the most expensive in the window, so the rate the budget projects from is the worst case rather than the best. A conservative projection refuses early; an optimistic one discovers the overrun after paying for it.

The sub-batch size is not a knob. The number of pairs in flight is how the budget is made checkable at the k real callers use; a caller who wants more work done raises budgetMs.

What a row is scored on — the shortest faithful text

In order:

  1. The row's summary, when its metadata carries a non-empty string one. A row whose author already answered "what is this?" in one line is scored on that line: the same order for a fraction of the wall.

  2. Otherwise the row's data text — the same text the row was embedded from, derived by the same function. Deriving it any other way would produce an order that is nearly right, which is the worst outcome, because nothing would fail.

As of 11.3, whichever text the rule picked is then capped to its first 128 tokens — a FIXED constant, the same on every host, applied ONCE in the native pass using the model's own tokenizer (never a character estimate, so the cut always lands on a real token boundary). This is independent of and happens BEFORE the [query, document] pair is built and truncated to the model's own 512-token window (longest-first) — the 128-token cap bounds what one candidate can cost regardless of query length; the 512-token window is the model's own positional ceiling on the pair as a whole, and with the cap in place it is rarely what actually limits a candidate's text anymore.

The rule is the same everywhere: a row with a summary is scored on its summary on every host, and a row over the cap is cut to the same 128 tokens on every host, so the answer does not move between machines. The stamp names the rule (textRule), the cap (textCapTokens), counts the rows read from their own summary (summaries) and counts the rows the cap actually shortened (capped).

Naming a field yourself. A caller who knows its rows better than the rule does can say so:

find({ query, limit: 20, rerank: true })   // the fixed rule; no numbers to passTS

The named field is read from every row's metadata. A row that does not carry it as a non-empty string refuses the call by field name and row id (RerankUnavailableError, reason: 'text-field-absent') — falling back per row would score one page against two different notions of what a row is.

The refusals

Never a silent pass-through. The failure this stage replaces was exactly that: a JS reranker whose loader "resolves to null rather than throwing" logged rerank=0 ms and served the un-reranked order for weeks, and nothing could tell.

RerankBudgetExceededError

The stage would spend more than budgetMs. It refuses with the whole arithmetic — k, pairsScored, perPairMs, projectedMs, elapsedMs, budgetMs — because a partial re-order is a page in an order nobody can describe, so none is returned.

The budget bounds the work, not just the report. After each sub-batch the rest of the work is projected from the rate just measured, over every pair still to score, and elapsedMs + projectedMs > budgetMs refuses before the next pair is paid for. Since the first sub-batch is a single pair, the worst a refusal can cost is one pair. (Before this, the window was scored in one 16-pair pass: at the k callers actually use the first budget decision point came after everything had been paid for, and on paragraph-class rows on a busy host a 500 ms budget refused at 11,629 ms.) The wall after the last sub-batch is checked too, for a host that ran far slower than the rate it had just shown.

The candidates come back with the refusal. error.candidates is the un-reranked page in the vector leg's own order — the work the cheap leg already paid for — so a fallback does not have to re-run the query:

try {
  rows = await brain.find({ query, limit: 20, rerank: true })
} catch (error) {
  if (error instanceof RerankBudgetExceededError) {
    rows = error.candidates          // the vector order, explicitly chosen
    servedUnreranked = true          // and say so downstream
  } else throw error
}TS

They are never served automatically: a page that reached a caller looking like a result is a page nobody can tell apart from a reranked one.

A host already measured too slow refuses immediately. The stage remembers the last per-pair wall it measured for each store, in memory, per process. When k x that rate cannot fit budgetMs, the call refuses with pairsScored: 0 and no model work at all, narrated once per store. The memo decides whether the stage runs on this machine under this load; it never touches what it returns — a run that completes returns the same order on every host, and raising budgetMs always gets the run.

budgetMs is the scoring wall. The one-time model load is not scoring and is not charged to it — a first query would otherwise refuse for a cost no later query pays. It is reported separately (report.phases.rerankLoad) and can be paid up front:

await brain.warmRerankStage()
// { model: { id, variant }, loadMs, assetsDir, weightsSource: 'mmapped-file' }TS

A serving process under a latency law should call it at start-up.

RerankStampMismatchError

The store was ordered by one cross-encoder and this process loaded another.

The first rerank a store ever runs writes _system/rerank-stamp.json:

{
  "version": 1,
  "model": { "id": "cross-encoder/ms-marco-MiniLM-L12-v2", "variant": "L12" },
  "weightsDigest": "1ed84b90…",
  "precision": "fp32",
  "stampedAt": "2026-09-04T22:11:03.418Z"
}JSON

Every later open reads it and checks it against the model actually loaded — model id and variant, the SHA-256 of the checkpoint's model.safetensors, and the precision those weights are held in. Any disagreement refuses, naming both sides and which fields differ.

Precision is in there because it is part of the answer, not part of the cost: the same pair scored in int8 and in fp32 emits two different logits, and a page ordered by one is not the page ordered by the other. A store that silently changed reranker is a store whose results changed with nothing to point at — the same class the re-embed ceremony closed for embedding spaces.

The stamp lives in the store, so a snapshot carries it and a restored copy is served by the model that ordered the original.

There is exactly one cure, and it is deliberate:

await brain.restampRerank()
// { version: 1, model: { id, variant }, weightsDigest, precision, stampedAt }TS

It rewrites the stamp to the loaded model and drops the store's score cache. Running it says, in as many words, that this store's future orders differ from its past ones. Nothing re-stamps on its own — a stamp that followed whichever model turned up would record nothing at all.

The gate runs after the model is loaded and before the first pair is scored, so a mismatched store pays one metadata read and one digest rather than a batch of forward passes it will not be allowed to serve.

RerankUnavailableError

Names one of five causes, because their cures differ:

reason

what happened

cure

engine-absent

no 'rerank' provider — this is not the native engine

install and activate the native engine; there is no JavaScript fallback, because one would score on the event loop

asset-absent

the weights are not on disk

reinstall the package (the DEFAULT checkpoint ships inside it); for L12 specifically, npm install @soulcraft/brainy-model-ms-marco-minilm-l12-v2 — see Models. Never fetched over the network either way

space-unstamped

the store declares no embedding space

run the re-embed ceremony, or open with a build that stamps at birth

space-mismatched

the store's vectors are in a space this cross-encoder is not paired with

converge the store, or omit rerank

text-field-absent

rerank.text named a metadata field a candidate row does not carry

write the field on every row the query can reach, name one they all carry, or omit rerank.text

The last two are the re-embed ceremony's law applied here: a space you cannot name is a space you must not score across. The stage is paired with the MiniLM family at 384 components, and it says so in the refusal — both the space the store is in and the one it expects.

Caller errors, refused before any work

rerank.text, when given, must be a non-empty metadata field name. When you pass rerank explicitly, topK must be a positive integer and budgetMs a positive number of milliseconds; an explicit option is all-or- nothing, because the two numbers together are what make the stage's promise checkable. (Omitting the option entirely is different — that is the fixed rule, which supplies both.) And an explicit rerank requires a text query: a cross-encoder scores pairs, so find({ vector, rerank }) has nothing to read a candidate against and refuses rather than silently skipping the stage. The default rule simply does not apply there, there being no query to score against.

The score cache

A bounded, per-brain, in-process cache of pair scores, keyed by

sha256(normalized query text) + row id + row version

A repeated or refined query never re-scores a pair it has already seen for an unchanged row.

It cannot change an answer, by construction. A pair's logit is a pure function of the query text, the row's text and the loaded model — so a cached score is the score, and reading one back skips arithmetic that would have produced the same number. Three rules keep that true:

  • The query is normalized, not folded. Trimmed, and runs of whitespace collapsed — a question retyped with a stray space is the same question and produces the same tokens. Nothing else is normalized: case and punctuation change what the tokenizer emits, so folding them would answer a question the model was never asked.

  • The key carries the row's version — its _rev and its updatedAt together, because every write moves at least one of them and add()'s overwrite semantics restamp _rev to 1. A written row is a different key, and its old score is unreachable.

  • A row with no version at all is not cached. One slot shared by every future revision of an id is the one way this cache could lie.

The bound is 8,192 entries per open brain, evicted least-recently-used, and there is no knob. One call inserts at most 128 entries (the rule's ceiling), so 8,192 holds the last ~64 full-width queries — the shape the cache exists for: a person refining a question, and a service paging through one. A full cache costs about 2 MB, which is 1.5% of the model it is saving forward passes on.

report.rerank and rerankStampOf(rows) both carry cacheHits / cacheMisses, so the saving is measurable rather than asserted. The re-stamp ceremony clears it: a new model makes every remembered score a number the new model did not produce.

Memory recall

The memory leg (recall()) runs its own budget — a single end-to-end wall across retrieve, fuse, expand, rerank and strengthen — and always reranks inside it. It does not inherit find()'s fixed rule: its rerank is one stage of a pipeline whose budget covers all of them, so it sizes and refuses its own window against that wall rather than against this one. Everything else here applies to it unchanged — the store stamp gates it, the score cache serves it, and a cached score is the same score whichever door asked for it.

What it costs

MEASURED on a 32-core CPU-only host on 2026-09-04, solo, with nothing else running on it, at commit 4bc4ce6, against L12 (the checkpoint shipped at that date). Every number below is a measurement; nothing here is projected. The 11.3 reversal below carries the current L12-vs-L6 comparison; this section's shape (row length dominates K, cores barely help) still holds for L6, just at roughly half the wall.

The rerank leg, by K and by row length

src/rerank/rerankLegBench (scripts/rerank-leg-bench.mjs), 5 runs per cell, one warm pass per shape discarded. The wall is the whole leg — tokenize, pad, forward pass, read the logits.

row length

K=16 p50 / p95

K=32 p50 / p95

K=64 p50 / p95

ms per pair

30 words

384 / 387 ms

769 / 775 ms

1,659 / 1,707 ms

~24

120 words

1,903 / 1,938 ms

4,229 / 4,233 ms

8,963 / 9,142 ms

~130

300 words

7,437 / 7,521 ms

15,185 / 15,773 ms

32,405 / 32,728 ms

~475

The fp32 stage does not meet the 300 ms law at K=32 on CPU. At the shortest row length it is ~2.5x over; at the row lengths a real brain holds it is tens of times over. That is why RERANK_DEFAULT_ON ships false, and it is exactly why budgetMs is a refusal and not a hint: a caller who sets a budget the box cannot meet gets a typed RerankBudgetExceededError naming the numbers, never a slow page pretending to be a fast one.

Set budgetMs from the table for your own row length, or lower topK. Row length dominates K: 16 rows of 120 words cost more than 64 rows of 30 words.

Why it costs that, measured

The forward pass barely scales with cores. At K=32, 30-word rows:

RAYON_NUM_THREADS

p50

1

1,068 ms

8

811 ms

32

794 ms

1.35x from 32 cores. The wall is essentially one core's fp32 matmul throughput, so the levers that would actually move it are quantized weights, a smaller checkpoint, or a device that is not a CPU — not more cores. None of those is in this release.

Precision: the int8 path, and why it is not the default

The section above names quantized weights as one of the three levers that could move this wall. It was built and re-measured against the checkpoint this package ships today (L6, not the L12 this was first tried against), and it still does not.

loadFromPathsWithPrecision() carries a second, gated path beside the f32 reference: the encoder's matmul weights — per layer attention.self.{query,key,value}, attention.output.dense, intermediate.dense, output.dense, over every loaded layer — held as GGML Q8_0 (32 weights per block, one f16 scale, symmetric, no zero-point), quantized in memory at load from the same shipped model.safetensors. There is no second asset and nothing is converted on disk. The three embedding tables, every LayerNorm, every bias and the two head layers (bert.pooler.dense, classifier) stay f32: they are lookups and elementwise passes rather than matmuls, so quantizing them would buy no wall, and the head is the last projection before the score. Whichever path served is named on the stamp:

rerankStampOf(rows).precision   // 'fp32' | 'int8'
report.rerank.precision         // the same valueTS

Not reachable by default: NativeCrossEncoder.getInstance() — the process singleton every production call takes — always loads 'fp32'. Reaching the int8 path means constructing NativeCrossEncoder.forPrecision('int8') by name, the same opt-in pattern Loading L12 explicitly already uses for a non-default checkpoint. It refuses to load on anything but a CPU device — the Q8_0 dot products are candle's CPU kernels, and a GPU device would silently take a different code path with different rounding.

What it costs — MEASURED, and it is a regression

scripts/rerank-precision-bench.mjs, batch of 16 pairs, RAYON_NUM_THREADS=8, 5 timed rounds after a warm one, on a 32-core CPU-only host, solo under the gate lock, against the default L6 checkpoint:

precision

shape

batch p50 (ms)

ms/pair

load (ms)

fp32

short (~34 tok)

163.3

10.21

107

int8

short (~34 tok)

2,330.4

145.65

89

fp32

long (~252 tok)

657.2

41.08

62

int8

long (~252 tok)

6,334.0

395.87

81

The int8 path is 14.3x SLOWER on short rows and 9.6x SLOWER on long rows — the same direction and the same rough magnitude the first attempt against the L12 checkpoint measured (14x-18x). The shape mismatch is the same one: the f32 matmuls go through a blocked, multi-threaded GEMM, while the quantized kernel is a per-row dot product written for single-row decoding, where the weights are streamed once and the arithmetic is memory bound. A cross-encoder batch is the opposite shape — several pairs of real text is many rows per matmul — so the blocked GEMM's cache blocking and threading win by more than 8-bit weights can give back, on this checkpoint as on the last one.

Weight-only quantization is the wrong lever for a batched encoder. The lever that would work is int8 GEMM with the ACTIVATIONS quantized too, which is a different kernel than the one available here.

And it does not preserve the order

scripts/rerank-precision-parity-bench.mjs — the SAME 384-row synthetic corpus and the same one-query, top-32 set-overlap-plus-Spearman measure The 11.3 reversal, measured uses for the L12-vs-L6 comparison, scored at both precisions instead of both checkpoints:

measure

value

top-32 set overlap (fp32 ∩ int8, of 32)

32/32 (100.0%)

top-32 EXACT order

diverges at rank 13

full-corpus Spearman ρ (fp32 vs int8)

1.0000

score delta, max

0.072442

score delta, p95

0.037045

score delta, mean

0.013140

Every row fp32 puts in its own top 32 is also in int8's — unlike the L12-vs-L6 comparison, where two different MODELS genuinely disagree at the margin, this is the SAME weights at two roundings, so the set stays identical. But the EXACT served order still moves: ranks 1-12 match and rank 13 does not, which is still a different answer for any caller reading past the top dozen. Hardware may change the wait, never the answer, so the int8 path is not the default and is not recommended: it costs more AND it does not preserve order. It stays kept, gated and named — a measured negative result is worth more than an untried idea, and the parity script above is what will judge the next attempt at a batched (rather than weight-only) quantized kernel.

Which checkpoint ships, and why (history through 11.2.5)

scripts/rerank-fixture-quality.mjs against a consumer's own graded recall fixture (4 queries, 15 candidates each, scored on the fixture's snippets — a lower bound on full-text quality, and identical inputs for every arm):

arm

NDCG@10

top-1

top-3 overlap

ms/query

the engine's vector order (no rerank)

0.4468

25%

25%

—

L12 (shipped through 11.2.5)

0.6528

25%

58%

520

L6 (alternate, this run)

0.6018

25%

58%

259

The stage earns its milliseconds: NDCG@10 rises from 0.4468 to 0.6528 (+46%) over the order the vector leg produced on the same rows, and the top-3 more than doubles its overlap with what the consumer's own pipeline judged relevant.

Through 11.2.5 this table was the whole answer: L12 shipped because L6 measured 8.5% worse on NDCG@10 here — not equal, and since neither checkpoint reached a 100 ms budget, halving a wall that was over by 8x did not buy the target, it only cost quality.

This fixture is 4 queries. It is real and it is honest, but it is small enough that one checkpoint's placement on 15 candidates can move the aggregate several points. See below for why the ruling that ships L6 in 11.3 does not treat this table as the last word.

The 11.3 reversal, measured

David ruled 2026-09-08: L6 becomes the default. Two things changed the answer from the table above, and both are stated here rather than asserted:

  1. The model's own benchmark, not this repo's 4-query fixture, is the number that should decide a default this consequential. Per the published model cards, MRR@10 on the full MS MARCO dev set (6,980 queries) is L6 39.01 vs L12 39.02 — a 0.01-point gap, not the 8.5% NDCG@10 gap the bespoke fixture showed. Both numbers are true; they are not the same measurement. The 4-query fixture remains useful for THIS package's own candidate shapes, but a fleet-wide default is set from the standard benchmark, not from four queries.

  2. CPU cost is real and recurring; the quality gap at fleet scale is not established to be. About half the wall, every call, on every host that reranks.

Per-pair cost, batch 16, RAYON_NUM_THREADS=8 (scripts/rerank-variant-bench.mjs, 5 rounds/cell):

checkpoint

shape

batch p50 (ms)

ms/pair

load (ms)

L12

short (~34 tok)

296.8

18.55

75

L6

short (~34 tok)

151.9

9.49

69

L12

long (~252 tok)

1,290.8

80.68

65

L6

long (~252 tok)

619.7

38.73

57

MEASURED speedup: 1.95x on short rows, 2.08x on long rows — squarely "about twice as fast," the number the ruling was made on.

The order does move, and here is by how much. scripts/rerank-parity-bench.mjs scores one query against a fixed, deterministic 384-row synthetic corpus with BOTH checkpoints and compares each model's own top-32:

measure

value

top-32 set overlap (L6 ∩ L12, of 32)

24/32 (75.0%)

full-corpus Spearman ρ (L12 vs L6 scores)

0.9906

Eight of the 384 rows are in exactly one checkpoint's own top-32 and not the other's on each side — a real reshuffle at the margin, not noise: the Spearman correlation over the WHOLE corpus is 0.99, so the two checkpoints agree almost everywhere, and the disagreement concentrates exactly where a cross-encoder's ranking is least confident — the boundary around rank 32, not the top few rows. Reported honestly, with no tolerance gate: this is the quality delta the ruling accepted, not a bug.

Reported honestly, with no tolerance gate: these are different models, so disagreement is expected. A store's rerank order under 11.3 is not identical to its order under 11.2.5, even on the same rows and the same query — see the CHANGELOG for the rollback-class statement.

L6 ships as the default from 11.3, inside the package. L12 stays vendored in this repo's source and loadable by name — see Loading L12 explicitly — for a caller that has already validated its own recall fixture against L12 specifically and wants to keep scoring with the exact checkpoint it measured against; from 11.4.0 it reaches a consumer as the optional package @soulcraft/brainy-model-ms-marco-minilm-l12-v2 rather than inside the main tarball (see Models) — this repo's own checkout still carries the source directory either way, which is what scripts/rerank-variant-bench.mjs and scripts/rerank-parity-bench.mjs score against. Both score BOTH vendored checkpoints in one run (no flag to choose one — the whole point is the side-by-side); re-run either yourself to reproduce the tables above on your own host.

What is coming: late interaction, measured

From 11.4 this stage gains a second, cheaper depth: late interaction, which encodes each row's tokens ONCE at index time and scores a query by MaxSim over the stored vectors. It is built and pinned but not yet wired into find() — nothing below changes what a find() does today, and this section is here so the cost comparison is in one place rather than two.

MEASURED on the build box in its exclusive lane, job li1-measure-1, on the vendored answerai-colbert-small-v1 checkpoint at the geometry that ships (96 components per token, a 128-token document cap, 32 query tokens):

this stage (L6 cross-encoder)

late interaction

per candidate, short rows

9.49 ms

50.79 µs

per candidate, long rows

38.73 ms

50.79 µs

fixed cost per find()

none

17.5 ms (one query encode)

K=32, long rows, all in

1,239 ms

19.1 ms

K=128, long rows, all in

4,957 ms

23.8 ms

paid at write time

none

22.3 ms/row short, 84.6 ms/row long

stored per row

none

11,600 bytes (12,288 at the cap)

The trade is stated plainly because it is the whole design: late interaction does not make the model cheap, it moves the model off the query path onto the write path, where it is paid once per row instead of once per query-row pair, and buys back the storage. A cross-encoder's score cannot be pre-computed — it does not exist until the query arrives — which is why this stage stays default-off and why the cheaper depth is the one that can be always-on.

Nothing about the numbers above makes this stage worse. It is the depth you reach for when you want a model to read the query and the row together, and from 11.4 it reads a much shorter list, because the list it receives has already been re-ordered by a stage that saw token-level interaction.

Model load and memory

The table below is L12's, from the 2026-09-04 measurement — historical, and a reasonable upper bound: L6 is the smaller checkpoint (~91 MB of safetensors against L12's 133,469,020 bytes), so its load and RSS are not larger. The per-checkpoint loadMs in the 11.3 table above is the current, checkpoint-specific number.

fact

measured (L12)

load (memory-mapped weights, cold)

292 ms

load (page cache warm)

~70 ms

RSS after load

243 MB

RSS after a full K sweep

621 MB

The load is not charged to budgetMs; brain.warmRerankStage() pays it up front. The weights are mapped, so a second process on the same box shares the file's pages rather than adding another copy.

Opted out is byte-identical

With rerank: false — and on any build whose stage does not apply by default, which is every build shipping today — no provider is looked up, no model is loaded, no timer is armed and no property is attached. find() returns exactly what it returned before this stage existed.

Even on a reranked page the wire format does not change: the stamp is attached under a Symbol.for key and is non-enumerable, so it is invisible to JSON.stringify, Object.keys and every for…in a caller already runs.

Fusion and purposes

The rerank stage re-orders a page. Rank-signal fusion orders the pool, and it is the stage that decides which rows the page holds at all.

const rows = await brain.find({
  query: 'what did we decide about pricing',
  limit: 10,
  purpose: 'conversation'
})

rows[0].fusion
// { fusedScore: 7_412_500_000,
//   purpose: 'conversation',
//   contributions: {
//     text:    { raw: 0.031, normalized: 1_000_000, weight: 6000, contribution: 6_000_000_000 },
//     graph:   { raw: 1,     normalized:   500_000, weight: 1000, contribution:   500_000_000 },
//     recency: { raw: 259_200_000, normalized: 365_000, weight: 2500, contribution: 912_500_000 },
//     boost:   { raw: 0,     normalized:         0, weight:  500, contribution:             0 } } }TS

The formula

fusedScore = w_text x text + w_graph x graph + w_recency x recency + w_boost x boostTEXT

Four signals, each normalised to an integer in [0, 1_000_000]; four weights, each an integer in [0, 10_000] summing to exactly 10_000. Ties break by id ascending. That is the whole of it — no learned weighting, no hidden normalisation step, no adaptive term.

Every number in a row's position is in the row. The four contribution values are weight x normalized and nothing else, so they sum to fusedScore exactly — a caller can add up what a row reports and get the number it was sorted by, to the digit. That is the difference between a ranked answer and a ranked answer anyone should trust.

Fixed point, and why

A sum is a reduction, and a floating-point reduction's answer depends on the order its terms were added in — which is a function of vectorisation width, chunking and thread count, none of which are part of the question you asked. So the combine is integer multiply-add: integer addition is associative, and no chunk width, evaluation order or core count can change the result.

One place a float appears at all: converting a signal provider's raw score onto the grid. That is a fixed sequence of exactly-rounded operations on a single row's value, with no reduction anywhere in it, so IEEE-754 makes it a function of the input bits alone.

Nothing in the normalisers calls exp, log or pow, and a pin reads the source to prove it. The natural spelling of a recency decay is exp(-age/tau), and exp's last bit is not specified by IEEE-754 — it genuinely differs between platforms and between libm versions. A decay built on it would make a page's order a property of the machine. The recency signal gets the same shape from an integer shift and one division instead: exactly 1/2 at one half-life, exactly 1/4 at two, straight-line between them.

The four signals

signal

what it reads

normalisation

absent means

text

the relevance score, from the store's stamped source

pool-max-relative (the leg's own score over the pool's largest) or maxsim-per-query-token (the late-interaction total over the query's token count)

a pool with no text evidence reads zero for every row, not one. On a store mid-backfill a row the tokens projection has not reached keeps its DENSE value — it is not scored zero and not swept down — and its contribution says source: 'dense' where a covered row's says 'maxsim'

graph

hops from the anchor set through the adjacency

hop-halving — 1_000_000 >> hops, and exactly zero past the purpose's cap

a row the walk never reached reads a raw hop count of -1 and a signal of zero: "no path within the cap", not a missing value

recency

system.updatedAt, else system.createdAt

age-halflife-piecewise-linear over the purpose's half-life

a row with neither timestamp reads -1 and zero

boost

your declared { field, equals, amount } rules

declared-boost-sum — matched amounts summed, then clamped

no rules means zero for every row, so no ordering effect at all

Two sources for text, and neither is the other's fallback: which one a store uses is stamped, and a build carrying the other refuses by name. The cross-encoder is deliberately not among them — its logit is the most faithful text signal this engine can produce and it is the one that cannot be a pool signal, at 25 to 130 ms per candidate. A fusion whose text term existed only for the page would order the page by four signals and the pool by one, which is the cosmetic reorder this stage exists to end. So the cross-encoder keeps its own place: a polish applied to the page after the fused order has chosen it.

The three purposes

purpose

serves

text

graph

recency

boost

half-life

hops

conversation

a chat turn's context bundle for one person

6000

1000

2500

500

7 d

2

workspace

an entered business, or a matter

2500

6000

1000

500

90 d

3

documents

passages over a book or a codebase

8000

1500

0

500

90 d

2

Every constant has a reason you can check rather than take:

conversation keeps text first and recency bounded. 6000 is the largest single weight, so no combination of the other three can lift a row the query does not match. Recency's whole range is 2500, which is 2500/6000 of the text range — so recency can overturn a text gap of at most 0.42, and never an exact match against nothing. "This week outranks last year at equal relevance" therefore holds; "last week outranks the answer" cannot. The half-life is seven days because the shape's own words are "what happened this week".

workspace makes "graph dominates text" an inequality. The ratio is w_graph > 2 x w_text, and that is the sentence turned into arithmetic at the hop that matters:

one hop:  6000 x 500,000 = 3.0e9  >  2500 x 1,000,000 = 2.5e9   -> the anchored row wins
two hops: 6000 x 250,000 = 1.5e9  <  2500 x 1,000,000 = 2.5e9   -> text can win againTEXT

A row one hop from what you entered outranks a row with a perfect text match and no path to it — for a business or a matter that is the right answer. At two hops the inequality reverses by design, so the pull is a neighbourhood rather than a flood. Both halves are pinned. The hop cap is three because a matter reaches its subjects through the business, and a cap of two would leave the matter's own rows outside the walk.

documents weights recency at exactly zero. Not merely small: the shape's word is "irrelevant", and a small weight is not irrelevance — it is a tiebreaker nobody asked for, which is how a chapter written last week quietly outranks the chapter that answers the question. The age is still read and still reported in each row's contributions, so you can see that it did not count.

The default purpose, and how it was chosen

documents — and the choice is measured rather than preferred.

MEASURED on the build box, job li2-fixture-1 (SUCCEEDED, exit 0), against a captured recall fixture: four real queries, ten rows each, in the order a real recall door actually served them. Each purpose ordered a pool built from that fixture's own relevance values, and was scored against the served order.

purpose

top-1 match

top-3 overlap

Kendall tau

conversation

4 / 4

0.92

0.856

documents

3 / 4

0.92

0.844

workspace

0 / 4

0.58

0.756

Read it as three facts, not one ranking:

  • conversation reproduces that bed best, which is the bed working as intended: the capture is a chat turn's recall, so the purpose built for that shape should win on it. Its weights are not a guess.

  • documents is within noise of it on top-3 and tau while agreeing on top-1 three times in four — which is exactly the property a DEFAULT needs. A caller who names no purpose gets an order very close to the one they are getting today, rather than a new one.

  • workspace does not reproduce it at all, and must not. Its graph term is deliberately able to overrule text; on a bed with no real graph that reads as disagreement, and on a business or a matter it is the whole point. A purpose that scored well here would be a purpose that was not doing its job.

The bed can speak to ORDERING and to nothing else — it carries relevance and the served order, and it carries no graph edges and no row timestamps, because the door that produced it ranked by neither. scripts/fusion-fixture-agreement.mjs says so in its own header and reports counts and ranks only, never row content.

Two more reasons agree with the measurement: documents is the shape the fleet itself named as the default, and it is the purpose closest to today's order — text dominant, recency exactly zero — so a caller who never asked for a recall shape does not silently acquire one.

The stage itself ships off by default in 11.4 — the build's switch reads false, exactly as the cross-encoder's does — and naming a purpose is how you get it.

Why it is still false, said plainly. Everything the flip needs is built and pinned: the rule, the three purposes, the stamp, every refusal, the cursor key, and a measured wall (the table under "What it costs"). What is NOT yet in hand is the one thing the precedent demands — a latency receipt taken on the FLEET's own hosts, at the pool sizes those stores actually assemble, with the adjacency I/O a real store pays included. Flipping the switch re-orders every existing find({ query }) on every store at once; the walls above were measured on the build box over synthetic pools with the graph served from memory, and a number taken on one box over a fabricated pool is not a receipt for that. It flips on a measurement, not on a hope — the same bar RERANK_DEFAULT_ON is held to, and for the same reason.

Three complete calls

// 1 — CONVERSATION. A chat turn's context bundle. Anchors inferred from the
//     query's own strongest matches; nothing else to say.
const turnRows = await brain.find({
  query: 'what did we decide about pricing',
  limit: 10,
  purpose: 'conversation'
})

// 2 — WORKSPACE. A business or a matter. The anchor set is EXPLICIT: entity
//     ids, given inside the purpose. Graph proximity is measured from these.
const roomRows = await brain.find({
  query: 'what did we agree',
  limit: 20,
  purpose: { name: 'workspace', near: [businessId, matterId] }
})

// 3 — DOCUMENTS. Passages over a corpus, with the caller's own boosts.
const passageRows = await brain.find({
  query: 'how does the flush path recover',
  limit: 10,
  purpose: {
    name: 'documents',
    boosts: [
      { field: 'source', equals: 'decision', amount: 1 },
      { field: 'vfsPath', equals: '/docs/architecture.md', amount: 0.5 }
    ]
  }
})TS

The three are a closed set, and you can reach it without quoting. The union type ships beside an as const object of the same name, so a caller autocompletes and a typo is a compile error rather than a runtime refusal:

import { RankingPurpose } from '@soulcraft/brainy'
await brain.find({ query, purpose: RankingPurpose.workspace })TS

docs/api-contract.json publishes the set under rankingPurposes — the option key, the three values in declared order, the default and the off switch — so a generated tool schema carries a JSON Schema enum rather than a free string, and a wire caller learns the allowed values from the contract instead of from a refusal.

An unrecognised value is refused, never ranked by the default. purpose: false is the only off switch; anything else the engine does not serve raises FusionOptionError listing the three. That is deliberate and it is pinned even with the stage switched on by default, because the alternative — falling through to the default purpose — serves a page that is indistinguishable from one that honoured the caller.

purpose.near is a list of entity ids, and it is NOT {@link FindParams.near}. FindParams.near is a pre-existing RETRIEVAL criterion with a different shape — { id, threshold? }, "find rows near this entity's vector" — which changes WHICH rows are candidates. purpose.near changes only how the candidates are ORDERED. They can be used in the same call and mean different things; the two shapes differ (an array of ids against one { id } object) precisely so that neither can be passed where the other was meant.

A boost's amount is a fraction of the boost signal's full range, in [0, 1]. Matched amounts are summed and then clamped, so writing more rules never buys more than the purpose's boost weight.

The legacy fusion option

FindParams.fusion — { strategy, weights: { vector, graph, field } } — is the open engine's hybrid-search blend, applied while the candidate set is being built. As of 11.4 it is superseded by purpose, and not removed:

  • A call that names only fusion behaves exactly as it always has. Nothing is logged, nothing is warned, and the default purpose rule stands down in front of it rather than adding a second weight system to a call the caller never changed.

  • A call that names both fusion and an explicit purpose is refused by name — ConflictingRankingOptionsError, naming both keys. They are two weight systems for one question: fusion.weights blends the candidate set as it is assembled, a purpose is a stamped fixed-point ordering over the finished pool. A call setting both would silently get one of them, and which one is an accident of stage order rather than a decision anybody made.

  • purpose: false alongside fusion is not a conflict. It is an opt-out.

Whether fusion is removed at all is a 12.0 decision, taken deliberately and in the open — never by quiet attrition.

The anchor rule

Explicit wins over inferred. purpose: { name: 'workspace', near: [businessId] } measures graph proximity from exactly those ids; an empty array is a statement ("no anchors") rather than an omission.

With no near, the anchor set is inferred from the head of the text ordering — the rows this query names, as this store understands it. That is the native form of what a consuming service does by hand today: run related() outward from the top hits and blend the neighbours in JavaScript.

One rule makes it informative rather than tautological: an inferred anchor gets no credit for being itself. It seeds the walk and scores zero on the graph term. Crediting it at hop 0 would hand the rows that already won on text the whole graph term as well — MEASURED on an eight-row corpus as exactly that: the top three each collected 1,500 basis points of free graph score and no boost could reach past them. An explicit anchor does sit at hop 0, because a caller naming it is evidence and a text ranking is not.

Paging a fused ordering

There is one ordering per (query, generation, clock). The stage does not read the wall clock directly — it reads it quantised to 60 seconds and stamps the value it used — so every call inside one tick computes the identical ordering, and offset paging reads that one ordering rather than a slightly different one per page.

limit no longer changes the pool. It used to: the candidate legs were bounded at limit * 2, so find({ limit: 8 }) assembled a wider pool than find({ limit: 3 }) and a paged walk at limit 3 walked a smaller ordering — comparing it against one wide call found rows missing. That was measured while writing this page's own paging pin, and the ranking seam fixed it at the source: RANK_CANDIDATE_POOL is a constant, so every ranked find asks its legs for the same width. A caller's limit chooses how much of the ordering they see, never what the ordering is, and both halves are pinned — the same pool at two limits, and a paged walk that concatenates to the wide call exactly.

fusionStampOf(rows).pool is how you see it, on every page. .nowMs is the clock the ordering was built against. With the query and the store's generation they are what the cursor's ordering key names.

The stamp, and the one cure

The weights are stamped per store at _system/fusion-stamp.json, beside the rerank stamp and under the same law. The first fused query a store ever runs writes it; every later one checks it, and a disagreement refuses by name with every differing field as a dotted path — turn.weights.recency, normalization.text — rather than saying "the weights".

await brain.restampFusion()                          // to this build's weights
await brain.restampFusion({ textSource: 'maxsim' })  // and to late interactionTS

It returns the stamp it wrote, as its receipt. Running it says, in as many words, that this store's future pages differ from its past ones. Nothing re-stamps on its own.

What it costs

MEASURED on the build box (32-core / 184 GB), exclusive lane, job li2-wall-1, over synthetic pools with real ages and a sparse graph — median of nine runs per cell, in milliseconds of stage wall:

candidate pool

conversation p50 / p95

workspace p50 / p95

documents p50 / p95

128

0.36 / 1.74

0.29 / 0.38

0.34 / 0.43

512

0.86 / 2.02

0.75 / 2.72

0.74 / 0.75

1,000

1.45 / 4.04

0.97 / 2.95

0.94 / 3.14

5,000

7.83 / 10.61

7.73 / 9.59

7.76 / 7.99

conversation is the most expensive of the three at every size, and that is the shape working: it is the only purpose that pays for both a graph walk and a recency term with a non-zero weight. documents walks a graph too — its co-mention pull is 1500 — but with recency at exactly zero it reads one fewer signal per row.

Growth is linear in the pool and in nothing else, measured on the same job: a pool of 500 costs 0.43 ms and a pool of 1,000 costs 0.86 ms — exactly 2.0x for exactly 2x the rows.

The one latency number any purpose carries is conversation's: under 100 ms server-side, for the whole recall. This stage is one term of that — the retrieval leg and, on a maxsim store, the query encode are the others — so the pin holds it to a tenth of the budget at the design's own candidate cap of 1,000, which is where the numbers above put it with room to spare.

What those numbers do and do not include. They are the stage's own arithmetic — the text normalisation, the combine, the sort — plus a column read and a graph walk served from memory. They do not include the adjacency I/O a real store pays: the walk reads one node's edges at a time through related(), so a frontier of N nodes is N edge reads, and on a cold store those are disk. The walk is bounded three ways so that cost cannot run away, but it is real, it is not in the table above, and the honest way to shrink it is a batched adjacency read rather than a wider frontier.

The batched adjacency read is 11.5's, deliberately deferred. It is a change to the walk's I/O shape, not to its arithmetic: one read per frontier instead of one per node, with the same bounds and the same by-id truncation. Doing it in 11.4 would mean landing a new native door on the release the fusion first ships on, and measuring its effect needs the same fleet-host receipt the default switch is waiting for — so the cost above stays honest about what it excludes, and the two land together rather than one of them landing unmeasured.

The wall grows with the pool and with nothing else. A cost that scaled with the store rather than with the work is the defect class that took two consumers down at 4.2.1, so it is pinned: the stage reads columns for exactly the pool's ids in one batched read, and walks a frontier bounded three ways — the purpose's hop cap, a 4,096-node frontier cap, and stopping the moment every pool row has a distance. When the frontier cap bites it truncates by id ascending, a deterministic choice rather than whatever order the adjacency answered in, and the page's stamp says it happened.

The refusals

reason

what happened

cure

engine-absent

no 'rank-fusion' provider — this is not the native engine

install and activate the native engine; there is no JavaScript fallback, because an approximate fused order is a page nobody can reconstruct from its own stamp

tokens-family-absent

the store is stamped to the late-interaction text signal and the tokens projection is not attached

run the re-index ceremony, or restamp to the vector source

tokens-backfill-incomplete

restampFusion({ textSource: 'maxsim' }) on a store whose projection is attached but whose ceremony has not FLIPPED — _cor_tokens_reindex/ready.json is absent, so the projection covers some prefix of the store rather than the store

run the re-index ceremony to completion — it resumes from where it stopped — and restamp once it reports its receipt. A family is attached from its first published installment, so "attached" is not "covered", and a maxsim stamp written over a partial projection would order every later page by a text signal most rows cannot produce

conflict

an explicit purpose and the legacy fusion option on one call

ConflictingRankingOptionsError names both keys — drop one; fusion alone stays unchanged

query-absent

a purpose was named with no text query

the first signal is text relevance; pass query, or drop purpose

boost-field-absent

a boost names a field the column index does not serve for any candidate

name a field the rows carry, or drop the boost — a boost that could never match would order the page as though it had not been written

recency-field-absent

the purpose weights recency and the store serves no timestamp column

repair the metadata index, or order with documents, whose recency weight is exactly zero

Malformed options — a purpose that is not one of the three, a near that is not a list of ids, a boost amount outside [0, 1], an aggregate — refuse before any work, with FusionOptionError.

Fused and reranked together

They compose, and the order is fixed: the fusion chooses the page out of the pool, then the cross-encoder orders the few rows inside it. A row's fusion contributions still describe the fused score that chose it; rerankScore describes the order it was finally served in.