Sizing a brainy server
In this section
How big a machine do you need? This page answers with measurements, says which numbers are measured and which are worked out from them, and ends with a command that measures your own machine so you never have to trust a table.
Every number is labelled. MEASURED means a run on real stores, with the setup named. PROJECTED means arithmetic on measured numbers, with the arithmetic shown. Nothing is a guess dressed as a figure.
1. What costs what
Three different things limit a server, and they are not the same thing.
Disk holds every brain. A brain is a directory. All of your brains live on disk whether anyone is using them or not. Disk is the cheap resource.
Memory holds the brains that are being used right now. A brain that no request touches costs memory for nothing. A brain that is being used costs the pages its requests actually read, plus a small fixed cost paid once for the whole server process (the embedding model and the engine's own pool) — not once per brain.
Work is the ceiling that arrives first. Today one server process answers all of its brains' requests. Before memory runs out on a large machine, the rate of requests one process can answer is the limit. The native server (
brainy serve) lifts that ceiling. This page does not state a request rate, because we have not measured one for this guide.
The most important sentence on this page: a brain's memory cost is what its requests touch, not its size on disk. A brain of 2.8 GiB on disk held 41 MiB in memory while it was in use. Measured below.
2. Measured rows
All rows: the 13.0 release build, the server run under a hard memory cap set by the operating system, on copies of real stores, driven by the same mix of requests that brainy measure uses (section 4): a vector search, a plain find, a count and a related lookup, one request every 10 to 30 seconds, for 10 to 20 minutes. "Attributed" is the memory the server itself says belongs to the brain; "process" is what the operating system says the whole server process holds.
The brains
Brain | On disk (allocated / files) | Attributed memory, settled | Process total | First request, cold start | Typical request (p50) |
|---|---|---|---|---|---|
Small shop brain | 596 MiB / 53,587 files | 73 MiB | 436 – 653 MiB | 1.9 s | 24 – 35 ms (p95 up to 222 ms) |
Team board brain | 2.80 GiB / 370,459 files | 41 MiB | 366 – 453 MiB | 0.7 s | 3 – 29 ms (p95 up to 126 ms) |
Large personal brain | 59.5 GiB / 2,044,860 files | 4,486 MiB (8,766 MiB in the first minute) | 5,328 MiB | 19.5 s | 3 – 64 ms by kind (65 of 65 requests answered, none slower than 1 s) |
Sources, all MEASURED on 2026-10-04 with the build described above. The small and team brains: the 10-minute runs under an 18 GiB cap (14.6 GB budget), nothing evicted, no memory pressure. The large brain: a 20-minute run after the plain-find fix, with no cap (a 48 GiB scope), 243 samples, zero evictions.
What the rows say:
Memory follows what is touched. 41 MiB for a 2.8 GiB brain, 73 MiB for a 0.6 GiB brain. The small shop brain is smaller on disk and costs more memory than the team brain because its requests touch a larger part of it. Disk size does not predict memory.
The fixed cost per process is a few hundred MiB, paid once: the process total minus the attributed memory is about 300 to 500 MiB in the two small rows. MEASURED, same runs.
A large brain costs more while its first minutes of work run, then less. The server's open of a very large brain does catch-up work that touches much of the store: attributed memory was 8,766 MiB in minute 1 and settled to about 5,000 MiB by minute 3 and 4,486 MiB by minute 17, with the process total steady at 5.3 GB (MEASURED, large-brain run). Plan for the settled figure, but do not be alarmed by the first minutes — and run the measuring command for at least ten minutes on a large brain.
The first request on a cold process took 0.7 to 1.9 s for the small brains and 19.5 s for the large one (MEASURED, same runs; the large brain took 28.8 s and 31.1 s on the earlier build). During that time the server answers "not ready yet" by name rather than waiting silently; a client that retries gets its answer when the open finishes.
A brain that does not fit its budget
The server's memory budget is the machine's limit, minus 20 % held back as headroom, minus the process's own baseline. We ran the large brain on an earlier build, which paged in about 17 GB, under caps that made it a squeeze (MEASURED, earlier build, so the memory figures are higher than the row above and are shown only for the behaviour):
Cap | Budget | Attributed memory | Result |
|---|---|---|---|
26 GiB | 21.2 GB | 17.0 GB | never flagged, p50 28 – 42 ms |
18 GiB | 14.6 GB | 17.0 GB (over budget) | named once in the server's log, kept serving, nothing evicted |
12 GiB | 9.7 GB | 10.7 GB | cap pressed (3,363 times), nothing killed, p50 13 – 44 ms |
10 GiB | 8.1 GB | 9.2 GB | cap pressed (2,033 times), nothing killed, p50 15 – 56 ms |
Two things to take from it. A brain that alone is bigger than the budget is named once and still served — it is never closed and reopened in a loop. And squeezing the memory it was given did not slow its typical request: the working set of this request mix was no larger than the 9.2 GB it held at the smallest cap.
3. What a 32 GB, 16-core machine holds
PROJECTED from the rows above, arithmetic shown.
The budget: 32 GiB = 32,768 MiB, minus 20 % headroom (6,554 MiB), minus the process baseline (127 MiB, from a measured run's own printout) = about 26,087 MiB.
One large brain active (4,486 MiB settled), and about 500 MiB of fixed process cost beyond the baseline: 26,087 − 4,486 − 500 = 21,101 MiB left.
Divided by a typical brain's measured cost: 21,101 ÷ 73 MiB = 289, and 21,101 ÷ 41 MiB = 514.
So memory alone would hold roughly 290 to 510 further active brains of the small and team kind, beside the large one. Without the large brain, roughly 350 to 620. This holds only for brains that cost what these cost under this request mix: measure a brain like yours (section 4) before relying on it.
On disk: a terabyte holds about 1,000 brains of 1 GB each (1,000 GB ÷ 1 GB, before leaving free space for growth and snapshots). Use the allocated size, not the apparent size — brains contain sparse files whose apparent size can be tens of times larger than what they occupy (du -sh, not ls -l).
The work ceiling comes first. One server process answers every brain's requests, so the number above is a ceiling on how many brains can be held active, not on how many requests per second the machine serves. Plan for several processes (section 5) before you plan for several hundred active brains in one.
4. Measure your own machine
brainy measure --brain /data/brains/my-brain --minutes 10
brainy measure --brain /data/brains/my-brain --minutes 10 --cap 12GIt starts the server over your brain, in place: it never copies the brain, never deletes anything and writes no row. Opening a brain takes its writer lease, so a brain another process has open is refused by name — stop that process first. It then sends the same mix of requests as above, samples the server's own memory accounting every 5 seconds, and prints:
The per-minute table, one row per minute:
evicted(budget),evicted(idle),readmitted— brains the server closed to stay under its budget, closed for being idle, or reopened. With one brain these should be 0.resident MiB min/max— the memory the server attributes to the brain that minute. In a large brain's first minutes this is the open's catch-up climbing and then falling.VmRSS MiB min/max— what the operating system says the whole process holds.scope MiB max— with--cap, the memory the operating system charges the capped group.n/awithout it.doors n p50 / p95 / max ms— how many requests that minute and how long they took.slow (>1 s) by kind— each request slower than a second, by kind.
The per-kind table — requests by kind with their slow counts and latencies.
The capacity reading, the part to act on:
the budget the host printed— the server's own line: the limit it found (the machine's, or the cap's), the headroom, and the baseline.this brain, active— the attributed memory once it settled, and the minute from which it stayed settled (when the open's catch-up ended). If it never settled the reading saysNOT SETTLEDand gives no count: run longer.the process's fixed cost— process memory minus attributed minus baseline: paid once, not per brain.about N more brains of X MiB fit the budget— the budget minus this brain's settled memory, divided by X, where X is this brain's own measured cost. It is never computed from disk size. A smaller or quieter brain costs less than X, so measure a brain like the ones you will add.
Then the verdict rows. doors.all-answered fails the run if any request did not answer 200. With --cap, the cap.* rows check the cap from the server's own output and its own cgroup file — not from what the launcher says — and scope.oom-kill-zero fails the run if the kernel killed anything. A brain over its budget, or an eviction, is a reading, not a failure: it is the answer you asked for. The exit code is 0 when every request answered and the cap, if any, held; otherwise 1 with the failing row named.
--cap needs systemd-run and refuses by name without it, rather than measuring uncapped and calling it capped.
5. How to scale
What it is | Exists today? | |
|---|---|---|
Up | A bigger machine, the same server. Memory and disk grow; the work ceiling does not. | Yes. |
Within one machine | Several server processes, each with its own brains directory and port, each with its own budget (set a memory cap per process so they cannot starve each other). | A pattern: the server runs as many times as you start it; nothing coordinates the processes for you. |
Out | A brain is a directory, so a brain can live on any machine. There is no query across brains: a request is for one brain. A map from brain to machine, kept behind one name, sends each request to the right one. | A pattern, not shipped: there is no multi-machine router in the package. |
Moving a brain | A snapshot taken on one machine and restored on another, then a normal open. See Backup and restore. | Yes. |
6. Running it in your own environment
Fast local disk is required. The design keeps brains on disk and pages them in as requests need them; a slow or network disk shows up directly in the first-request wall and the large-brain catch-up above.
The embedding service runs on your hardware. Its model is part of the fixed per-process cost in section 2.
Your data never leaves the machine. The server makes no outside contact.
Backups are the data root plus the snapshot directory. See Backup and restore.
The license key travels with the package. It is checked when the engine opens, and at no other moment.
Next: brainy serve for the server's doors, and Connect your terminal for the plugin every door serves.