Roadmap
Four phases, each ending with measurements that decide whether the next one is worth starting. No later phase begins without the numbers from the one before it — and two of the four were closed by their own gate: the measurements said the work was not worth doing.
Measured figures are marked as such and come from benchmarks in the repository. The rest — phases 3 and 4, and the table further down — are still projections from the design, and none of it is confirmed until each phase runs its own exit measurements.
Core
Phase 1 — Core
Done · measuredEverything in the engine design specification's file format, copy-on-write B+tree, free
list, transactions, secondary indexes, delete_range and encryption sections. 4,000-6,000 lines.
Exit criteria, all met and measured:
- Property-based tests and fault injection green.
- At most 40 bytes per log entry over a million synthetic records — 34.14 measured, and 33.94 at 63 million.
- A freshly created database under 2 KB — 1,536 bytes.
- Constant-cost opening with 30 KB and 2 GB files — 4 physical reads at both, and at 14 MB in between.
- A comparative benchmark against SQLite on quota's real schema — 39% less disk and pruning 123× faster.
- Zero runtime dependencies and no
unsafeof our own, both verified by tooling.
RAM and CPU
Phase 2 — RAM and CPU
Done · measuredWithout touching the on-disk format. The plan listed six changes; measuring it first retired half of them and surfaced one that was not on the list and mattered more than any of them: writing each modified page once per commit instead of once per row. Copy-on-write has to copy the root-to-leaf path on every write — that is what buys crash safety without a journal — but it never had to send that path to the medium every time. A thousand-row commit was issuing 2,009 physical writes for 13 distinct pages.
Exit criteria: 40 KB or less per open database on a pro node, 8 KB or less on a micro node, write amplification of 1.2× or less.
- RAM per open database: 19,744 B on a pro node and 4,864 B on a micro one — met by lowering one constant, once the measurements showed a bigger cache buys nothing on a scattered workload.
- Write amplification: 335× down to 1.71× at the batch size the edge uses, and 1.10× at larger ones. Met at larger batches; what is left at the smaller ones is page granularity, not waste.
- Of the six changes originally planned for this phase, five were retired by measurement: the structure they were meant to shrink turns out to be 2,400 bytes per open database, and a shared page pool would reclaim 6.7 MB across 2,700 tenants. The sixth — closing idle databases — depends on how the edge is deployed, which is not the engine's decision.
One measurement contradicted another, and the page keeps both: the cache curve is flat on scattered lookups but not on a hot working set, where 32 pages do nine times fewer reads than 4. The default stays at 4 because writes outnumber reads at the edge — but a deployment serving heavy reports should raise it, and the 8 KB per-database target this phase was written against turned out to protect nothing: it was set when an open database was estimated at 832 KB, and the real floor is 2.4 KB.
Segments
Phase 3 — Append-only segments
Not openingWould replace the B+tree with sequential segments and a sparse index in the log keyspace. About 700 lines, and the only phase that changes the engine's on-disk format — which is why it was the one phase specified with a gate: measure before committing to it.
The gate closed it. Two of its three exit criteria were already met by work from phases 1 and 2: leaves sit at 99.5% fill, and pruning half a million rows takes 2.1 ms. The third — write amplification near 1.0× — would take a 2,700-tenant fleet from 4.1 to 2.4 GB of writes a day, on a machine that already writes 9.1 GB a day on its own.
Columnar
Phase 4 — Columnar encoding
Not openingDelta and bitpacking on timestamps, RLE on low-cardinality columns, frame-of-reference on latencies and sizes, dictionary and bitpacking on interned identifiers, LZ4 over the block. Target: 13 bytes or less per log entry, against 34 today.
It needs phase 3's sealed segments, and its own prize does not carry the two of them: across
2,700 tenants it would take the fleet from 6.06 to 4.35 GB. That is 1.71 GB saved on a disk
with 59 GB free, in exchange for 1,300 lines of codecs and an LZ4 written by hand — the
engine forbids the crate everyone uses, because it contains unsafe.
What would reopen both: a fleet an order of magnitude larger, where 1.71 GB becomes 17. A tier that keeps raw records for months rather than days, where bytes per entry starts mattering again. Or a deployment where writing is the expensive resource — an SD card in a real IoT node, not a server volume. None of the three is far-fetched. None is true today.
Open work
What is actually next
In progressWith phases 3 and 4 closed by measurement, the next work is not an engine phase. The engine is finished for the workload it was built for; what is left is putting it under that workload, plus three pieces of engine surface that are named, deliberate absences rather than oversights.
- Under quota's edge. One file per tenant, with the configuration the
measurements settled: 512 B pages on micro, 2 KB on starter, 4 KB on pro and
enterprise; a 4-page cache;
Sync::Commit; and a page cap per tier so one tenant cannot eat the disk of the rest. The one deployment detail this forces: page size cannot be changed on an existing file, so a node that moves up a tier keeps the geometry it was created with unless re-provisioning recreates the database — cheap in exactly the small tiers that move up most. - An explicit index rebuild. A version mismatch between the stored extractor and the registered one is always an error today, on purpose: rebuilding means re-reading and re-writing an entire keyspace, triggered by the act of opening the database — the last thing a node with a page cap can afford quietly. What is missing is the verb a caller invokes on purpose, at a time it chooses.
- Range deletion over indexed keyspaces. Rejected today, because the property that makes range deletion fast — it never reads the leaves — is the same one that makes it impossible to know which index entries to retire. The way through exists on paper: let an index declare that its key starts with the same time bucket as the primary key, which makes the index deletion another contiguous range. What is missing is the declaration mechanism, and it needs its own design, because a false declaration corrupts an index silently.
- Compaction.
reclaim_tailreturns the free tail of a file to the filesystem; it does not move live pages down. A tenant that permanently drops a tier gets its tail back today. One whose live pages happen to sit high does not.
None of the three engine items is on a date. They are written down so that the absence reads as a decision with a reason attached, which is the only kind of absence you can plan around.
Per-phase numbers
Numbers per phase, pro node
The first two rows are measured; the last two are what phases 3 and 4 would have delivered, and they are the reason neither is being built. Disk is the steady state of a real tenant over 30 simulated days, not a calculation — which is why phase 1 reads 6.3 MB and not the 46 MB the design expected. That gap, found by measuring, is what made the last two phases unnecessary.
| Phase | Disk | Write amplification | RAM per database |
|---|---|---|---|
| 1 — core | 6.3 MB | 1.71× | 19.7 KB |
| 2 — RAM and CPU | 6.3 MB | 1.10× at larger batches | 19.7 KB |
| 3 — segments | 40 MB | ~1.0× | ~34 KB |
| 4 — columnar | ~15 MB | ~1.0× | ~34 KB |
Disk figures are reliable within 10%. CPU figures are order-of-magnitude estimates, with a 50% margin. RAM figures depend entirely on how many distinct users a node sees within its retention window — the table above assumes 10,000 per pro node.
A server with 4,500 micro or starter nodes and 500 pro nodes: about 640 MB of RAM after phase 1, about 61 MB after phase 2 with every database open, and about 12 MB if 95% are idle and closed.