MVCC row heap
Fresh writes land here first. Snapshot isolation and savepoints ride on MVCC, and pages live in a buffer pool with clock-sweep eviction.
Database
Fresh writes land in an MVCC row heap tuned for OLTP. A background thread compacts committed rows into .zyr segments with per-column encoding, and analytical queries run directly on the encoded data with predicate pushdown and late materialization.
Storage
The row heap takes transactional writes, the .zyr columnar format takes analytical scans, and automatic background compaction moves rows between them. Queries never choose a side.
A background thread compacts committed rows into encoded segments while queries keep running. HybridScan reads whichever side holds the rows, in the same pass.
Fresh writes land here first. Snapshot isolation and savepoints ride on MVCC, and pages live in a buffer pool with clock-sweep eviction.
A background thread compacts committed rows into .zyr segments with per-column encoding, built for analytical scans.
One operator reads the heap and the segments together, measured at ~1.3% overhead vs a heap-only scan.
Heap and lake tables are first-class in the same SQL, joined by the same HybridScan operator, whether they live on the local node, in a shared lake, or across the mesh.
One write, end to end
01 / 04
A write is appended to the write-ahead log and fsynced through group commit. The client is acknowledged after that one fsync, before any page is written.
02 / 04
The row lives in the MVCC heap with its own visibility header. An update writes a new version and closes the old one, so readers under snapshot isolation never block a writer. A schema epoch in the slot lets ALTER TABLE change the shape without rewriting the row.
03 / 04
A background thread compacts committed rows into a .zyr segment. Each column is encoded on its own, dictionary, RLE, bit-pack, or FastLanes, and the segment keeps a zone map, the minimum and maximum of every column.
04 / 04
A query's predicate is pushed down to the scan. Segments whose zone map cannot match are skipped, the rest are filtered while still encoded, and only the rows that survive are materialized.
Keep scrolling to follow the write
Durability
Every commit fsyncs the WAL through group commit before the client is acknowledged. Dirty heap and index pages are written by a background page writer and fsynced at explicit durability barriers, so an acknowledged transaction survives a power loss without paying a per-page fsync on the OLTP hot path.
~778Ktxn/sec
Durable group-commit peak
~69.9µs
Durable commit floor, device write
~83.0x
Group-commit amplification, c=1 to c=512
These paths avoid mutexes on the query and write path entirely. Locks exist only on single-owner background threads.
Every figure here is extracted from committed benchmark result files, alongside the full tables on the performance page.
Query path
A query enters as text and leaves as batches. The path runs in three stages, and each stage lives in its own crate.
A recursive descent parser with a Pratt expression core turns SQL text into a typed AST.
The cost-based optimizer reorders joins with dynamic programming, pushes predicates and projections into the scan, and decorrelates subqueries.
Vectorized, morsel-parallel operators move batches of rows through the plan instead of one row at a time.
Online DDL
A schema epoch in the tuple slot lets ALTER TABLE add, drop, or retype a column without rewriting a row. Index builds publish, wait for in-flight writers, scan, and flip without blocking the table, and an incompatible type change runs as a shadow rewrite behind live traffic.
A column change advances the epoch in the tuple slot and touches no row. An index build publishes, waits for in-flight writers, scans, and flips while the table stays open to writes.
Each tuple slot carries a schema epoch, so ALTER TABLE can add, drop, or retype a column without rewriting a row.
An index build publishes, waits for in-flight writers, scans, and flips. The table stays open to writes the whole way through.
An incompatible type change runs as a shadow rewrite behind live traffic, so reads and writes keep going while the rewrite lands.
On a Raft group, DDL runs on the connection so the client sees the real command tag, and the applier picks it up if the client drops. Every statement in the grammar replicates the same way, covered on the consensus page. The catalog rows behind a schema migrate at upgrade time through a schema evolution registry, on the upgrades page.
Grammar
The surface fills in with row expansion, pivots, joins keyed on time, and the statement forms around them. The SQL reference is generated from the grammar registry, so a statement without documentation fails the build instead of shipping.
UNNEST, FLATTEN, and UNPIVOT run as one expand-rows operator, with PIVOT alongside.
A join keyed on time order, and a time-series operator that fills the gaps in a series.
Temporary tables are node-local, and every statement form accepts a schema-qualified name.
Prepared statements, holdable cursors, COPY, and CANCEL BACKEND round out what a session can do.
The registry that defines the grammar is the source of the SQL reference. A statement that has no documentation fails the build, so the reference cannot drift from what the parser accepts.
Change streams add statement forms of their own, CREATE CHANGE STREAM and APPLY CHANGES, with the position advanced in the consumer's own transaction. They are covered on the streams page.
Legibility
One walk of the grammar writes both the reference pages and a dataset embedded in the binary, so the help a running server gives and the reference pages come from the same source.
HELP answers in every form, with docs_search over statements, functions, clauses, and prose.
VALIDATE checks a statement without running it, so a mistake surfaces before anything is touched.
Every error carries a stable error code, and the code rides on every transport.
A parse error carries the expected-token set, names the statement, and says where to read more.
Reference examples are rewritten against the caller's schema, so they name tables the caller actually has.
Plain-language recipes sit over the function registry, so the way to do something is written beside the functions that do it.
Deprecations get the same treatment, with time-bounded windows, rate-limited warnings, and auto-generated migration guides, covered on the upgrades page.
Columnar
Encoded columns are filtered without being decoded. The scan works directly on the compressed representation and decodes only what survives.
~11.9GB/sec
.zyr scan throughput
~36.8x
Metadata-aggregate pruning speedup
Bloom filters and zone maps ride on segment metadata, so scans rule out data before touching it.
Predicates run on encoded columns, and only the rows that survive are materialized.
SQL surface
Standard SQL, plus native extensions that go beyond it. The analytical surface is part of the engine, not an add-on.
Underneath it all is the PostgreSQL wire protocol v3 over both TCP and QUIC, so psql and existing drivers connect unchanged. Holdable cursors, COPY, and CANCEL BACKEND are part of that same surface, and every error that comes back carries a stable error code.
The full set of SQL highlights is in the README.
-- Read a table as it was at a point in time
SELECT * FROM orders AS OF TIMESTAMP '2026-05-06 09:00:00';
-- Branch the database, experiment in isolation, then merge or drop
CREATE BRANCH experiment FROM main;
USE BRANCH experiment;-- Native vector search
CREATE VECTOR INDEX ON docs (embedding) WITH (metric = 'cosine');
SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10;Storage tiers
Five tiers, each with one backing and one job. The engine moves rows between them so queries never have to.
| Tier | Backing | Purpose |
|---|---|---|
| B+ tree index | Resident in RAM, persisted via WAL + checkpoint | Point lookups, range scans |
| Row heap (OLTP) | Buffer pool with clock-sweep eviction | Recent and transactional rows |
| Columnar (.zyr) | Disk, hot segments cached in the buffer pool | Encoded analytical scans |
| ZyronLake (.zyr + log) | Object store or shared filesystem, versioned per commit, manifest-checkpointed | Analytical tables on shared storage. Branches, time travel, cross-engine reads |
| Write-ahead log | Append-only, group commit, fsync on commit | Durability and crash recovery |
| Raft log | Group-committed on the leader, quorum-replicated to followers, byte-aware residency cap with older entries paged from disk | Consensus record of DML and DDL for replication and failover |
The ZyronLake tier is a full table format of its own, covered on the lake page, and a single query can join any of these tiers to remote publications or a shared lake across the mesh. Every table also carries a change log that a consumer reads inside its own transaction, on the streams page.
Architecture
Zyron is a Cargo workspace where every crate depends only on the layers beneath it. A query enters at the wire and falls straight down the stack.
Clients
unchanged Postgres tooling
Connectivity
PostgreSQL v3, TCP + QUIC, TLS 1.3
Cluster
consensus and cross-node coordination
Orchestration
sessions, background workers, backup, load adaptation
Query path
cost-based, vectorized, morsel-parallel
Native subsystems
in the engine, not an app tier
Storage engine
heap, B+ tree, .zyr, MVCC, WAL, lake
Foundation
errors, pages, hashing, PRNG
The full architecture, storage tiers, and crate map are in the README.
Prebuilt binaries for Linux and Windows, no external dependencies, one command to a running server.