Performance

Benchmarked in the open

47 benchmark suites run against the real server over the wire. Every number on this page is extracted from result files committed to the repository and regenerated by script, trade-offs included.

GROUP COMMIT · MANY TRANSACTIONS, ONE FSYNCClient 1Client 2Client 3Client 4Client 5Client 6Client 7Client 8Commit batchWAL deviceOne fsyncFor the whole batchConcurrent transactions share one durable writeLatency stays flat while the batch grows

Performance

Numbers from committed benchmarks, not a lab

Two views of the same engine, what a client gets from a running server over the wire, and the raw subsystem numbers underneath it.

OLTP throughput by concurrent clients

Thousand transactions per second over the wire, higher is better.

141664256

concurrent clients

View the numbers
ClientsThroughputp99 latency
113.6K tps234 µs
439K tps283 µs
1665.6K tps476 µs
6462.9K tps1.7 ms
25655K tps8.2 ms

1.8Btuples/s

MVCC GC sweep

11.9GB/s

.zyr scan throughput

12.3Mrows/s

Compaction pipeline

1816Mrows/s

Picosecond timestamp decode

69.9µs

Durable commit floor

778Ktxn/s

Durable group commit peak

216ms

Leader election after a kill

99.8%

Follower keep-up vs leader

The trade-offs are published too. Bulk load, trickle load, and indexed point lookups are cases where the row heap beats the lake, and they are in the same tables. See benchmarks/ for the raw result files.

Tail latency under load

p99 stays under half a millisecond through 16 concurrent clients, then grows as saturation sets in. The full curve is published, not just the flattering part.

OLTP p99 latency by concurrent clients

Milliseconds over the wire, lower is better.

141664256

concurrent clients

View the numbers
Clientsp99 latency
1234 µs
4283 µs
16476 µs
641.7 ms
2568.2 ms

Durability is not the bottleneck

Every one of those transactions was durably committed, the WAL is fsynced through group commit before the client sees an acknowledgement. Group commit amortizes the device write across concurrent transactions instead of skipping it.

~778Ktxn/s

Durable group commit peak

~69.9µs

Durable commit floor, device write

~83.0x

Group commit amplification, c=1 to c=512

The durability model is covered on the database page.

End to end, cold start to shutdown

What a client actually experiences against a running server, not a microbenchmark.

35 ms

Cold boot to accepting queries

0.77 ms

First ReadyForQuery

11.2 ms

Schema DDL bootstrap

237K rows/sec

Seed insert

0.57 ms

Analytical query, median

64 ms

Graceful shutdown

Consensus and replication

A three-node Raft group with the leader accepting writes, followers applying committed entries through the same operator path the leader used, and reads served through a ReadIndex round trip so they stay linearizable. How the group works is on the consensus page.

~216ms

Leader election after a kill

~0.19µs

Single log append

~99.8%

Follower keep-up vs the leader

~97entries

Worst follower lag

~1.13s

Snapshot 1GB, create

~2.72s

Snapshot 1GB, transfer

Row heap vs ZyronLake, honestly

Same workload, same rows, both formats. The write-heavy and indexed-point cases where the heap wins are published as numbers, the scan-side cases where the lake wins are in the repository's cross-format charts.

MetricResultFavors
Point lookup with a heap B+ tree index, lake vs heap~0.6x, heap wins indexed pointsRow heap
Bulk load to queryable, lake vs heap~1.5x, heap wins large batchesRow heap
Trickle load to queryable, lake vs heap~9.7x, heap wins tiny commitsRow heap

Engine internals, the full table

Raw subsystem throughput and hot-path latency under microbenchmark, straight from the committed result files.

SubsystemMetricResult
MVCCGC sweep~1.8B tuples/sec
Columnar.zyr scan throughput~11.9 GB/sec
ColumnarCompaction pipeline~12.3M rows/sec
ColumnarHybridScan overhead vs heap-only~1.3%
ColumnarMetadata-aggregate pruning speedup~36.8x
TemporalPicosecond timestamp decode~1816M rows/sec
VersioningTime-travel scan overhead~24%
WireQUIC connection handshake~5 µs
TransactionsDurable commit floor, device write~69.9 µs
TransactionsDurable group-commit peak~778K txn/sec
TransactionsGroup-commit amplification, c=1 to c=512~83.0x
LakeCommit rate, insert~2403 commits/sec
LakeCommit latency, insert~2.40 ms
LakeCommit rate, delete predicate~1972 commits/sec
LakeDerived clustering expression files pruned~93%
LakeLoad with clustering expression~739K rows/sec
ConsensusLeader election after a kill~216 ms
ConsensusSingle log append~0.19 µs
ConsensusSnapshot 1GB, create~1.13 s
ConsensusSnapshot 1GB, transfer~2.72 s
ReplicationFollower keep-up vs leader~99.8%
ReplicationWorst follower lag~97 entries

Methodology

All figures come from a release build on a single machine, an Intel Core Ultra 7 270K Plus with 24 cores and 31.4 GB RAM on windows/x86_64. End-to-end suites talk to the real server over the wire, internals suites microbenchmark one subsystem at a time.

The 47 suites cover storage, executor, optimizer, encoding, wire, search, analytics, CDC, versioning, transactions, temporal, columnar, lake, cross-format, raft, replication, types, lifecycle, gateway, Zyron-to-Zyron, and end-to-end. Each run writes a timestamped JSON and TXT pair under benchmarks/, and the tables and charts in the README are regenerated from those files by script. Nothing on this page is hand-entered.

Anyone can rerun them. The suites are cargo tests, so reproducing a number takes one command.

Reproduce a suite

# Run one suite, results land in benchmarks/<suite>/
cargo test -p zyron-search --test search_bench --release -- --nocapture

# Regenerate the README charts from the latest result files
python scripts/gen_bench_charts.py

Running in the next five minutes

Prebuilt binaries for Linux and Windows, no external dependencies, one command to a running server.