Benchmark methodology¶
How the figures on Performance were measured, what the published data
file contains, and how to measure your own machine. The benchmark suite lives in the
repository's benchmarks/ directory. It is not part of the installed package, so every
command below runs from a checkout.
The gate, precisely¶
The gate's method was written down in benchmarks/gate_preregistration.md before the first
measurement and not changed afterwards. In summary:
Workload. Events are generated before timing, all in bounds, one event per microsecond of
event time, t increasing across the whole stream; each call gets distinct events. Per call,
x and y are drawn uniformly over the sensor, or, for the clustered distribution, around 8
fixed centres with a Gaussian spread of 2% of the smaller sensor side; p is uniform over
{0, 1}. Seeds are fixed per resolution, batch size and distribution, and every run records a
SHA-256 of the bytes it fed: the runs of a cell must agree.
Kernel level times kernel.begin_call(state) plus kernel.accumulate(events, state,
watermark) per call, through the public kernel classes. Validation, range and bounds
checks, reading and result checks are outside the timed region. Statistic per run: events
per call divided by the median call time.
Engine level times Engine.ingest(events) per call, publication included when the Engine
decides on one. The Engine's clock is replaced, for its construction and each call, by a
virtual clock that reads the arrival time of the call's events at 20M events/s (0.5 ms per
10k call, 5 ms per 100k call, 50 ms per 1M call), so the Engine's own interval check makes
exactly the publications a live 20M events/s stream would cause. Statistic per run: events
divided by the sum of the call times, in wall-clock time.
Calls. One warm-up call per run, discarded. Timed calls per run: 50 / 20 / 7 at 10k / 100k
/ 1M events at kernel level; at Engine level the same at 0 ms, and at 16 ms enough calls for
at least 10 publications (320 / 40 / 10). The garbage collector is disabled during timed
calls; each call is one time.perf_counter_ns() interval.
Runs and classification. 5 runs per cell and level, each a separate Python process. The cell's value is the median of the 5 per-run values. At or above 22M events/s is a clear pass, below 18M a clear miss; in between, 20 more runs decide (median at or above 20M and at least 18 of the 20 at or above 20M). No outlier removal; no run discarded or repeated. A cell passes only if both levels pass; a runtime meets the gate only if all 150 cells pass on it.
Validity. A result counts only if every run passed its result checks (the final state equals a reference built with other NumPy primitives; at Engine level also the counters, the publication schedule, the sequence numbers and each publication's watermark), the runtime matched, the fed bytes matched, and the document came from a clean working tree on the reference machine.
Characterisation, not classification. Latency percentiles (nearest rank over the pooled
calls) and memory (tracemalloc in a separate untimed pass) are recorded for every cell but
never change a verdict.
Conditions of the published run. Commit 6a0fa27, clean tree. Apple M4 (4P + 6E), 16 GB,
macOS 15.7.7, mains power, Low Power Mode off. Runtime A: CPython 3.11.14, NumPy 2.4.6.
Runtime B: CPython 3.14.2 free-threaded with the GIL disabled, NumPy 2.4.6 (pinned by hand so
NumPy is the same on both runtimes).
The gate run predates the later sleep guard (below) and entered a low-power state near the end of the final 3.14t Engine-level run. A low-power state can only slow a run, so no pass verdict depends on it; in that run's last ten cells, median call times were 0.89-1.08x those of the other four runs. The recorded background environment was not controlled.
The data file¶
benchmarks/results/gate_v1.csv
holds the published gate run: 600 rows, one per cell, level and runtime (150 cells × 2
levels × 2 runtimes). It was derived from the run's result and verdict documents, with local
paths and machine-specific details left out.
| column | meaning |
|---|---|
runtime |
cpython-3.11.14 or cpython-3.14.2t-gil-disabled |
python, free_threaded_build, gil_enabled, numpy |
as recorded inside every run's own process |
commit |
the measured commit, 6a0fa27 |
kernel, kernel_parameters |
the kernel, and its parameter if it has one |
width, height |
the sensor size |
events_per_call, interval_ms |
the batch size and snapshot_interval_ms (at kernel level the interval doesn't apply; the 0 and 16 ms rows are separate measurements of the same work) |
distribution |
uniform or clustered |
level |
kernel or engine |
statistic |
median_call (kernel) or sustained (engine), as defined above |
events_per_s |
the cell's value: the median of the per-run statistics, events per second |
run_min_events_per_s, run_max_events_per_s |
the range of the 5 per-run values |
runs, timed_calls_per_run |
5, and the timed calls in each run |
latency_samples, latency_p50_us, latency_p95_us, latency_p99_us, latency_max_us |
per-call time over the pooled timed calls of all runs, microseconds |
peak_temporary_mib |
the largest tracemalloc peak temporary allocation of one call, MiB |
stage1 |
this level's classification: clear_pass, borderline or clear_miss |
level_result, cell_verdict |
this level's result, and the cell's verdict over both levels |
Measuring your own machine¶
From a checkout, with the development environment (uv sync), writing results outside the
repository:
RESULTS="${TMPDIR:-/tmp}/frames2py-gate" && mkdir -p "$RESULTS"
uv run python -m benchmarks list --suite gate # the 150 cells
uv run python -m benchmarks run --suite gate --target v1-kernel --out "$RESULTS/kernel.json"
uv run python -m benchmarks run --suite gate --target v1-engine --out "$RESULTS/engine.json"
uv run python -m benchmarks stage2 "$RESULTS/kernel.json" --out "$RESULTS/kernel_stage2.json"
uv run python -m benchmarks stage2 "$RESULTS/engine.json" --out "$RESULTS/engine_stage2.json"
uv run python -m benchmarks gate --kernel-level "$RESULTS/kernel.json" \
--engine-level "$RESULTS/engine.json" \
--kernel-stage2 "$RESULTS/kernel_stage2.json" \
--engine-stage2 "$RESULTS/engine_stage2.json" \
--out "$RESULTS/gate.json"
uv run python -m benchmarks report "$RESULTS/engine.json" # a table of one document
A result document records the commit and the working tree's state, and the gate does not
classify a document from a dirty tree, so keep the results out of the checkout. stage2
measures only the borderline cells, and nothing if there are none. An existing result file
is not overwritten unless you pass run the --overwrite flag.
For the free-threaded runtime, the published run used a separate environment synced from
the lockfile for CPython 3.14.2t, with NumPy then pinned to 2.4.6, and invoked its Python
directly (uv run would re-sync NumPy):
UV_PROJECT_ENVIRONMENT="$RESULTS/venv-314t" uv sync --frozen --python 3.14.2+freethreaded
uv pip install --python "$RESULTS/venv-314t/bin/python" numpy==2.4.6
"$RESULTS/venv-314t/bin/python" -m benchmarks run --suite gate --target v1-kernel --out "$RESULTS/314t_kernel.json"
On any machine other than an Apple M4 with 16 GB, gate reports NOT ON THE REFERENCE
MACHINE instead of a verdict: the published gate is tied to that machine, and a result on
another machine is a measurement of that machine, not a re-run of the gate. The per-cell
numbers are still comparable with the CSV.
To measure a subset, run takes --kernel, --resolution WxH, --batch-size, --interval
and --distribution, each repeatable. The other suites are adapters, recorder, viewer
and replay (uv run python -m benchmarks --help).
Hygiene that mattered¶
On this machine, differences under about 10% between measurements of the same configuration were not treated as meaningful, and the suite never compares single runs. Conditions that changed results enough to matter:
- System sleep. The Mac used idle-sleeps after a minute and can also run in a low-power state; runs that overlapped either were slow. Since the published gate run, the runner holds its own idle-sleep assertion on macOS, refuses to start outside full wake, records any sleep during the run, and the gate treats cells from a run that slept as invalid.
- Background load. Other work on the machine competes for CPU. The published gate run recorded the machine's load but did not control it, so read small differences between cells as noise.
- Separate processes. Repeated runs inside one process measured slower than the first, so each run is its own process by default.