Recorder¶
frames2py.recorder writes the events you give it to an HDF5 file that
frames2py.adapters.hdf5 reads back unchanged. It records events, not frames:
what you get back is the event stream, which you can feed to any kernel afterwards. It is a
sink your own loop calls, next to Engine.ingest(); the Engine knows nothing about it.
Install frames2py[recorder] (h5py and hdf5plugin, the same backends as frames2py[hdf5]).
import frames2py.recorder imports neither; recorder.open() without them raises
ImportError naming the extra.
Record, then read back¶
# Needs frames2py[recorder]. Reading back with frames2py.adapters.hdf5 needs the same two libraries
# (h5py and hdf5plugin), so frames2py[recorder] alone is enough for this example.
import tempfile
from pathlib import Path
import numpy as np
import frames2py
from frames2py import recorder
from frames2py.adapters import evt, hdf5
with tempfile.TemporaryDirectory() as directory:
path = Path(directory) / "session.h5"
engine = frames2py.Engine((640, 480), "event_count")
written = []
with evt.open("sparklers_100k.evt2.raw", sensor_size=(640, 480)) as reader, \
recorder.open(path, sensor_size=(640, 480)) as rec:
for events in reader:
rec.write(events) # compression and file I/O on this thread
engine.ingest(events) # the Engine never calls the recorder
written.append(events)
engine.stop()
with hdf5.open(path, group="events", sensor_size=(640, 480)) as reader:
read_back = np.concatenate(list(reader))
original = np.concatenate(written)
print("events recorded:", len(read_back))
print("same dtype:", read_back.dtype == original.dtype == frames2py.EVENT_DTYPE)
print("same values, in order:", all(np.array_equal(original[f], read_back[f]) for f in "txyp"))
The API is recorder.open(path, *, sensor_size, group="events", compression="blosc",
overwrite=False), which returns a recorder with write(events) and close(), used as a
context manager. Full signatures: API reference.
What is stored¶
A recording is one HDF5 file:
/events a group (the `group` argument; "/" puts it at the root)
t uint64 µs, absolute, as given
x uint16
y uint16
p uint8 as given: 0 is OFF, anything else ON
attributes (int64):
sensor_width from sensor_size
sensor_height
frames2py_format_version 1
The four datasets are 1-D, equally long, one element per event, in the order write() got
them, chunked 65,536 events at a time. There is no t_offset: t is stored as the absolute
microsecond value you passed in. Nothing else is in the file.
Events are recorded exactly as given. Out-of-order and backward timestamps, events
outside sensor_size and repeated events are all kept; nothing is sorted, clamped,
deduplicated or repaired. Timestamp discontinuities stay the caller's responsibility, as for
the Engine.
Validation. write() checks events the way Engine.ingest() does, with the same code.
A malformed array (wrong dtype for t, x, y or p, byte order included, not 1-D, not
C-contiguous) raises TypeError; extra fields are ignored. Any event with t >= 2**63
rejects the whole call with ValueError, and nothing of that call is recorded; earlier and
later calls are unaffected.
Reading it back takes hdf5.open(path, group="events"): the reader's group has no
default. The reader doesn't take the geometry from the file either, so pass
sensor_size=... as well if you want reader.sensor_size (for example to build an Engine
from it); without it, reader.sensor_size is None.
The attributes. sensor_width and sensor_height record the geometry you declared.
They are informative: the HDF5 adapter doesn't read them. They are not checked against the
events either. frames2py_format_version is the
version of this layout, 1; the HDF5 adapter refuses a group whose version isn't exactly 1.
Compression¶
compression |
filters | who can read it | measured cost |
|---|---|---|---|
"blosc" (default) |
Blosc, LZ4, level 5, byte shuffle, through hdf5plugin | needs the Blosc plugin | fast, about 3 bytes per event |
"gzip" |
gzip level 4 and the shuffle filter, built into HDF5 | any HDF5 reader | slowest, about 2 bytes per event |
None |
none | any HDF5 reader | fastest, 13 bytes per event |
Blosc is the default because it was the one option both compact (about a quarter of the raw 13 bytes per event) and fast enough to keep up with the recordings measured (Performance). The cost is portability: a plain HDF5 install doesn't have the Blosc filter. From Python, importing hdf5plugin registers it:
# Sketch (not runnable): needs a recording called session.h5.
import h5py
import hdf5plugin # noqa: F401 registers Blosc with HDF5
with h5py.File("session.h5") as f:
t = f["events/t"][:]
Outside Python, HDF5 loads filter plugins from the directories in HDF5_PLUGIN_PATH, and
hdf5plugin ships Blosc as such a plugin:
That was checked with h5py reading a Blosc recording without importing hdf5plugin; not with
h5dump, HDFView or other tools. Use compression="gzip" when the file has to be read by
tools you don't control.
Identical bytes¶
The recorder buffers events until it has a whole chunk of 65,536 and writes each chunk once.
The bytes of a recording therefore depend only on the events and the open() arguments, not
on how the events were split across write() calls: one call with a million events and a
million calls with one event each give the same file. The test suite checks this for each
compression across several irregular splits. The buffer holds at most one chunk (65,535
events, about 0.85 MB); the last partial chunk is written at close().
Finishing a recording¶
The recorder writes to a hidden temporary file next to the target
(.session.h5.<random>.partial) and moves it onto the target with os.replace() only once
the recording has been finished and closed. A file at the target path is therefore always a
complete, closed recording.
close(), or leaving thewithblock normally: the buffered events are written, the file is closed, and it replaces the target.- An exception leaving the
withblock,KeyboardInterruptincluded: the recording is finished the same way and becomes the output. Ctrl-C is a normal way to end a live recording; the file then holds every event of everywrite()that completed. If the interrupt lands inside awrite(), the part of that call's events already written stays in the recording, possibly none of it: the recording is always a prefix of everything passed towrite(). A rejected call (malformed, ort >= 2**63) is different: nothing of it is buffered or written. overwrite=False(the default):open()raisesFileExistsErrorif the target exists. The check is repeated just before the final move; if a file appeared at the target in the meantime,close()raisesFileExistsErrorand leaves the finished recording at the temporary path its message names.overwrite=True: the old target stays untouched while recording and is replaced when the new recording is finished, an interrupted one included.- If finishing fails (the disk fills while the last chunk is written, or a second
Ctrl-C lands inside
close()), the temporary file is removed and the error raised; the target is unchanged.
A crash or power loss is not covered. If the process is killed or the machine loses
power, the target is untouched (or absent) and a .partial file may be left behind. It is
not a valid recording: HDF5 files that were never closed generally can't be opened, and
Frames2Py makes no attempt to recover them. There is no journal, no SWMR mode and no fsync;
whether the moved file has reached the disk is up to the operating system.
Other errors from open(): IsADirectoryError for a directory, FileNotFoundError if the
target's directory doesn't exist, PermissionError if it can't be written; TypeError and
ValueError for bad arguments, before any file is created. write() after close()
raises ValueError.
Threads and blocking¶
write() does its work on the thread that calls it: it validates, copies into the buffer,
and whenever a chunk is full, compresses and writes it. So write() blocks its caller on
compression and file I/O. The recorder starts no thread and has no queue, and the Engine
never calls it or waits for it.
The consequence: if one thread both records and ingests, as in the example, that loop
sustains at most what the two together sustain. The measured rates are on the
Performance page. If the recorder's share matters,
call write() from a thread of your own; that thread and its hand-off are yours, not
Frames2Py's.