Handing snapshots to PyTorch¶
Frames2Py doesn't depend on PyTorch, doesn't import it, and has no PyTorch API: no extra, no module, no tensor type. A consumer that wants tensors converts snapshots itself. This page is the recipe for doing that without touching the frame other consumers share:
engine.snapshot() → snapshot.copy() → torch.from_numpy(...) → .to(dtype) → .to(device)
shared, read-only yours shares the copy as needed copies
# Needs PyTorch, installed by you: Frames2Py neither depends on it nor imports it.
import numpy as np
import torch
import frames2py
events = np.zeros(6, dtype=frames2py.EVENT_DTYPE)
events["t"] = [3, 8, 14, 21, 27, 33]
events["x"] = [0, 1, 2, 3, 1, 2]
events["y"] = [0, 0, 1, 2, 1, 2]
events["p"] = [1, 0, 1, 1, 0, 1]
# A 2-D frame: counts, (H, W) uint32.
engine = frames2py.Engine((4, 3), "event_count", snapshot_interval_ms=0)
engine.ingest(events)
snapshot = engine.snapshot()
counts = torch.from_numpy(snapshot.copy()) # copy first: the tensor shares the copy's memory, which is yours
print("counts:", tuple(counts.shape), counts.dtype)
counts = counts.to(torch.int64) # torch has few uint32 ops; int64 holds every uint32 exactly
counts += 1 # yours to change: the published frame doesn't see it
print("published frame unchanged:", int(snapshot.frame.sum()) == len(events))
# The voxel grid: (bins, H, W) float32, time-first. As NCHW, the bins are the channels.
engine = frames2py.Engine((4, 3), frames2py.VoxelGrid(bins=3, bin_us=10), snapshot_interval_ms=0)
engine.ingest(events)
voxels = torch.from_numpy(engine.snapshot().copy())
batch = voxels.unsqueeze(0) # (1, bins, H, W), a view: nothing transposed
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = torch.nn.Conv2d(in_channels=3, out_channels=8, kernel_size=3, padding=1).to(device)
with torch.no_grad():
features = model(batch.to(device)) # .to() copies when the device changes
print("voxels:", tuple(voxels.shape), voxels.dtype, "-> features:", tuple(features.shape))
# The histogram: (2, bins, H, W) uint32. Merging polarity and time into 2 * bins channels is a
# reshape, polarity-major: channel p * bins + b is frame[p, b].
engine = frames2py.Engine((4, 3), frames2py.StackedHistogram(bins=3, bin_us=10), snapshot_interval_ms=0)
engine.ingest(events)
histogram = torch.from_numpy(engine.snapshot().copy())
channels = histogram.reshape(1, 2 * 3, 3, 4).to(torch.float32) # exact while every count is below 2**24
print("histogram:", tuple(histogram.shape), histogram.dtype, "-> channels:", tuple(channels.shape), channels.dtype)
# Without an allocation per read: one buffer of your own, refilled in place. The tensor is that
# buffer, so each refill changes it; clone() what you need to keep.
engine = frames2py.Engine((4, 3), "event_count", snapshot_interval_ms=0)
buffer = np.zeros((3, 4), dtype=np.uint32)
frame = torch.from_numpy(buffer)
for batch_of_events in (events[:2], events[2:]):
engine.ingest(batch_of_events)
engine.snapshot().copy(out=buffer)
print("window total:", int(frame.sum()))
counts: (3, 4) torch.uint32
published frame unchanged: True
voxels: (3, 3, 4) torch.float32 -> features: (1, 8, 3, 4)
histogram: (2, 3, 3, 4) torch.uint32 -> channels: (1, 6, 3, 4) torch.float32
window total: 2
window total: 4
Who owns what¶
- Frames2Py owns the published frame.
snapshot.frameis shared by every consumer that reads that publication, and Frames2Py never writes it again (Snapshots). snapshot.copy()gives you a new NumPy array, writable and C-contiguous, that nothing else references.torch.from_numpy()doesn't copy. The tensor uses the array's memory and keeps the array alive for as long as the tensor lives. A write through either one shows in the other.- So the tensor is your copy. Change it, in place or otherwise, and no other consumer sees it. Frames2Py never sees it either: it keeps no reference to the copy.
- With
copy(out=buffer)the tensor is your buffer. Each refill changes the tensor's values, as the last part of the example shows;clone()a tensor you need to keep, and don't refill the buffer while other code may still be reading the tensor.
Don't convert snapshot.frame itself¶
torch.from_numpy(snapshot.frame) succeeds, but it is not a read-only handoff. PyTorch
doesn't support read-only tensors: the tensor shares the published frame's memory and is
writable. A write
through it, an in-place operation included, changes the frame every other consumer reads,
even though NumPy still reports the frame as read-only. PyTorch emits a UserWarning ("The
given NumPy array is not writable...") for the first such conversion in a process only.
torch.as_tensor(snapshot.frame) without a dtype change shares the memory the same way.
DLPack doesn't help. NumPy marks a read-only array's DLPack export as read-only, but the
tensor torch.from_dlpack(snapshot.frame) returns is writable, shares the frame's memory,
and comes without a warning.
# Sketch (not runnable): what the recipe avoids
alias = torch.from_numpy(snapshot.frame) # shares the published frame: one warning per process
alias = torch.from_dlpack(snapshot.frame) # shares it too, writable, no warning
alias.zero_() # every consumer of this publication now reads zeros
A consumer that only reads could use either call, but nothing would stop a later in-place operation from writing through. Frames2Py makes no claim that a tensor sharing the published frame is safe. Copy first.
What copies, and what doesn't¶
| step | copies? |
|---|---|
snapshot.copy(), snapshot.copy(out=buffer) |
yes: the whole frame, into a new array or your buffer |
torch.from_numpy(array) |
no: the tensor shares array's memory |
.unsqueeze(), .reshape() of a contiguous tensor, .permute() |
no: views |
.contiguous() after .permute() |
yes |
.to(dtype) |
yes, when the dtype changes; otherwise it returns the same tensor |
.to(device) |
yes, when the device changes; otherwise it returns the same tensor |
The one copy the recipe always makes is snapshot.copy(). None of this is free, and none of
it happens on the producer's thread: conversion is consumer work.
Shapes, dtypes and layouts¶
torch.from_numpy keeps the frame's shape, its memory layout (C-contiguous) and its dtype.
| kernel | frame | tensor | before arithmetic |
|---|---|---|---|
event_count |
(H, W) uint32 |
torch.uint32 |
.to(torch.int64) or a float type |
polarity |
(H, W, 2) uint32 |
torch.uint32 |
as above; .permute(2, 0, 1) for channels first |
time_surface |
(H, W) uint64 |
torch.uint64 |
.to(torch.int64) |
exp_decay |
(H, W) float32 |
torch.float32 |
none |
timestamp_decay |
(H, W) float32 |
torch.float32 |
none |
StackedHistogram |
(2, bins, H, W) uint32 |
torch.uint32 |
.to(torch.int64) or a float type |
VoxelGrid |
(bins, H, W) float32 |
torch.float32 |
none |
Unsigned integers: convert first. Most uint32 arithmetic is unsupported in PyTorch;
convert first. PyTorch's documentation says unsigned types other than uint8 "are currently
planned to only have limited support in eager mode"
(Tensor Attributes, PyTorch 2.14).
.to(torch.int64)is exact for everyuint32count. It is exact for everytime_surfacevalue too, because the event contract rejects timestamps of 2^63 and above..to(torch.float32)is exact only up to 2^24 (16,777,216). Above that, not every integer has a float32: 16,777,217 becomes 16,777,216, and 4,294,967,295 becomes 4,294,967,296. Counts that stay below 2^24 convert exactly. Microsecond timestamps pass 2^24 after about 16.8 seconds, so a time surface converted straight to float32 loses precision; subtract a reference time in int64 first if your model wants small floats..to(torch.float64)is exact for everyuint32.
Counts wrap modulo 2^32 in Frames2Py (Kernels), so a converted count is the wrapped value.
Layouts stay as Frames2Py publishes them. The temporal kernels are time-first, and the recipe doesn't transpose them:
VoxelGrid(bins, H, W)is already channels-first.unsqueeze(0)gives the NCHW batch(1, bins, H, W)thattorch.nn.Conv2d(in_channels=bins, ...)takes, without a copy.StackedHistogram(2, bins, H, W):reshape(1, 2 * bins, H, W)merges polarity and time into channels, polarity-major, so channelp * bins + bisframe[p, b]. A model that expects another channel order needs its own permute.polarity(H, W, 2)is channels-last.permute(2, 0, 1)gives(2, H, W), channel 0 OFF and 1 ON, as a non-contiguous view;.contiguous()copies it if a model needs a contiguous tensor.- 2-D frames
(H, W)become(1, 1, H, W)with[None, None].
Devices¶
.to(device) is PyTorch's: it copies the tensor to the device when the device differs and
returns the same tensor when it doesn't. Frames2Py's frames live in CPU memory, and
Frames2Py makes no GPU claim. In CI, the test that a device transfer copies runs where PyTorch
reports a device: Apple's MPS on the macOS runners. The Linux runners have none and skip it.
No CUDA device is tested.
Tested versions¶
The recipe's tests and the example above run in a separate CI workflow
(.github/workflows/torch.yml), apart from the main CI, which never installs PyTorch:
- torch 2.14.1, from PyTorch's CPU wheel index (
https://download.pytorch.org/whl/cpu; on macOS that index serves the standard macOS wheel); - CPython 3.11, and 3.14t with the GIL disabled, on Linux x86_64, Linux ARM64 and macOS ARM64.
In those 3.14t environments, importing torch 2.14.1 did not re-enable the GIL: the job checks the GIL after the import and at the start and end of the test run. That is the whole finding. It is not a statement about PyTorch on free-threaded Python in general.
Other PyTorch versions, CPython 3.12 to 3.14, and CUDA builds of PyTorch are not tested.