Buckets:
EventActivityNet Dataset Generation
Source Video Lineage
Both representations were generated from the non-interpolated/original-rate ActivityNet video lineage. Source frame rate varies by video, no source video is resampled or interpolated, and source spatial resolution is preserved.
Large Subset Curation
The canonical Large set contains 3,263 unique videos. The recovered curation implementation:
- merged ActivityNet Captions train and validation metadata before sampling;
- used seed
2025; - used initial class-balanced sampling with
max(5, int(class_ratio * class_count))andclass_ratio=0.2; - used 33% and 66% duration quantiles;
- enriched event-friendly examples using caption keywords or first-frame darkness.
The keyword list is run, fast, sprint, night, dark, and
slow-motion; the darkness rule is normalized first-frame mean brightness
below 0.4.
HDF5 Generation
The recovered implementation follows:
activitynet.pyvideo loading;mp4_to_h5.mp4_to_h5_stream();EventEmulatorGPU.video_to_voxel();- temporal grouping;
- HDF5 writing.
| Parameter | 5-bin | 9-bin |
|---|---|---|
num_bins |
5 | 9 |
frames_per_bin |
1 | 1 |
| Source resizing | none | none |
events dtype |
int16 |
int16 |
| Start/count dtypes | int64 / int32 |
int64 / int32 |
| Event compression | gzip+shuffle | gzip+shuffle |
| Auxiliary compression | LZF+shuffle | LZF+shuffle |
| Learned checkpoint | none | none |
The canonical 9-bin lineage was generated with a frozen deterministic
configuration using RNG seed 42. The emulator creates and seeds its generator
for each streaming batch while carrying the previous frame and returned
potential across batches. This is recorded as technical reproducibility
metadata; it is not a comparative quality claim.
Generic Grouping Semantics
For bin count B and N decoded source frames:
T_B = ceil((N - 1) / B)
Each generated event slice represents one adjacent-frame transition.
events[t] groups up to B consecutive slices;
voxel_event_start[t] = B*t; and voxel_event_count[t] records the valid
slice count. The final group is zero-filled beyond its valid count.
Timing Basis
Seconds-level conversion uses each video's released rational nominal or average frame rate. Construction follows decoded frame order and does not consume per-frame presentation timestamps, so mapping is approximate for within-video variable-frame-rate streams. No fixed 25-fps or 240-fps assumption should be used.
Release Scales
Large is the recovered historical set. Medium and Small are deterministic
nested v1.0 scales selected with seed 2025 and strata
(split, class_label, duration_bucket, event_friendly). Small is a strict
subset of Medium, and Medium is a strict subset of Large. Both event voxel
representations share these scale manifests.
Validation
Final validation confirmed exactly 3,263 readable files per representation, the same 2,316/947 split assignment, expected tensor length and metadata arrays, and zero remaining structural failures. Canonical HDF5 payload sizes are 4,355,745,895,245 bytes for 5-bin and 4,214,122,096,103 bytes for 9-bin.
Xet Storage Details
- Size:
- 3.2 kB
- Xet hash:
- f29f48f26bfcd651a0fee9ffc9221b301ceb7e910bdc1f0be301ea987359a741
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.