HuyMichel/EventActivityNet-bucket / docs /DATASET_GENERATION.md
HuyMichel's picture
|
download
raw
3.2 kB

EventActivityNet Dataset Generation

Source Video Lineage

Both representations were generated from the non-interpolated/original-rate ActivityNet video lineage. Source frame rate varies by video, no source video is resampled or interpolated, and source spatial resolution is preserved.

Large Subset Curation

The canonical Large set contains 3,263 unique videos. The recovered curation implementation:

  • merged ActivityNet Captions train and validation metadata before sampling;
  • used seed 2025;
  • used initial class-balanced sampling with max(5, int(class_ratio * class_count)) and class_ratio=0.2;
  • used 33% and 66% duration quantiles;
  • enriched event-friendly examples using caption keywords or first-frame darkness.

The keyword list is run, fast, sprint, night, dark, and slow-motion; the darkness rule is normalized first-frame mean brightness below 0.4.

HDF5 Generation

The recovered implementation follows:

  1. activitynet.py video loading;
  2. mp4_to_h5.mp4_to_h5_stream();
  3. EventEmulatorGPU.video_to_voxel();
  4. temporal grouping;
  5. HDF5 writing.
Parameter 5-bin 9-bin
num_bins 5 9
frames_per_bin 1 1
Source resizing none none
events dtype int16 int16
Start/count dtypes int64 / int32 int64 / int32
Event compression gzip+shuffle gzip+shuffle
Auxiliary compression LZF+shuffle LZF+shuffle
Learned checkpoint none none

The canonical 9-bin lineage was generated with a frozen deterministic configuration using RNG seed 42. The emulator creates and seeds its generator for each streaming batch while carrying the previous frame and returned potential across batches. This is recorded as technical reproducibility metadata; it is not a comparative quality claim.

Generic Grouping Semantics

For bin count B and N decoded source frames:

T_B = ceil((N - 1) / B)

Each generated event slice represents one adjacent-frame transition. events[t] groups up to B consecutive slices; voxel_event_start[t] = B*t; and voxel_event_count[t] records the valid slice count. The final group is zero-filled beyond its valid count.

Timing Basis

Seconds-level conversion uses each video's released rational nominal or average frame rate. Construction follows decoded frame order and does not consume per-frame presentation timestamps, so mapping is approximate for within-video variable-frame-rate streams. No fixed 25-fps or 240-fps assumption should be used.

Release Scales

Large is the recovered historical set. Medium and Small are deterministic nested v1.0 scales selected with seed 2025 and strata (split, class_label, duration_bucket, event_friendly). Small is a strict subset of Medium, and Medium is a strict subset of Large. Both event voxel representations share these scale manifests.

Validation

Final validation confirmed exactly 3,263 readable files per representation, the same 2,316/947 split assignment, expected tensor length and metadata arrays, and zero remaining structural failures. Canonical HDF5 payload sizes are 4,355,745,895,245 bytes for 5-bin and 4,214,122,096,103 bytes for 9-bin.

Xet Storage Details

Size:
3.2 kB
·
Xet hash:
f29f48f26bfcd651a0fee9ffc9221b301ceb7e910bdc1f0be301ea987359a741

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.