You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

These weights are trained on NVIDIA PhysicalAI-AV data (TanitAD research program). Access is granted per request for research/evaluation use only; you agree not to redistribute.

Log in or Sign Up to review the conditions and access this model content.

TanitAD β€” REF-B v2 (arch-v2 --refbpatch, step 29,999 / 30,000 FINAL)

Reference arm B of the TanitAD 3-arm study: a hierarchical planner with a from-scratch trained encoder (the "trained-encoder" reference vs REF-A's frozen encoder). This is the v2 milestone-final checkpoint (271.6 M params).

Architecture (v2)

  • Trained ViT encoder: 9-channel, 256 px, patch 16, depth 25, d_model 768 (grad-checkpointed).
  • operative (d 768, depth 6, action_seq 5) + tactical + strategic (4 nav commands) + confidence head + OOD monitor.
  • v2 --refbpatch levers:
    • B1 β€” TIME-anchored tactical decoder (anchor space=time, 128 anchors from a 4096 pool, 4 layers / 8 heads) producing waypoints @ 5/10/15/20 steps (0.5/1/1.5/2 s).
    • B2 β€” proprioceptive ego conditioning [v0, yr0]: ego speed + backward-diff yaw-rate (raw rad/s), added to the nav embedding with ego_dropout 0.5 (anti-shortcut).
    • aux-yaw + aux-accel heads, path heads at fixed distances [2, 5, 10, 20] m, jerk penalty 0.02.

Training

  • Data: PhysicalAI-AV front-wide (phase-0 epcache).
  • Step 29,999 / 30,000 (FINAL). Optimizer AdamW (lr 3e-4, cosine), weight_decay 0.05, warmup 2000, AMP.

Evaluation (TanitEval, held-out PhysicalAI-AV val, n=881 windows; open-loop)

Planner-native trajectory = tactical waypoint heads (direct regression, nav=follow). Bootstrap 95% CI.

horizon ADE (m) CV baseline
0.5 s 0.103
1.0 s 0.217
1.5 s 0.379
2.0 s 0.592 Β± 0.069 0.825
  • fde@2s 1.23 (CV 1.71), miss@2m 0.203 (CV 0.313). Beats the constant-velocity baseline (0.592 vs 0.825, βˆ’28%).
  • By speed (model / CV, ADE@2s): high 0.43 / 0.65 Β· med 0.52 / 0.93 Β· low 0.82 / 0.93 β€” beats CV in every band.
  • By curvature (model / CV, ADE@2s): gentle 0.86 / 1.36 Β· sharp 0.90 / 2.38 (large margin) Β· straight 0.48 / 0.44 (CV is near-optimal on straights).
  • Error is ~79 % longitudinal at 2 s (along-track / speed-timing dominant); becomes lateral-dominant only on sharp curves.
  • Vision use: zeroing the camera frames worsens ADE@2s 0.591 β†’ 0.965 (β‰ˆ39 % of accuracy is vision-driven); the remainder rides the [v0, yr0] ego state.

Milestone / cross-arm A/B (paired, same windows)

  • vs REF-B v2 @20k (prior milestone): 0.591 vs 0.643 β€” v2-30k wins, significant (Ξ”CI [βˆ’0.077, βˆ’0.026]); largest gain on sharp curves (βˆ’0.135 m).
  • vs REF-B v1 (step-6k, ADE@2s 0.868): v2-30k wins decisively (Ξ”CI [βˆ’0.32, βˆ’0.22]); sharp-curve gain βˆ’1.19 m.
  • vs Flagship-4B v1 @30k (grounded world-model): flagship better on open-loop ADE (0.427 vs 0.591).

Generalization (cross-corpus, planner-direct)

corpus kind ADE@2s CV
PhysicalAI-AV in-dist 0.592 0.825
comma2k19 OOD (real highway) 0.515 0.646
Cosmos-Drive-Dreams OOD synthetic 0.495 0.451

The edge over CV shrinks OOD (in-dist βˆ’28 % β†’ comma βˆ’20 % β†’ cosmos +10 %, i.e. below CV). Cosmos numbers are PROVISIONAL (a subset of clips carry a ground-truth timing desync; a rebuilt cache is pending). Absolute OOD ADE is lower mainly because those corpora are geometrically easier (highway / synthetic, low curvature).

Imagination and hierarchy panels are N/A for REF-B (a direct planner with no grounded rollout head). Trained on PhysicalAI-AV derived data. Gated for research/eval use.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support