JEPA World Model

Each camera casts 16 DINOv2 patch columns as rays through a spatial world model. The model predicts entrance locations even when they are occluded by trees, cars, or awnings in every available image. Trained on 1,668 POIs from Boulder County with image quality filtering. Evaluated against RTK ground truth: 1.73m MAE, 0.78m median, 52% better than facade midpoint baseline.

Context Encoder DINOv2 patches + facade_t PE
Target Encoder Facade geometry + entrance_t
Predictor z_visual → z_geo (AdaLN)
Entrance Head z_geo → t ∈ [0, 1]
4.2M params 300 epochs, ~9 min on A10G
JEPA predicted entrance
RTK ground truth (cm-level)
Ray hitting facade (≥4/16 required)
Ray missing facade
Camera position
cycle POIs   1-9 select image

World Model Entrance Prediction — Val MAE 1.73m | Median 0.78m | P90 4.70m

Facade Edge Position
Viewpoints