Skip to main content

Lund University Publications

LUND UNIVERSITY LIBRARIES

GASP : Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving

Ljungbergh, William ; Lilja, Adam ; Tonderski, Adam LU orcid ; Ling, Arvid Laveno ; Lindstrom, Carl ; Verbeke, Willem ; Fu, Junsheng ; Petersson, Christoffer ; Hammarstrand, Lars and Felsberg, Michael (2026) 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 In Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026 p.3077-3087
Abstract

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy,... (More)

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see our project page.

(Less)
Please use this url to cite or link to this publication:
author
; ; ; ; ; ; ; ; and
organization
publishing date
type
Chapter in Book/Report/Conference proceeding
publication status
published
subject
host publication
Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
series title
Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
pages
11 pages
publisher
IEEE - Institute of Electrical and Electronics Engineers Inc.
conference name
2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026
conference location
Tucson, United States
conference dates
2026-03-06 - 2026-03-10
external identifiers
  • scopus:105041302017
ISBN
9798331555115
DOI
10.1109/WACV61042.2026.00301
language
English
LU publication?
yes
id
0461cf55-44f0-460f-a076-60bd001b9ea3
date added to LUP
2026-09-17 15:37:51
date last changed
2026-09-17 15:38:09
@inproceedings{0461cf55-44f0-460f-a076-60bd001b9ea3,
  abstract     = {{<p>Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied at scale. Similarly, autonomous driving generates vast amounts of spatiotemporal data, alluding to the possibility of harnessing scale to learn the underlying geometric and semantic structure of the environment and its evolution over time. In this direction, we propose a geometric and semantic self-supervised pre-training method, GASP, that learns a unified representation by predicting, at any queried future point in spacetime, (1) general occupancy, capturing the evolving structure of the 3D scene; (2) ego occupancy, modeling the ego vehicle path through the environment; and (3) distilled high-level features from a vision foundation model. By modeling geometric and semantic 4D occupancy fields instead of raw sensor measurements, the model learns a structured, generalizable representation of the environment and its evolution through time. We validate GASP on multiple autonomous driving benchmarks, demonstrating significant improvements in semantic occupancy forecasting, online mapping, and ego trajectory prediction. Our results demonstrate that continuous 4D geometric and semantic occupancy prediction provides a scalable and effective pre-training paradigm for autonomous driving. For code and additional visualizations, see our project page.</p>}},
  author       = {{Ljungbergh, William and Lilja, Adam and Tonderski, Adam and Ling, Arvid Laveno and Lindstrom, Carl and Verbeke, Willem and Fu, Junsheng and Petersson, Christoffer and Hammarstrand, Lars and Felsberg, Michael}},
  booktitle    = {{Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026}},
  isbn         = {{9798331555115}},
  language     = {{eng}},
  pages        = {{3077--3087}},
  publisher    = {{IEEE - Institute of Electrical and Electronics Engineers Inc.}},
  series       = {{Proceedings - 2026 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2026}},
  title        = {{GASP : Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving}},
  url          = {{http://dx.doi.org/10.1109/WACV61042.2026.00301}},
  doi          = {{10.1109/WACV61042.2026.00301}},
  year         = {{2026}},
}