Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Explicit Model Pre-training for 3D Occupancy Prediction

Brandin, Pontus LU and Mattsson, Albert Henrik LU (2026) In Master’s Theses in Mathematical Sciences FMAM05 20261
Mathematics (Faculty of Engineering)
Abstract
Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned... (More)
Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned representation is evaluated through binary occupancy prediction on Occ3D-nuScenes. Our results show that meaningful explicit 3D representations can be learned without vision foundation models at inference, and that fully self-supervised pre-training remains viable, although with reduced occupancy performance compared to VFM-assisted baselines. When adapted for semantic occupancy prediction, the proposed model achieves competitive results, including stronger mIoU than the original GaussTR baseline in our comparison. The findings suggest that geometric understanding and feature separation are closely linked. Improved geometry facilitates semantic learning, while stronger features support more accurate spatial representations. (Less)
Please use this url to cite or link to this publication:
author
Brandin, Pontus LU and Mattsson, Albert Henrik LU
supervisor
organization
course
FMAM05 20261
year
type
H2 - Master's Degree (Two Years)
subject
publication/series
Master’s Theses in Mathematical Sciences
report number
LUTFMA-3618-2026
ISSN
1404-6342
other publication id
E:32
language
English
id
9236227
date added to LUP
2026-06-12 10:59:40
date last changed
2026-06-12 10:59:40
@misc{9236227,
  abstract     = {{Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned representation is evaluated through binary occupancy prediction on Occ3D-nuScenes. Our results show that meaningful explicit 3D representations can be learned without vision foundation models at inference, and that fully self-supervised pre-training remains viable, although with reduced occupancy performance compared to VFM-assisted baselines. When adapted for semantic occupancy prediction, the proposed model achieves competitive results, including stronger mIoU than the original GaussTR baseline in our comparison. The findings suggest that geometric understanding and feature separation are closely linked. Improved geometry facilitates semantic learning, while stronger features support more accurate spatial representations.}},
  author       = {{Brandin, Pontus and Mattsson, Albert Henrik}},
  issn         = {{1404-6342}},
  language     = {{eng}},
  note         = {{Student Paper}},
  series       = {{Master’s Theses in Mathematical Sciences}},
  title        = {{Explicit Model Pre-training for 3D Occupancy Prediction}},
  year         = {{2026}},
}