Explicit Model Pre-training for 3D Occupancy Prediction
(2026) In Master’s Theses in Mathematical Sciences FMAM05 20261Mathematics (Faculty of Engineering)
- Abstract
- Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned... (More)
- Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned representation is evaluated through binary occupancy prediction on Occ3D-nuScenes. Our results show that meaningful explicit 3D representations can be learned without vision foundation models at inference, and that fully self-supervised pre-training remains viable, although with reduced occupancy performance compared to VFM-assisted baselines. When adapted for semantic occupancy prediction, the proposed model achieves competitive results, including stronger mIoU than the original GaussTR baseline in our comparison. The findings suggest that geometric understanding and feature separation are closely linked. Improved geometry facilitates semantic learning, while stronger features support more accurate spatial representations. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9236227
- author
- Brandin, Pontus LU and Mattsson, Albert Henrik LU
- supervisor
- organization
- course
- FMAM05 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- publication/series
- Master’s Theses in Mathematical Sciences
- report number
- LUTFMA-3618-2026
- ISSN
- 1404-6342
- other publication id
- E:32
- language
- English
- id
- 9236227
- date added to LUP
- 2026-06-12 10:59:40
- date last changed
- 2026-06-12 10:59:40
@misc{9236227,
abstract = {{Understanding 3D scene structure from camera input is a central challenge in autonomous driving. Recent methods often rely on vision foundation models or annotated occupancy labels to provide strong semantic and geometric supervision, but these dependencies can limit scalability and introduce external biases. This thesis investigates whether an explicit 3D Gaussian scene representation can be pre-trained in a more self-contained manner from multi-view camera data and optional LiDAR supervision. Building on GaussTR, we replace foundation-model feature inputs with a ResNet backbone and replace depth supervision from Metric3D with self-supervised geometric consistency losses, optionally supported by LiDAR depth supervision. The learned representation is evaluated through binary occupancy prediction on Occ3D-nuScenes. Our results show that meaningful explicit 3D representations can be learned without vision foundation models at inference, and that fully self-supervised pre-training remains viable, although with reduced occupancy performance compared to VFM-assisted baselines. When adapted for semantic occupancy prediction, the proposed model achieves competitive results, including stronger mIoU than the original GaussTR baseline in our comparison. The findings suggest that geometric understanding and feature separation are closely linked. Improved geometry facilitates semantic learning, while stronger features support more accurate spatial representations.}},
author = {{Brandin, Pontus and Mattsson, Albert Henrik}},
issn = {{1404-6342}},
language = {{eng}},
note = {{Student Paper}},
series = {{Master’s Theses in Mathematical Sciences}},
title = {{Explicit Model Pre-training for 3D Occupancy Prediction}},
year = {{2026}},
}