Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Combining Monocular Geometry and Semantic Segmentation for Structural Plane Extraction in Indoor Scenes

Strand, Erik LU (2026) In Master’s Theses in Mathematical Sciences FMAM05 20261
Mathematics (Faculty of Engineering)
Abstract
Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each... (More)
Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each candidate region via a semantic majority vote.

The pipeline is evaluated on 34 annotated images from 21 unique indoor scenes using Intersection over Union (IoU), which measures overlap between predicted and ground-truth regions. The strongest backbone combination tested achieves a floor instance IoU of 0.945 and a wall instance IoU of 0.854 across all 34 images. Real scenes captured with Axis network cameras scored higher (0.975 floor, 0.900 wall) than the synthetic Hypersim subset (0.938 floor, 0.844 wall). Wall detection was more sensitive to backbone choice than floor detection. (Less)
Please use this url to cite or link to this publication:
author
Strand, Erik LU
supervisor
organization
course
FMAM05 20261
year
type
H2 - Master's Degree (Two Years)
subject
publication/series
Master’s Theses in Mathematical Sciences
report number
LUTFMA-3619-2026
ISSN
1404-6342
other publication id
2026:E33
language
English
id
9235070
date added to LUP
2026-06-15 13:04:23
date last changed
2026-06-15 13:04:23
@misc{9235070,
  abstract     = {{Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each candidate region via a semantic majority vote.

The pipeline is evaluated on 34 annotated images from 21 unique indoor scenes using Intersection over Union (IoU), which measures overlap between predicted and ground-truth regions. The strongest backbone combination tested achieves a floor instance IoU of 0.945 and a wall instance IoU of 0.854 across all 34 images. Real scenes captured with Axis network cameras scored higher (0.975 floor, 0.900 wall) than the synthetic Hypersim subset (0.938 floor, 0.844 wall). Wall detection was more sensitive to backbone choice than floor detection.}},
  author       = {{Strand, Erik}},
  issn         = {{1404-6342}},
  language     = {{eng}},
  note         = {{Student Paper}},
  series       = {{Master’s Theses in Mathematical Sciences}},
  title        = {{Combining Monocular Geometry and Semantic Segmentation for Structural Plane Extraction in Indoor Scenes}},
  year         = {{2026}},
}