Combining Monocular Geometry and Semantic Segmentation for Structural Plane Extraction in Indoor Scenes
(2026) In Master’s Theses in Mathematical Sciences FMAM05 20261Mathematics (Faculty of Engineering)
- Abstract
- Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each... (More)
- Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each candidate region via a semantic majority vote.
The pipeline is evaluated on 34 annotated images from 21 unique indoor scenes using Intersection over Union (IoU), which measures overlap between predicted and ground-truth regions. The strongest backbone combination tested achieves a floor instance IoU of 0.945 and a wall instance IoU of 0.854 across all 34 images. Real scenes captured with Axis network cameras scored higher (0.975 floor, 0.900 wall) than the synthetic Hypersim subset (0.938 floor, 0.844 wall). Wall detection was more sensitive to backbone choice than floor detection. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9235070
- author
- Strand, Erik LU
- supervisor
- organization
- course
- FMAM05 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- publication/series
- Master’s Theses in Mathematical Sciences
- report number
- LUTFMA-3619-2026
- ISSN
- 1404-6342
- other publication id
- 2026:E33
- language
- English
- id
- 9235070
- date added to LUP
- 2026-06-15 13:04:23
- date last changed
- 2026-06-15 13:04:23
@misc{9235070,
abstract = {{Network cameras observe physical spaces but lack an inherent understanding of the scenes they see. This thesis presents a pipeline for floor and wall instance segmentation from a single uncalibrated RGB image, requiring no scene-specific prior information such as camera intrinsics, depth sensors, floor plans, or labelled floor and wall data for training. The detected instances are intended to support downstream applications such as room-boundary estimation, object detector validation, and camera placement verification. The method combines monocular geometry estimation from MoGe with semantic segmentation from Mask2Former, clustering surface normals by orientation, splitting co-oriented regions by projected depth, and classifying each candidate region via a semantic majority vote.
The pipeline is evaluated on 34 annotated images from 21 unique indoor scenes using Intersection over Union (IoU), which measures overlap between predicted and ground-truth regions. The strongest backbone combination tested achieves a floor instance IoU of 0.945 and a wall instance IoU of 0.854 across all 34 images. Real scenes captured with Axis network cameras scored higher (0.975 floor, 0.900 wall) than the synthetic Hypersim subset (0.938 floor, 0.844 wall). Wall detection was more sensitive to backbone choice than floor detection.}},
author = {{Strand, Erik}},
issn = {{1404-6342}},
language = {{eng}},
note = {{Student Paper}},
series = {{Master’s Theses in Mathematical Sciences}},
title = {{Combining Monocular Geometry and Semantic Segmentation for Structural Plane Extraction in Indoor Scenes}},
year = {{2026}},
}