4D RECONSTRUCTION IN SPARSE CAMERA SETTINGS FOR VR
(2026) MAMM15 20261Department of Design Sciences
Ergonomics and Aerosol Technology
- Abstract
- This thesis explores 4D reconstruction in sparse fixed-camera settings for virtual reality visualization. Gaussian Splatting provides an efficient representation for static and dynamic scenes, but most existing pipelines rely on dense multi-view capture or sufficient camera motion. This project focuses on a constrained setting with only four fixed cameras observing a dynamic human object, where limited view coverage makes stable geometry and novel-view rendering difficult. We first evaluate vanilla 3DGS and 4DGS pipelines on public datasets and self-recorded data to identify their limitations under sparse fixed views. The baseline experiments show that the original pipelines struggle to obtain reliable camera poses and stable geometry.... (More)
- This thesis explores 4D reconstruction in sparse fixed-camera settings for virtual reality visualization. Gaussian Splatting provides an efficient representation for static and dynamic scenes, but most existing pipelines rely on dense multi-view capture or sufficient camera motion. This project focuses on a constrained setting with only four fixed cameras observing a dynamic human object, where limited view coverage makes stable geometry and novel-view rendering difficult. We first evaluate vanilla 3DGS and 4DGS pipelines on public datasets and self-recorded data to identify their limitations under sparse fixed views. The baseline experiments show that the original pipelines struggle to obtain reliable camera poses and stable geometry. Geometry-based improvements are then explored, including depth-prior supervision and cross-view geometric consistency. These constraints provide additional supervision, but they do not recover unobserved regions or produce a stable full 360° reconstruction. Based on these observations, the final pipeline adopts a generative completion approach. The four cameras are calibrated, foreground masks are generated to isolate the target human object, and Diffuman4D is used to synthesize novel-view human videos. The generated multi-view data is adapted to vanilla 4DGS and trained with modified data loading to handle the larger dataset. The dynamic Gaussian representation is then exported as per-frame splat files and visualized in a Unity-based VR environment. The results show that generative completion is more suitable than geometry-only stabilization for this setting. The final reconstruction enables recognizable 360° dynamic observation, but remains limited by temporal inconsistencies in generated images and residual foreground noise. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9230317
- author
- Bati, Ezgi Aysel LU and Chen, Yiran LU
- supervisor
-
- Günter Alce LU
- organization
- course
- MAMM15 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- keywords
- 4DGS, Gaussian Splatting, Novel View Synthesis, Virtual Reality
- language
- English
- id
- 9230317
- date added to LUP
- 2026-06-04 10:21:33
- date last changed
- 2026-06-04 10:21:33
@misc{9230317,
abstract = {{This thesis explores 4D reconstruction in sparse fixed-camera settings for virtual reality visualization. Gaussian Splatting provides an efficient representation for static and dynamic scenes, but most existing pipelines rely on dense multi-view capture or sufficient camera motion. This project focuses on a constrained setting with only four fixed cameras observing a dynamic human object, where limited view coverage makes stable geometry and novel-view rendering difficult. We first evaluate vanilla 3DGS and 4DGS pipelines on public datasets and self-recorded data to identify their limitations under sparse fixed views. The baseline experiments show that the original pipelines struggle to obtain reliable camera poses and stable geometry. Geometry-based improvements are then explored, including depth-prior supervision and cross-view geometric consistency. These constraints provide additional supervision, but they do not recover unobserved regions or produce a stable full 360° reconstruction. Based on these observations, the final pipeline adopts a generative completion approach. The four cameras are calibrated, foreground masks are generated to isolate the target human object, and Diffuman4D is used to synthesize novel-view human videos. The generated multi-view data is adapted to vanilla 4DGS and trained with modified data loading to handle the larger dataset. The dynamic Gaussian representation is then exported as per-frame splat files and visualized in a Unity-based VR environment. The results show that generative completion is more suitable than geometry-only stabilization for this setting. The final reconstruction enables recognizable 360° dynamic observation, but remains limited by temporal inconsistencies in generated images and residual foreground noise.}},
author = {{Bati, Ezgi Aysel and Chen, Yiran}},
language = {{eng}},
note = {{Student Paper}},
title = {{4D RECONSTRUCTION IN SPARSE CAMERA SETTINGS FOR VR}},
year = {{2026}},
}