Towards Autonomous Sports Production: A Multimodal Pipeline for Multi-Angle Basketball Highlight Generation
(2026) In Master's Theses in Mathematical Sciences FMAM02 20261Mathematics (Faculty of Engineering)
- Abstract
- In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead.
This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.
The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against... (More) - In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead.
This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.
The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against emergent foundation models. Furthermore, we assess the capability of MLLMs to perform high-level decision-making in optimal view selection. By conducting a performance analysis on real-world multi-view datasets, this study validates the potential for multimodal architectures to achieve human-like directing logic with significantly reduced manual intervention and computational cost. (Less) - Popular Abstract
- The Challenge: High Computing Costs in Video Analysis
In multi-camera sports broadcasting, extracting highlights and selecting the optimal camera angle relies heavily on manual directing. While Artificial Intelligence (AI) can potentially automate this process, processing multiple high-resolution video streams concurrently creates severe computational bottlenecks. Running complex video models on multiple live feeds simultaneously demands immense computer processing power, leading to high hardware costs and long processing delays on standard computers.
Our Work: A Step-by-Step Data Filtering System
To resolve this efficiency bottleneck, this thesis introduces an automated, end-to-end video processing pipeline. Instead of forcing heavy AI... (More) - The Challenge: High Computing Costs in Video Analysis
In multi-camera sports broadcasting, extracting highlights and selecting the optimal camera angle relies heavily on manual directing. While Artificial Intelligence (AI) can potentially automate this process, processing multiple high-resolution video streams concurrently creates severe computational bottlenecks. Running complex video models on multiple live feeds simultaneously demands immense computer processing power, leading to high hardware costs and long processing delays on standard computers.
Our Work: A Step-by-Step Data Filtering System
To resolve this efficiency bottleneck, this thesis introduces an automated, end-to-end video processing pipeline. Instead of forcing heavy AI models to endlessly analyze every second of all camera streams, our system utilizes a step-by-step filtering strategy to progressively remove unnecessary video data. First, the system watches the scoreboard, monitoring the stadium scoreboard and activating deep video analysis only when a point change is detected, which immediately narrows the focus to a 10-second window before the basket. Next, it finds the active half-court: a lightweight tracking module scans the main view to identify where players are clustered, instantly turning off the video streams from the empty or irrelevant side of the court. The system then pinpoints the action, precisely cropping the remaining video to find the exact fraction of a second where the scoring action occurs. Finally, it chooses the best angle, using a smart vision model to evaluate a single snapshot from each active camera and select the optimal view based on visual clarity, player visibility, and the absence of obstructions.
Results and Performance
We evaluated our system on real-world multi-angle basketball match footage. The pipeline successfully replaces manual logging while operating with minimal computing overhead. In aesthetic evaluations, the camera perspectives chosen by the AI achieved an 81.82% alignment accuracy with the choices of human sports-directing experts. Crucially, the entire system runs locally on a standard portable laptop without requiring expensive server clusters. It processes raw multi-angle footage and outputs a fully directed highlight clip in approximately 20 seconds. This demonstrates that smart data filters allow advanced video models to be deployed on everyday hardware for affordable and accessible automated sports production. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9236679
- author
- Bian, Yuhe LU and Zhou, Chun LU
- supervisor
-
- Viktor Larsson LU
- Gustav Hanning LU
- organization
- course
- FMAM02 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- keywords
- Temporal Sentence Grounding, Video Foundation Models, Vision-Language Models, Cross-Modal Retrieval, Multi-View Video Understanding
- publication/series
- Master's Theses in Mathematical Sciences
- report number
- LUTFMA-3617-2026
- ISSN
- 1404-6342
- other publication id
- 2026:E31
- language
- English
- id
- 9236679
- date added to LUP
- 2026-06-12 13:22:09
- date last changed
- 2026-06-12 13:22:09
@misc{9236679,
abstract = {{In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead.
This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.
The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against emergent foundation models. Furthermore, we assess the capability of MLLMs to perform high-level decision-making in optimal view selection. By conducting a performance analysis on real-world multi-view datasets, this study validates the potential for multimodal architectures to achieve human-like directing logic with significantly reduced manual intervention and computational cost.}},
author = {{Bian, Yuhe and Zhou, Chun}},
issn = {{1404-6342}},
language = {{eng}},
note = {{Student Paper}},
series = {{Master's Theses in Mathematical Sciences}},
title = {{Towards Autonomous Sports Production: A Multimodal Pipeline for Multi-Angle Basketball Highlight Generation}},
year = {{2026}},
}