Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Towards Autonomous Sports Production: A Multimodal Pipeline for Multi-Angle Basketball Highlight Generation

Bian, Yuhe LU and Zhou, Chun LU (2026) In Master's Theses in Mathematical Sciences FMAM02 20261
Mathematics (Faculty of Engineering)
Abstract
In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead.

This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.

The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against... (More)
In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead.

This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.

The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against emergent foundation models. Furthermore, we assess the capability of MLLMs to perform high-level decision-making in optimal view selection. By conducting a performance analysis on real-world multi-view datasets, this study validates the potential for multimodal architectures to achieve human-like directing logic with significantly reduced manual intervention and computational cost. (Less)
Popular Abstract
The Challenge: High Computing Costs in Video Analysis
In multi-camera sports broadcasting, extracting highlights and selecting the optimal camera angle relies heavily on manual directing. While Artificial Intelligence (AI) can potentially automate this process, processing multiple high-resolution video streams concurrently creates severe computational bottlenecks. Running complex video models on multiple live feeds simultaneously demands immense computer processing power, leading to high hardware costs and long processing delays on standard computers.
Our Work: A Step-by-Step Data Filtering System
To resolve this efficiency bottleneck, this thesis introduces an automated, end-to-end video processing pipeline. Instead of forcing heavy AI... (More)
The Challenge: High Computing Costs in Video Analysis
In multi-camera sports broadcasting, extracting highlights and selecting the optimal camera angle relies heavily on manual directing. While Artificial Intelligence (AI) can potentially automate this process, processing multiple high-resolution video streams concurrently creates severe computational bottlenecks. Running complex video models on multiple live feeds simultaneously demands immense computer processing power, leading to high hardware costs and long processing delays on standard computers.
Our Work: A Step-by-Step Data Filtering System
To resolve this efficiency bottleneck, this thesis introduces an automated, end-to-end video processing pipeline. Instead of forcing heavy AI models to endlessly analyze every second of all camera streams, our system utilizes a step-by-step filtering strategy to progressively remove unnecessary video data. First, the system watches the scoreboard, monitoring the stadium scoreboard and activating deep video analysis only when a point change is detected, which immediately narrows the focus to a 10-second window before the basket. Next, it finds the active half-court: a lightweight tracking module scans the main view to identify where players are clustered, instantly turning off the video streams from the empty or irrelevant side of the court. The system then pinpoints the action, precisely cropping the remaining video to find the exact fraction of a second where the scoring action occurs. Finally, it chooses the best angle, using a smart vision model to evaluate a single snapshot from each active camera and select the optimal view based on visual clarity, player visibility, and the absence of obstructions.
Results and Performance
We evaluated our system on real-world multi-angle basketball match footage. The pipeline successfully replaces manual logging while operating with minimal computing overhead. In aesthetic evaluations, the camera perspectives chosen by the AI achieved an 81.82% alignment accuracy with the choices of human sports-directing experts. Crucially, the entire system runs locally on a standard portable laptop without requiring expensive server clusters. It processes raw multi-angle footage and outputs a fully directed highlight clip in approximately 20 seconds. This demonstrates that smart data filters allow advanced video models to be deployed on everyday hardware for affordable and accessible automated sports production. (Less)
Please use this url to cite or link to this publication:
author
Bian, Yuhe LU and Zhou, Chun LU
supervisor
organization
course
FMAM02 20261
year
type
H2 - Master's Degree (Two Years)
subject
keywords
Temporal Sentence Grounding, Video Foundation Models, Vision-Language Models, Cross-Modal Retrieval, Multi-View Video Understanding
publication/series
Master's Theses in Mathematical Sciences
report number
LUTFMA-3617-2026
ISSN
1404-6342
other publication id
2026:E31
language
English
id
9236679
date added to LUP
2026-06-12 13:22:09
date last changed
2026-06-12 13:22:09
@misc{9236679,
  abstract     = {{In multi-view sports broadcasting, extracting high-quality highlights and selecting the optimal camera angle from concurrent video streams is traditionally a labor-intensive process with significant computational overhead. 

This thesis explores the feasibility of automating the production of multi-view sports highlights using Multimodal Large Language Models (MLLMs). To manage the high computational demands of multi-camera environments, we investigate a hierarchical data-reduction strategy that progressively filters temporal and spatial redundancies.

The research focuses on evaluating the effectiveness of vision-language alignment for precise action localization, comparing traditional video understanding architectures against emergent foundation models. Furthermore, we assess the capability of MLLMs to perform high-level decision-making in optimal view selection. By conducting a performance analysis on real-world multi-view datasets, this study validates the potential for multimodal architectures to achieve human-like directing logic with significantly reduced manual intervention and computational cost.}},
  author       = {{Bian, Yuhe and Zhou, Chun}},
  issn         = {{1404-6342}},
  language     = {{eng}},
  note         = {{Student Paper}},
  series       = {{Master's Theses in Mathematical Sciences}},
  title        = {{Towards Autonomous Sports Production: A Multimodal Pipeline for Multi-Angle Basketball Highlight Generation}},
  year         = {{2026}},
}