Skip to main content

LUP Student Papers

LUND UNIVERSITY LIBRARIES

Reinforcement Learning-Based Ordering

Dahn, Samuel LU and Hallbeck, Theodor LU (2026) MIOM05 20261
Department of Industrial and Mechanical Sciences
Production Management
Abstract
This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience... (More)
This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience buffer modification, and early stopping to improve
training efficiency and stability.
The performance of the RL-model is benchmarked against a standard peri-
odic review (R,Q)-policy using backorder and holding cost. Results indicate
that while RL demonstrates potential in capturing complex system dynamics
and achieving competitive performance in certain scenarios, its effectiveness is
highly dependent on data availability, demand characteristics, and computa-
tional constraints. In particular, challenges related to intermittent demand,
policy convergence, and generalization to unseen environments restrict consis-
tent outperformance of traditional methods.
The findings suggest that RL based approaches can complement existing inven-
tory control methods but cannot yet fully replace them in practical applications.
This highlights the importance of hybrid decision frameworks that leverage the
strengths of both data-driven and analytical methods. The thesis concludes with
recommendations for future development, including model scalability, improved
data utilization, and more robust evaluation across varying demand patterns. (Less)
Popular Abstract
The handling of spare parts is a crucial and costly operation for manufacturing companies. Providing customers with quick and reliable access to spare parts while minimizing the tied up capital and required storage capacity, is a balancing worth the extra effort. With the introduction of Deep Reinforcement Learning (DRL), inventory management has great potential of improvement. Although still in quick development, DRL has the possibility of considering additional information in optimization and reduce inventory costs. Still, the best approach for this has not yet been found and revolutionary advantages lie in the future.
Please use this url to cite or link to this publication:
author
Dahn, Samuel LU and Hallbeck, Theodor LU
supervisor
organization
course
MIOM05 20261
year
type
H2 - Master's Degree (Two Years)
subject
keywords
Machine Learning, Reinforcement Learning, Inventory Management, Inventory Control, Spare Parts, Ordering, Demand Forecasting
other publication id
26/5355
language
English
id
9238531
date added to LUP
2026-06-17 16:10:04
date last changed
2026-06-17 16:10:04
@misc{9238531,
  abstract     = {{This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience buffer modification, and early stopping to improve
training efficiency and stability.
The performance of the RL-model is benchmarked against a standard peri-
odic review (R,Q)-policy using backorder and holding cost. Results indicate
that while RL demonstrates potential in capturing complex system dynamics
and achieving competitive performance in certain scenarios, its effectiveness is
highly dependent on data availability, demand characteristics, and computa-
tional constraints. In particular, challenges related to intermittent demand,
policy convergence, and generalization to unseen environments restrict consis-
tent outperformance of traditional methods.
The findings suggest that RL based approaches can complement existing inven-
tory control methods but cannot yet fully replace them in practical applications.
This highlights the importance of hybrid decision frameworks that leverage the
strengths of both data-driven and analytical methods. The thesis concludes with
recommendations for future development, including model scalability, improved
data utilization, and more robust evaluation across varying demand patterns.}},
  author       = {{Dahn, Samuel and Hallbeck, Theodor}},
  language     = {{eng}},
  note         = {{Student Paper}},
  title        = {{Reinforcement Learning-Based Ordering}},
  year         = {{2026}},
}