Reinforcement Learning-Based Ordering
(2026) MIOM05 20261Department of Industrial and Mechanical Sciences
Production Management
- Abstract
- This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience... (More) - This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience buffer modification, and early stopping to improve
training efficiency and stability.
The performance of the RL-model is benchmarked against a standard peri-
odic review (R,Q)-policy using backorder and holding cost. Results indicate
that while RL demonstrates potential in capturing complex system dynamics
and achieving competitive performance in certain scenarios, its effectiveness is
highly dependent on data availability, demand characteristics, and computa-
tional constraints. In particular, challenges related to intermittent demand,
policy convergence, and generalization to unseen environments restrict consis-
tent outperformance of traditional methods.
The findings suggest that RL based approaches can complement existing inven-
tory control methods but cannot yet fully replace them in practical applications.
This highlights the importance of hybrid decision frameworks that leverage the
strengths of both data-driven and analytical methods. The thesis concludes with
recommendations for future development, including model scalability, improved
data utilization, and more robust evaluation across varying demand patterns. (Less) - Popular Abstract
- The handling of spare parts is a crucial and costly operation for manufacturing companies. Providing customers with quick and reliable access to spare parts while minimizing the tied up capital and required storage capacity, is a balancing worth the extra effort. With the introduction of Deep Reinforcement Learning (DRL), inventory management has great potential of improvement. Although still in quick development, DRL has the possibility of considering additional information in optimization and reduce inventory costs. Still, the best approach for this has not yet been found and revolutionary advantages lie in the future.
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9238531
- author
- Dahn, Samuel LU and Hallbeck, Theodor LU
- supervisor
- organization
- course
- MIOM05 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- keywords
- Machine Learning, Reinforcement Learning, Inventory Management, Inventory Control, Spare Parts, Ordering, Demand Forecasting
- other publication id
- 26/5355
- language
- English
- id
- 9238531
- date added to LUP
- 2026-06-17 16:10:04
- date last changed
- 2026-06-17 16:10:04
@misc{9238531,
abstract = {{This thesis investigates the applicability and performance of reinforcement learn-
ing (RL) for inventory management of spare parts in a single-echelon setting.
Traditional inventory control methods, such as (R,Q)-policies, typically decou-
ple forecasting and ordering decisions, which may lead to suboptimal system-
wide performance, particularly in environments characterized by intermittent
and stochastic demand.
To address these limitations, a reinforcement learning-based ordering model is
developed and evaluated using both simulated and real-world data provided
by Syncron AB. The proposed model is a Dueling Double Deep Q-Network
(DDDQN) architecture incorporating enhancements such as prioritized expe-
rience replay, experience buffer modification, and early stopping to improve
training efficiency and stability.
The performance of the RL-model is benchmarked against a standard peri-
odic review (R,Q)-policy using backorder and holding cost. Results indicate
that while RL demonstrates potential in capturing complex system dynamics
and achieving competitive performance in certain scenarios, its effectiveness is
highly dependent on data availability, demand characteristics, and computa-
tional constraints. In particular, challenges related to intermittent demand,
policy convergence, and generalization to unseen environments restrict consis-
tent outperformance of traditional methods.
The findings suggest that RL based approaches can complement existing inven-
tory control methods but cannot yet fully replace them in practical applications.
This highlights the importance of hybrid decision frameworks that leverage the
strengths of both data-driven and analytical methods. The thesis concludes with
recommendations for future development, including model scalability, improved
data utilization, and more robust evaluation across varying demand patterns.}},
author = {{Dahn, Samuel and Hallbeck, Theodor}},
language = {{eng}},
note = {{Student Paper}},
title = {{Reinforcement Learning-Based Ordering}},
year = {{2026}},
}