@misc{9248572,
  abstract     = {{Purpose: The purpose of this project was to examine how the degree of cancer suspicion
evaluated by an AI system diﬀers when examining repeated images of a realis=c breast
phantom acquired with DM and DBT. This project is a part of a larger project to develop quality
assurance systems and methods for AI systems in breast cancer screening.

Background: AI systems are currently being introduced in breast cancer screening with
promising results, especially with respect to reduc=on of radiologist workload, but also
increased sensi=vity at constant or reduced recall rates. Similarly to other technology, solid
quality assurance programs for AI systems are needed to guarantee a good and constant
performance. Current mammography imaging QA workflows with technical phantoms are
not adapted to test AI systems. Previously, baseline varia=ons of the output of a cancer
detec=on AI system with repeated imaging of an anthropomorphic breast phantom under
constant imaging condi=ons was inves=gated. To establish reliable QC of AI, these baseline
varia=ons warrant a more thorough examina=on of AI output at repeated imaging both for
digital mammography (DM) and digital breast tomosynthesis (DBT).

Methods: An anthropomorphic, physical 3D breast phantom with a simulated lesion was
repeatedly imaged using automa=c exposure control (AEC) with both DM and DBT. AEC
imaging was done using three diﬀerent mammography systems: Siemens Mammomat
B.brilliant, GE Senographe Pris=na and Hologic Selenia Dimensions. The diﬀerence in central
tendency and variance between DM and DBT were examined using Mann-Whitney U tests
and Levene’s test. Furthermore, the tube voltage and tube loading of the Siemens
Mammomat B.brilliant were varied, based on exposure seWngs determined by the AEC
system. For each set of exposure parameters, repeated imaging was performed. The images
were analysed using an AI-based cancer detec=on system. Correla=ons between exposure
seWngs and the AI cancer-risk assessment were studied and results for DM and DBT were
compared.

Results: The results when varying exposure settings generally point towards a
significant increase of the central tendency of region scores as both tube loading and
tube voltage increase, for all three systems.

Furthermore, the comparison of DM and DBT using AEC showed a significant diWerence
in central tendency of AI risk scores (Mann-Whitney U test, α=0.05), the significance
held for all systems. The variance of the AI risk score when comparing DM and DBT using
AEC was significantly diWerent for the Siemens and Hologic systems, while no
significant diWerence in variance was found for the risk score of the GE system (Levene’s
test, α=0.05).

Conclusion: The results point towards higher repeatability of AI risk scores when using DBT
imaging compared to DM. The results also point towards a higher sensi=vity for DBT imaging
than DM}},
  author       = {{Magnusson, Jonas}},
  language     = {{eng}},
  note         = {{Student Paper}},
  title        = {{AI Lesion Risk Score with DBT and DM}},
  year         = {{2026}},
}

