Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound : a consensus framework
(2026) In Anaesthesia p.1-7- Abstract
INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.
METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.
RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus.... (More)
INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.
METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.
RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p < 0.001).
DISCUSSION: The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.
(Less)
- author
- Berggreen, Johan LU ; Johansson, Anders LU ; Möller, Sebastian ; Jansson, Tomas LU ; Augustinsson, Annelie LU and Jildenstål, Pether LU
- organization
- publishing date
- 2026-09-01
- type
- Contribution to journal
- publication status
- epub
- subject
- in
- Anaesthesia
- pages
- 1 - 7
- publisher
- Wiley-Blackwell
- external identifiers
-
- pmid:42680190
- ISSN
- 0003-2409
- DOI
- 10.1111/anae.70364
- language
- English
- LU publication?
- yes
- additional info
- © 2026 The Author(s). Anaesthesia published by John Wiley & Sons Ltd on behalf of Association of Anaesthetists.
- id
- 685a2eaf-c4a3-40fc-a410-95848b397ea6
- date added to LUP
- 2026-09-03 08:30:11
- date last changed
- 2026-09-03 08:30:11
@article{685a2eaf-c4a3-40fc-a410-95848b397ea6,
abstract = {{<p>INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.</p><p>METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.</p><p>RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p < 0.001).</p><p>DISCUSSION: The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.</p>}},
author = {{Berggreen, Johan and Johansson, Anders and Möller, Sebastian and Jansson, Tomas and Augustinsson, Annelie and Jildenstål, Pether}},
issn = {{0003-2409}},
language = {{eng}},
month = {{09}},
pages = {{1--7}},
publisher = {{Wiley-Blackwell}},
series = {{Anaesthesia}},
title = {{Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound : a consensus framework}},
url = {{http://dx.doi.org/10.1111/anae.70364}},
doi = {{10.1111/anae.70364}},
year = {{2026}},
}