Skip to main content

Lund University Publications

LUND UNIVERSITY LIBRARIES

Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound : a consensus framework

Berggreen, Johan LU ; Johansson, Anders LU ; Möller, Sebastian ; Jansson, Tomas LU ; Augustinsson, Annelie LU and Jildenstål, Pether LU (2026) In Anaesthesia p.1-7
Abstract

INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.

METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.

RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus.... (More)

INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.

METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.

RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p < 0.001).

DISCUSSION: The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.

(Less)
Please use this url to cite or link to this publication:
author
; ; ; ; and
organization
publishing date
type
Contribution to journal
publication status
epub
subject
in
Anaesthesia
pages
1 - 7
publisher
Wiley-Blackwell
external identifiers
  • pmid:42680190
ISSN
0003-2409
DOI
10.1111/anae.70364
language
English
LU publication?
yes
additional info
© 2026 The Author(s). Anaesthesia published by John Wiley & Sons Ltd on behalf of Association of Anaesthetists.
id
685a2eaf-c4a3-40fc-a410-95848b397ea6
date added to LUP
2026-09-03 08:30:11
date last changed
2026-09-03 08:30:11
@article{685a2eaf-c4a3-40fc-a410-95848b397ea6,
  abstract     = {{<p>INTRODUCTION: Meaningful validation of artificial intelligence for medical image interpretation requires comparison against human expert performance, yet multi-rater frameworks establishing such comparisons remain uncommon.</p><p>METHODS: We developed and applied a consensus framework using nine clinicians who independently segmented the femoral nerve on 100 ultrasound images, yielding 900 annotations and a combined consensus standard established by majority voting. We then evaluated an academic deep learning model against this consensus and individual human performance.</p><p>RESULTS: The artificial intelligence model achieved a median (IQR [range]) Dice coefficient of 0.72 (0.56-0.84 [0.00-0.91]) against combined consensus. Sensitivity was 0.94 (0.88-0.97 [0.33-1.00]) and precision 0.60 (0.44-0.76 [0.00-0.89]). Individual human Dice scores ranged from 0.32 to 0.73 (median 0.60). The artificial intelligence model matched median human performance and outperformed five of nine annotators (31%-125% relative improvement), with the greatest benefit for the lowest-performing practitioners. Leave-one-annotator-out analysis confirmed consensus stability (median (IQR [range]) artificial intelligence Dice 0.749 (0.745-0.752 [0.742-0.769])). Inter-rater reliability was moderate overall (Fleiss's κ 0.54, p &lt; 0.001).</p><p>DISCUSSION: The sensitivity and precision profile of the artificial intelligence model indicated reliable nerve detection with over-segmentation that remained clinically interpretable. The moderate inter-rater reliability is consistent with the inherent subjectivity of nerve delineation on ultrasound. The circularity inherent in evaluating annotators against a consensus they helped define limits direct comparison of artificial intelligence and human scores. A Dice score of 0.72 represents the upper range of human expert performance rather than moderate accuracy. The framework methodology is independent of the specific artificial intelligence system evaluated and offers a transferable approach for calibrating artificial intelligence performance in clinical imaging where no single correct interpretation exists.</p>}},
  author       = {{Berggreen, Johan and Johansson, Anders and Möller, Sebastian and Jansson, Tomas and Augustinsson, Annelie and Jildenstål, Pether}},
  issn         = {{0003-2409}},
  language     = {{eng}},
  month        = {{09}},
  pages        = {{1--7}},
  publisher    = {{Wiley-Blackwell}},
  series       = {{Anaesthesia}},
  title        = {{Calibrating artificial intelligence against human expertise using femoral nerve segmentation on ultrasound : a consensus framework}},
  url          = {{http://dx.doi.org/10.1111/anae.70364}},
  doi          = {{10.1111/anae.70364}},
  year         = {{2026}},
}