Skip to main content

Lund University Publications

LUND UNIVERSITY LIBRARIES

Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes

Gullstrand Heander, Lo LU orcid ; Sergeyuk, Agnia ; Zakharov, Ilya ; Söderberg, Emma LU orcid and Mukhortov, Nikita (2026) In Leibniz International Proceedings in Informatics (LIPIcs) 394. p.1-89
Abstract
Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.

Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.

Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a... (More)
Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.

Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.

Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.

Results. Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50-3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.

Conclusions. Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention. (Less)
Please use this url to cite or link to this publication:
author
; ; ; and
organization
publishing date
type
Chapter in Book/Report/Conference proceeding
publication status
published
subject
host publication
20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)
series title
Leibniz International Proceedings in Informatics (LIPIcs)
editor
Feld, Robert ; Paasivaara, Maria ; Mendez, Daniel ; Wagner, Stefan and Muñoz Barón, Marvin
volume
394
article number
89
pages
1 - 89
publisher
Schloss Dagstuhl - Leibniz-Zentrum für Informatik
ISSN
1868-8969
ISBN
978-3-95977-450-5
DOI
10.4230/LIPIcs.ESEM.2026.89
project
DAPPER: Seamless, Tailored Code Review
language
English
LU publication?
yes
id
fc180a98-fbda-4ba0-8d24-53d6e1737b22
date added to LUP
2026-10-05 12:00:30
date last changed
2026-10-05 15:43:42
@inproceedings{fc180a98-fbda-4ba0-8d24-53d6e1737b22,
  abstract     = {{Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.<br/><br/>Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.<br/><br/>Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.<br/><br/>Results. Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50-3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.<br/><br/>Conclusions. Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention.}},
  author       = {{Gullstrand Heander, Lo and Sergeyuk, Agnia and Zakharov, Ilya and Söderberg, Emma and Mukhortov, Nikita}},
  booktitle    = {{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)}},
  editor       = {{Feld, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Muñoz Barón, Marvin}},
  isbn         = {{978-3-95977-450-5}},
  issn         = {{1868-8969}},
  language     = {{eng}},
  month        = {{10}},
  pages        = {{1--89}},
  publisher    = {{Schloss Dagstuhl - Leibniz-Zentrum für Informatik}},
  series       = {{Leibniz International Proceedings in Informatics (LIPIcs)}},
  title        = {{Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes}},
  url          = {{http://dx.doi.org/10.4230/LIPIcs.ESEM.2026.89}},
  doi          = {{10.4230/LIPIcs.ESEM.2026.89}},
  volume       = {{394}},
  year         = {{2026}},
}