Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes
(2026) In Leibniz International Proceedings in Informatics (LIPIcs) 394. p.1-89- Abstract
- Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.
Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.
Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a... (More) - Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.
Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.
Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.
Results. Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50-3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.
Conclusions. Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/record/fc180a98-fbda-4ba0-8d24-53d6e1737b22
- author
- Gullstrand Heander, Lo
LU
; Sergeyuk, Agnia
; Zakharov, Ilya
; Söderberg, Emma
LU
and Mukhortov, Nikita
- organization
- publishing date
- 2026-10-05
- type
- Chapter in Book/Report/Conference proceeding
- publication status
- published
- subject
- host publication
- 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)
- series title
- Leibniz International Proceedings in Informatics (LIPIcs)
- editor
- Feld, Robert ; Paasivaara, Maria ; Mendez, Daniel ; Wagner, Stefan and Muñoz Barón, Marvin
- volume
- 394
- article number
- 89
- pages
- 1 - 89
- publisher
- Schloss Dagstuhl - Leibniz-Zentrum für Informatik
- ISSN
- 1868-8969
- ISBN
- 978-3-95977-450-5
- DOI
- 10.4230/LIPIcs.ESEM.2026.89
- project
- DAPPER: Seamless, Tailored Code Review
- language
- English
- LU publication?
- yes
- id
- fc180a98-fbda-4ba0-8d24-53d6e1737b22
- date added to LUP
- 2026-10-05 12:00:30
- date last changed
- 2026-10-05 15:43:42
@inproceedings{fc180a98-fbda-4ba0-8d24-53d6e1737b22,
abstract = {{Background. Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.<br/><br/>Aims. We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.<br/><br/>Method. In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.<br/><br/>Results. Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50-3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.<br/><br/>Conclusions. Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention.}},
author = {{Gullstrand Heander, Lo and Sergeyuk, Agnia and Zakharov, Ilya and Söderberg, Emma and Mukhortov, Nikita}},
booktitle = {{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)}},
editor = {{Feld, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Muñoz Barón, Marvin}},
isbn = {{978-3-95977-450-5}},
issn = {{1868-8969}},
language = {{eng}},
month = {{10}},
pages = {{1--89}},
publisher = {{Schloss Dagstuhl - Leibniz-Zentrum für Informatik}},
series = {{Leibniz International Proceedings in Informatics (LIPIcs)}},
title = {{Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes}},
url = {{http://dx.doi.org/10.4230/LIPIcs.ESEM.2026.89}},
doi = {{10.4230/LIPIcs.ESEM.2026.89}},
volume = {{394}},
year = {{2026}},
}