Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes
(2026)- Abstract
- Background: Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.
Aims: We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.
Method: In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a... (More) - Background: Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.
Aims: We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.
Method: In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.
Results: Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50--3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.
Conclusions: Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/record/198a12d0-d377-4077-aa00-6fea1fef0541
- author
- Gullstrand Heander, Lo
LU
; Sergeyuk, Agnia
; Zakharov, Ilya
; Söderberg, Emma
LU
and Mukhortov, Nikita
- organization
- publishing date
- 2026-06-02
- type
- Working paper/Preprint
- publication status
- published
- subject
- keywords
- Code review, Participatory design, LLM-generated code, Trust calibration, Software development tools
- pages
- 19 pages
- publisher
- arXiv.org
- DOI
- 10.48550/arXiv.2606.01969
- project
- DAPPER: Seamless, Tailored Code Review
- How can code reviews be made fit-for-purpose?
- language
- English
- LU publication?
- yes
- id
- 198a12d0-d377-4077-aa00-6fea1fef0541
- date added to LUP
- 2026-06-02 10:39:37
- date last changed
- 2026-08-12 12:00:09
@misc{198a12d0-d377-4077-aa00-6fea1fef0541,
abstract = {{Background: Developers increasingly review multi-file code changes generated by LLM-based agents, yet no validated end-to-end workflow or IDE tooling design exists for this scenario.<br/><br/>Aims: We investigate (RQ1) the challenges developers face when reviewing LLM-generated multi-file changes and (RQ2) how developers envision effective workflows for this task.<br/><br/>Method: In collaboration with JetBrains, we conducted a participatory design study structured using the double-diamond design process with Discover, Define, Develop, and Deliver phases. Industry practitioners participated in the Discover phase (N=17); seven of these returned for the Develop phase. The Define phase was an author-led synthesis. The Deliver phase produced a conceptual design and a high-fidelity semi-interactive prototype evaluated through a follow-up survey with N=43 practitioners.<br/><br/>Results: Participants identified trust-calibration as the central challenge. The study yielded a three-level review workflow (overview, file-analysis, code snippet review) supported by seven design constructs (chunk, risk-per-line, risk-per-file, judge, walk-through, zooming in/out, and security cage). In the validation survey, all three workflow levels scored above the neutral midpoint (means 3.50--3.91 on a five-point scale). Of the respondents, 63% expected reduced overall review effort, and 52% reduced trust-assessment effort, relative to their current tools. These findings suggest that the design constructs indicate a positive direction for future tool development.<br/><br/>Conclusions: Reviewing LLM-generated multi-file changes is a trust-calibration problem rather than a diffing problem. The three-level workflow and the seven constructs we report give tool designers a conceptual framework for building AI-ready code review tools that surface risk and confidence signals at the granularity at which developers allocate attention.}},
author = {{Gullstrand Heander, Lo and Sergeyuk, Agnia and Zakharov, Ilya and Söderberg, Emma and Mukhortov, Nikita}},
keywords = {{Code review; Participatory design; LLM-generated code; Trust calibration; Software development tools}},
language = {{eng}},
month = {{06}},
note = {{Preprint}},
publisher = {{arXiv.org}},
title = {{Trust-calibrated code review : a participatory design study of review workflows for LLM-generated multi-file changes}},
url = {{http://dx.doi.org/10.48550/arXiv.2606.01969}},
doi = {{10.48550/arXiv.2606.01969}},
year = {{2026}},
}