Skip to main content

Lund University Publications

LUND UNIVERSITY LIBRARIES

LUCIA : a scalable and versatile 0.82 TOPS/W/mm2 chiplet architecture for CNN in 22 nm FDSOI

Prieto Llorens, Arturo LU ; Westring, Kristoffer LU ; Allfjord, Alex LU ; Nouripayam, Masoud LU ; Castillo Mohedano, Sergio LU ; Marnfeldt, William LU ; Rodrigues, Joachim LU ; Svensson, Linus ; Andersson, Per LU and Åberg, Victor LU orcid (2026)
Abstract
This paper presents LUCIA, an architecture tailored for multi-chiplet systems optimized for convolutional neural network (CNN) acceleration through configurable hardware engines and distributed workload mapping. The architecture accommodates multiple independent near-memory computing (NMC) units in close proximity to their respective memory banks, tightly coupled with a low-power RISC-V processor. Dedicated memory control units facilitate efficient intra-chiplet and inter-chiplet data transfers, reducing execution time by 34% and consequently improving bandwidth utilization. Fabricated in 22nm FDSOI technology, LUCIA achieves a peak clock frequency of 565MHz and demonstrates an energy-area efficiency of 0.82TOPS/W/mm2, outperforming... (More)
This paper presents LUCIA, an architecture tailored for multi-chiplet systems optimized for convolutional neural network (CNN) acceleration through configurable hardware engines and distributed workload mapping. The architecture accommodates multiple independent near-memory computing (NMC) units in close proximity to their respective memory banks, tightly coupled with a low-power RISC-V processor. Dedicated memory control units facilitate efficient intra-chiplet and inter-chiplet data transfers, reducing execution time by 34% and consequently improving bandwidth utilization. Fabricated in 22nm FDSOI technology, LUCIA achieves a peak clock frequency of 565MHz and demonstrates an energy-area efficiency of 0.82TOPS/W/mm2, outperforming state-of-the-art implementations in comparable or more advanced technologies. (Less)
Please use this url to cite or link to this publication:
author
; ; ; ; ; ; ; ; and
organization
publishing date
type
Chapter in Book/Report/Conference proceeding
publication status
in press
subject
host publication
2026 IEEE European Solid-State Electronics Research Conference (ESSERC)
language
English
LU publication?
yes
id
1c6757c5-c2cd-42a4-a991-b0f0e92fbcfa
date added to LUP
2026-08-25 11:56:59
date last changed
2026-09-09 15:22:49
@inproceedings{1c6757c5-c2cd-42a4-a991-b0f0e92fbcfa,
  abstract     = {{This paper presents LUCIA, an architecture tailored for multi-chiplet systems optimized for convolutional neural network (CNN) acceleration through configurable hardware engines and distributed workload mapping. The architecture accommodates multiple independent near-memory computing (NMC) units in close proximity to their respective memory banks, tightly coupled with a low-power RISC-V processor. Dedicated memory control units facilitate efficient intra-chiplet and inter-chiplet data transfers, reducing execution time by 34% and consequently improving bandwidth utilization. Fabricated in 22nm FDSOI technology, LUCIA achieves a peak clock frequency of 565MHz and demonstrates an energy-area efficiency of 0.82TOPS/W/mm2, outperforming state-of-the-art implementations in comparable or more advanced technologies.}},
  author       = {{Prieto Llorens, Arturo and Westring, Kristoffer and Allfjord, Alex and Nouripayam, Masoud and Castillo Mohedano, Sergio and Marnfeldt, William and Rodrigues, Joachim and Svensson, Linus and Andersson, Per and Åberg, Victor}},
  booktitle    = {{2026 IEEE European Solid-State Electronics Research Conference (ESSERC)}},
  language     = {{eng}},
  title        = {{LUCIA : a scalable and versatile 0.82 TOPS/W/mm2 chiplet architecture for CNN in 22 nm FDSOI}},
  year         = {{2026}},
}