No title
This paper presents LUCIA, an architecture tailored for multi-chiplet systems optimized for convolutional neural network (CNN) acceleration through configurable hardware engines and distributed workload mapping. The architecture accommodates multiple independent near-memory computing (NMC) units in close proximity to their respective memory banks, tightly coupled with a low-power RISC-V processor.
