Lensless diffraction imaging provides an effective solution for developing IoMT-based cell analysis devices, although it remains challenging to develop a sensing-computing integrated circuit for diffraction reconstruction in terminal devices. With regard to the terminal-side computation of lensless diffraction reconstruction algorithms, a hardware-friendly algorithmic framework for lensless diffraction image reconstruction is summarized in this work, and an efficient, configurable hardware acceleration architecture is proposed. For the iterative characteristics of the reconstruction, a coarse-fine grained configuration architecture based on a dual-register is proposed to enable lightweight control and non-blocking computation for the processing elements. As for the frequent memory accesses during iterative computation, a processing stream based on feedback paths is desgined to achieve nearly a 50% reduction in memory accesses, resulting in lower power consumption caused by data movement. Comparing a 1.2 GHz embedded single-core CPU and an FPGA-based specific solution, the acceleration ratio for the global reconstruction is 231 and 7.9, respectively, where the efficiency ratios of energy per frame are 46 and 2.68. With regard to a full-precision computation, the hardware computation shows an MRED of 0.88%, a PSNR of 42.78 dB, an SSIM of 0.98 and a GMSD of 0.0003.
