In plain words: A flat camera swaps the lens for a thin mask that scrambles incoming light, then a trained image generator rebuilds the picture from the sensor's raw data, optionally guided by a text description. Its reconstructions beat earlier recovery algorithms on quality and naturalness.
Abstract
The flat lensless camera design reduces the camera size and weight significantly. In this design, the camera lens is replaced by another optical element that interferes with the incoming light. The image is recovered from the raw sensor measurements using a reconstruction algorithm. Yet, the quality of the reconstructed images is not satisfactory. To mitigate this, we propose utilizing a pre-trained diffusion model with a control network and a learned separable transformation for reconstruction. This allows us to build a prototype flat camera with high-quality imaging, presenting state-of-the-art results in both terms of quality and perceptuality. We demonstrate its ability to leverage also textual descriptions of the captured scene to further enhance reconstruction. Our reconstruction method which leverages the strong capabilities of a pre-trained diffusion model can be used in other imaging systems for improved reconstruction results.
Erez Yosef, Raja Giryes
arXiv:2408.07541 · cs.CV, cs.AI, eess.IV · submitted Aug 14, 2024
abstract · pdf · html
https://waller-lab.github.io/DiffuserCam/ https://waller-lab.github.io/DiffuserCam/tutorial.html includes instructions and code to build your own https://ieeexplore.ieee.org/abstract/document/8747341 https://ieeexplore.ieee.org/document/7492880