In plain words: By tuning into electromagnetic waves leaking from HDMI cables, a neural network turns those signals back into the displayed picture, trained on simulated signals so no real snooping rig is needed. It misread text characters over 60 percentage points less often than earlier setups.
Abstract · Deep-TEMPEST: Using Deep Learning to Eavesdrop on HDMI from its Unintended Electromagnetic Emanations
In this work, we address the problem of eavesdropping on digital video displays by analyzing the electromagnetic waves that unintentionally emanate from the cables and connectors, particularly HDMI. This problem is known as TEMPEST. Compared to the analog case (VGA), the digital case is harder due to a 10-bit encoding that results in a much larger bandwidth and non-linear mapping between the observed signal and the pixel's intensity. As a result, eavesdropping systems designed for the analog case obtain unclear and difficult-to-read images when applied to digital video. The proposed solution is to recast the problem as an inverse problem and train a deep learning module to map the observed electromagnetic signal back to the displayed image. However, this approach still requires a detailed mathematical analysis of the signal, firstly to determine the frequency at which to tune but also to produce training samples without actually needing a real TEMPEST setup. This saves time and avoids the need to obtain these samples, especially if several configurations are being considered. Our focus is on improving the average Character Error Rate in text, and our system improves this rate by over 60 percentage points compared to previous available implementations. The proposed system is based on widely available Software Defined Radio and is fully open-source, seamlessly integrated into the popular GNU Radio framework. We also share the dataset we generated for training, which comprises both simulated and over 1000 real captures. Finally, we discuss some countermeasures to minimize the potential risk of being eavesdropped by systems designed based on similar principles.
Santiago Fernández, Emilio Martínez, Gabriel Varela, Pablo Musé, Federico Larroca
arXiv:2407.09717 · cs.CR, cs.CV, cs.LG · submitted Jul 12, 2024
abstract · pdf · html · Submitted to LADC '24
Second, the paper was written (and tech implemented) by people with significant signals experience - quite a lot of thought went into the design, and a CNN (the part they trained) is just one component of the stack - for instance, they run the output image through Tesseract at the end to do character recognition. I’m not sure how they manage gradient descent end to end, although they talk about it in the paper.
So, this is a practitioner’s paper, using some modern techniques for ‘the hard bit’ -> taking radio waves and turning them into an image.
I’d be really interested in seeing someone do this again ‘the dumb way’ by just creating a full end to end autodifferentiable stack and running it for longer. I’m sure it would take more training time, but the number of people in the world who could have come up with this idea and done the implementation is small, probably in the single digit thousands.
Using, e.g. Sonnet 3.5 or Lllama 3.1 to be like ‘design and implement an autodifferentiable tempest attacker for me’, and seeing where results are right now is the sort of benchmark that I think matters a lot to track progress on the ‘leverage’ part of AI — basically can tech like this be delivered to, e.g. 1 million people worldwide, rather than thousands, with the help of a large model?
Anyway, very cool.
Finally, I’ll point out the two mitigations they mention don’t seem likely to be successful to me: they suggest adding Gaussian noise to the signal, or adding more gradients in colors for images. The second is not going to happen, except in very high security environments. I don’t believe the first is resistant to extra network training against the mitigation.