In plain words: Malware is tucked into the numbers that make up a neural network, so the file works as an image-recognition model and looks unchanged to scanners. They hid 36.9MB of malware inside a 178MB model with under 1% accuracy loss, and antivirus tools missed it.
Abstract · EvilModel: Hiding Malware Inside of Neural Network Models
Delivering malware covertly and evasively is critical to advanced malware campaigns. In this paper, we present a new method to covertly and evasively deliver malware through a neural network model. Neural network models are poorly explainable and have a good generalization ability. By embedding malware in neurons, the malware can be delivered covertly, with minor or no impact on the performance of neural network. Meanwhile, because the structure of the neural network model remains unchanged, it can pass the security scan of antivirus engines. Experiments show that 36.9MB of malware can be embedded in a 178MB-AlexNet model within 1% accuracy loss, and no suspicion is raised by anti-virus engines in VirusTotal, which verifies the feasibility of this method. With the widespread application of artificial intelligence, utilizing neural networks for attacks becomes a forwarding trend. We hope this work can provide a reference scenario for the defense on neural network-assisted attacks.
Zhi Wang, Chaoge Liu, Xiang Cui
arXiv:2107.08590 · cs.CR, cs.AI · submitted Jul 19, 2021 · updated Aug 5, 2021
abstract · pdf · html · To be appear at 26th IEEE Symposium on Computers and Communications (ISCC 2021)
Maybe the amount of data you can hide is higher, but that's primarily because they're storing all their weights as 32-bit floats which is overkill for inference.
...and I guess the fact you can retrain after hiding your malware to increase your inference accuracy again is maybe interesting?