Deep Backdoors in Deep Reinforcement Learning Agents

No ratings

Presented at Black Hat USA 2024

Deep Reinforcement Learning (DRL) introduces a novel learning paradigm where AI agents learn by interacting with their environment. Due to its ability to make critical decisions at superhuman speeds, it is already transforming sectors such as autonomous driving, healthcare, cybersecurity, and nuclear fusion control. However, this introduces a new threat as these agents are typically assumed benign (even though their training is typically outsourced or even downloaded from the Internet e.g., HuggingFace.co)This talk will explore backdoors in DRL. We will first address how the resource-heavy training and the limited explainability of AI models expose users to supply chain attacks while hindering effective countermeasures. We will then dissect DRL backdoors, and show how sophisticated adversaries can embed backdoors in models while minimizing their footprint. Through practical demonstrations, we will illustrate backdoor functionality first in simpler settings and then in a simulation of a nuclear fusion reactor (KSTAR) Korea Superconducting Tokamak Advanced Research. To conclude, we will present a novel class of techniques for detecting DRL backdoors in real-time, enabling operators to intervene before any damage occurs.