Glass-Box Security: Operationalizing Mechanistic Interpretability for Defending AI Agents

No ratings

Presented at [un]prompted 2026 by

Perimeter defenses are failing against the next generation of AI agents. This talk introduces "Glass-Box Security," a paradigm shift that utilizes Mechanistic Interpretability and Latent Space Geometry to monitor a model’s internal state for malicious intent and data exfiltration. We will explore why true observability requires a return to self-hosted infrastructure and present the Starseer architecture—a technical reference for building an "Internal EDR." Attendees will learn to replace fragile regex filters with "semantic tripwires" that detect deception and code leakage at the neuron level, long before the model generates output.