Prototype
Inside the folder
Project Lead, Agentic Tool Security
Tool inspectionPrompt injectionEvaluation
Description
An explainable prototype for inspecting suspicious tool behavior before it reaches an AI agent.
This is a prototype. The results below describe its evaluation.
My contribution
- Formalized prompt injection, tool poisoning, privilege escalation, and unauthorized data access as observable attack classes
- Combined tool schema, metadata, and runtime signals in a detection layer
- Used Bernoulli naive Bayes with deterministic policy gates to expose behavior signals and feature contributions
Outcomes & evidence
- Evaluated detection rate, false positives, and latency across 20 adversarial scenarios
- Demonstrated real-time blocking in the prototype
- Exposed false positives and missed detections rather than treating inspection as complete protection