Sharing my first research paper as an independent researcher — I evaluated Kavach, a policy-enforcement runtime for AI coding agents, against IssueTrojanBench’s real prompt-injection attack corpus (42 malicious tool-call actions from GitHub issue payloads).
Key finding: coverage comes from adapter canonicalization under a default-deny policy, not from benchmark-specific policy tuning — a distinction that matters for anyone evaluating similar guardrails.
Feedback welcome, and if anyone can point me to a cs.CR arXiv endorser, that’d help a lot — this is my first submission there.