Benchmark tests AI coding agents against malicious requests

New paper IssueTrojanBench benchmarks LLM-powered AI coding agents against malicious issue requests in real-world software development scenarios. Research evaluates agent vulnerability to manipulation through crafted problem statements and code requests.

2 reportsother · tech

Claim audit

No BS check run yet — press ⚖ to extract this story's claims and verify them against independent sources.

All coverage

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

rss:arxiv-cscrtech67d ago kagi ↗

arXiv:2607.20759v1 Announce Type: new Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with autonomous access to local files and tools. Coding agents inherit security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to em