SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks https:// arxiv.org/html/2603.24755v1#A7

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks https:// arxiv.org/html/2603.24755v1#A7

1 reportother

Claim audit

No BS check run yet — press ⚖ to extract this story's claims and verify them against independent sources.

All coverage