Most coding agent benchmarks skip large-scale refactoring. Not this one.

The New Stack The New Stack

https://cdn.thenewstack.io/media/2026/08/2bd0685f-paris-bilal-9w7tpcva0xk-unsplash-1024x768.jpg" style="display: block; margin: auto; margin-bottom: 20px;" width="1024" />

AI coding agents still struggle with large-scale refactoring, with the best model achieving only a 41.2% resolve rate on a


The post https://thenewstack.io/ai-agents-refactoring-benchmarks/">Most coding agent benchmarks skip large-scale refactoring.

Not this one. appeared first on https://thenewstack.io">The New Stack.

Read full article at The New Stack →