AI coding agents still struggle with large-scale refactoring, with the best model achieving only a 41.2% resolve rate on a The post Most coding agent benchmarks skip large-scale refactoring. Not this one. appeared first on The New Stack.
The New Stack