An AI Model Programmed Nonstop for 19 Days on a Single MirrorCode Task That Cost $2,600
1 min read AI for Software Engineering (Copilots, SDLC, Testing) -/5
In short
  • Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code.
  • Claude Opus 4.7 leads with a 56 percent solve rate, successfully rebuilding a 16,000-line toolkit in just 14 hours.
  • However, every model tested still fails on the most complex tasks.
-/5 (0)
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, successfully rebuilding a 16,000-line toolkit in just 14 hours. However, every model tested still fails on the most complex tasks. This development is noteworthy but should be assessed in the context of previous market dynamics. A final assessment would be premature at this point.