AI scores a ‘C-’ on its hardest math test yet
The second batch of “First Proof” problems is meant to evaluate AI’s usefulness for research-level math.
The best model got six or seven of the ten questions right.
The second batch of “First Proof” problems is meant to evaluate AI’s usefulness for research-level math.
The best model got six or seven of the ten questions right.