🎓 EduPathHub

If AI solves the problem, what are we grading?

study-help ▲ 29 13 views 2026-07-18

I've been testing the latest frontier models (Gemini 3, Claude Opus 4.5, etc.) on some harder Linear Algebra problems from standard textbooks. They are now passing with flying colours.

This has forced me to ask: If the "answer" is free, what is the specific value-add of a human learning this material in 2025?

I see four possible "New Whys" for Math Education.

  1. The Auditor: We learn math so we can spot when the AI is hallucinating. (Focus on Verification)
  2. The Architect: We learn math so we can translate messy reality into the prompts/structures the AI can solve. (Focus on Modeling)
  3. The Athlete: The content is irrelevant; the rigor is the point. We are building "cognitive stamina" for other fields. (Focus on Neural Training)
  4. The Artist: We learn it because it is beautiful and human, regardless of utility. (Focus on Appreciation)

Personally, I want to believe in #4, but I fear universities require #1 or #2 to justify tuition.

Fellow math teachers, which one/ones do you think we should optimising our curriculum for?


Update: December 12, 2025 OpenAI has officially released GPT-5.2. According to their announcement, this model demonstrates significant performance improvements over GPT-5.1 across several key benchmarks.

Benchmark Category GPT-5.2 GPT-5.1
AIME 2025 (no tools) Competition Math 100.0% 94.0%
FrontierMath (Tier 1–3) Advanced Mathematics 40.3% 31.0%
FrontierMath (Tier 4) Advanced Mathematics 14.6% 12.5%

I included this data because it shows that suggesting LLMs cannot solve non-standard problems is a bit simplistic.


Update: March 18, 2026

There has been a recent effort in evaluating the mathematical ability of frontier models called First Proof, which challenges these models to prove 10 lemmas at the research math level. The proofs are known to the authors of this project but were not yet published initially. All in all, about 6-8 of these are solved by these models. See here for a comment by Daniel Litt.


Update: As Math Stack Exchange and Math Overflow both forbid AI generated content, I started my own forum to discuss the applications of AI in math: https://ai-for-math.discourse.group/

1 Answer

Shift your grading focus from the final result to the "decision trail." Instead of giving points for the correct answer, grade the logic used to get there. Ask students to explain why they chose a specific theorem or why a certain path was more efficient than another. This turns the assignment into a defense of a strategy rather than a calculation exercise. You should also introduce "critique assignments." Give students a solution generated by an AI that contains a subtle, logical flaw. Their grade depends on finding the error and explaining exactly where the reasoning broke down. This moves the goalpost from solving the problem to analyzing the process, which is a higher-level cognitive skill. Another practical step is to move toward oral exams or live demonstrations. If a student can talk through the conceptual connections in real-time, they've internalized the material. It's much harder to fake a deep understanding during a conversation than it is in a written paper. By grading the "how" and the "why" instead of the "what," you're assessing the student's mental model. This approach ensures that the human is still the one driving the logic, even if the AI is doing the heavy lifting.

Have a similar question?

Ask the community →
Share this question: Share Reddit