InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs
Original titleinternlm/AdvancedMathBench-AutoVerifier
AISummary
InternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step.
It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.
Source: InternLM (Shanghai AI Lab) · new models on Hugging Face · huggingface.coPublished · added here