Model assessment ranking
ARR asks the strongest suitable frontier models available for each assessment round to attack a paper, not merely summarize it. They search for counterexamples, hidden assumptions, proof gaps, unsupported novelty and reproducibility failures on the exact hashed version. ARR promises no fixed provider, model, report count or reasoning tier; every published score names the model, artifact, version and date, and unresolved material objections block admission.
Passing this unusually hard filter is meaningful positive evidence that a paper deserves serious attention. It is not infallibility: models can share blind spots, and a score cannot replace domain-expert review, formal proof or later correction.
- 01ARR-2026-5QQF95VHTC9GABH8 · v1
Sharp Onset and Unbounded Growth of Norm-Optimal Self-Commutator Rank
7.00Exceptional · n=1 · 7.00–7.00 - 024.20Strong · n=1 · 4.20–4.20
- 03ARR-2026-7NPRNBW4488HG90K · v1
One-Spike Inverse Self-Commutators and Exact Three-versus-Four-Kick Curvature Synthesis
4.10Strong · n=1 · 4.10–4.10
Five is very good, not a failing grade.
The 0.00–10.00 Millennium scale is not a school percentage or a probability of correctness. Three stars is the acceptable publication floor. Ten is the top comparison anchor and cannot be established by a model alone.
| Stars | Display | Meaning | Anchor |
|---|---|---|---|
| 1 | Critical concerns | ||
| 2 | Substantial revision needed | ||
| 3 | Acceptable | Publication floor | |
| 4 | Strong | ||
| 5 | Very good | Very good is deliberately above the publication floor | |
| 6 | Excellent | ||
| 7 | Exceptional | ||
| 8 | Potentially field-shaping | ||
| 9 | Potentially historic | ||
| 10 | Millennium-resolution benchmark | Unconditional recognized Millennium Problem solution after extraordinary verification |
Criterion profile
Each report also supplies one to five stars, with a written basis, for correctness confidence, rigor, novelty, significance and reproducibility. These diagnostic ratings are shown separately and are not silently averaged into the headline score.
Only assessments marked independent of manuscript creation enter the median. ARR shows the count and range, never pools different paper versions and preserves later reassessments so future systems can be compared with earlier ones.
Download the versioned machine-readable assessment registry · JSON Schema