ARR-ASSESS-1.0 · exact-version evidence

Model assessment ranking

Full policy

ARR asks the strongest suitable frontier models available for each assessment round to attack a paper, not merely summarize it. They search for counterexamples, hidden assumptions, proof gaps, unsupported novelty and reproducibility failures on the exact hashed version. ARR promises no fixed provider, model, report count or reasoning tier; every published score names the model, artifact, version and date, and unresolved material objections block admission.

Passing this unusually hard filter is meaningful positive evidence that a paper deserves serious attention. It is not infallibility: models can share blind spots, and a score cannot replace domain-expert review, formal proof or later correction.

3 rated current versions167 not yet rated4 preserved reports0 signed highlights
RankPaperMedian assessment
  1. 01
    ARR-2026-5QQF95VHTC9GABH8 · v1

    Sharp Onset and Unbounded Growth of Norm-Optimal Self-Commutator Rank

    Lluis Eriksson

    7.00Exceptional · n=1 · 7.00–7.00
  2. 02
    ARR-2026-1D2QV1RP1292JREW · v1

    Sharp Rank-Adaptive Bounds for Inverse Self-Commutators

    Lluis Eriksson

    4.20Strong · n=1 · 4.20–4.20
  3. 03
    ARR-2026-7NPRNBW4488HG90K · v1

    One-Spike Inverse Self-Commutators and Exact Three-versus-Four-Kick Curvature Synthesis

    Lluis Eriksson

    4.10Strong · n=1 · 4.10–4.10
High-ceiling research scale

Five is very good, not a failing grade.

The 0.00–10.00 Millennium scale is not a school percentage or a probability of correctness. Three stars is the acceptable publication floor. Ten is the top comparison anchor and cannot be established by a model alone.

StarsDisplayMeaningAnchor
1Critical concerns
2Substantial revision needed
3AcceptablePublication floor
4Strong
5Very goodVery good is deliberately above the publication floor
6Excellent
7Exceptional
8Potentially field-shaping
9Potentially historic
10Millennium-resolution benchmarkUnconditional recognized Millennium Problem solution after extraordinary verification

Criterion profile

Each report also supplies one to five stars, with a written basis, for correctness confidence, rigor, novelty, significance and reproducibility. These diagnostic ratings are shown separately and are not silently averaged into the headline score.

Only assessments marked independent of manuscript creation enter the median. ARR shows the count and range, never pools different paper versions and preserves later reassessments so future systems can be compared with earlier ones.

Download the versioned machine-readable assessment registry · JSON Schema