Back to Heatmap

Statistics

Live distributions and correlations for all 16 rating dimensions across the full heatmap dataset. Statistics recompute automatically whenever new papers are rated.

64,498 papers · data updated 9/9/2026, 10:05:04 AM · auto-refreshes every minute

Impact

Composite overall predicted scientific impact.

Mean: 5.54Std: 1.30Range: [1, 10]n=64,498
1.02.03.04.05.06.07.08.09.0030006000900012000μ=5.5

Significance

Likely influence on future research, practice or applications.

Mean: 5.63Std: 1.42Range: [0.5, 10]n=64,491
1.02.03.04.05.06.07.08.09.0030006000900012000μ=5.6

Rigor

Soundness of the paper's design — proofs, baselines, controls — independent of the results obtained.

Mean: 6.05Std: 1.56Range: [1, 10]n=64,483
1.02.03.04.05.06.07.08.09.0025005000750010000μ=6.0

Novelty

Originality of the idea, method or framing.

Mean: 5.63Std: 1.45Range: [0.5, 9.5]n=64,492
1.02.03.04.05.06.07.08.09.0025005000750010000μ=5.6

Clarity

Writing quality and logical organization.

Mean: 7.00Std: 1.00Range: [1.5, 9.5]n=64,492
1.02.03.04.05.06.07.08.09.00450090001350018000μ=7.0

Difficulty

How much specialist background knowledge is needed to understand this paper.

Mean: 6.51Std: 1.19Range: [1, 9.5]n=21,627
1.02.03.04.05.06.07.08.09.001500300045006000μ=6.5

Surprising

How unexpected the results are vs. current understanding.

Mean: 4.14Std: 1.23Range: [1, 9]n=21,293
1.02.03.04.05.06.07.08.09.00950190028503800μ=4.1

Reproducible

Could an independent researcher replicate the main results from this paper alone?

Mean: 6.53Std: 1.52Range: [1, 10]n=19,629
1.02.03.04.05.06.07.08.09.00750150022503000μ=6.5

Translational

How close this work is to real-world, commercial, or economic application.

Mean: 3.78Std: 1.73Range: [1, 9]n=21,621
1.02.03.04.05.06.07.08.09.00700140021002800μ=3.8

Evidence

How well experiments, proofs and ablations support the claims.

Mean: 5.76Std: 1.47Range: [1, 9.5]n=21,262
1.02.03.04.05.06.07.08.09.00900180027003600μ=5.8

Generalisable

How broadly the findings apply beyond the tested conditions.

Mean: 4.63Std: 1.10Range: [1, 9]n=21,551
1.02.03.04.05.06.07.08.09.001500300045006000μ=4.6

Interdisciplinary

How many distinct research communities would benefit from this work.

Mean: 3.36Std: 1.04Range: [1, 8]n=21,627
1.02.03.04.05.06.07.08.09.002000400060008000μ=3.4

Refutation

Whether this paper meaningfully challenges or overturns a prior finding.

Mean: 2.32Std: 1.21Range: [1, 10]n=21,617
1.02.03.04.05.06.07.08.09.001500300045006000μ=2.3

Replication

Whether this paper independently corroborates a contested prior finding.

Mean: 2.05Std: 0.88Range: [1, 7.5]n=21,445
1.02.03.04.05.06.07.08.09.002000400060008000μ=2.0

Resources

The scale of resources needed to produce or extend this work.

Mean: 3.27Std: 1.67Range: [1, 10]n=21,627
1.02.03.04.05.06.07.08.09.001500300045006000μ=3.3

Foundational

Whether this is a reusable building block vs. a complete-in-itself result (empirically correlates strongly with significance).

Mean: 4.31Std: 0.95Range: [1, 9.5]n=21,627
1.02.03.04.05.06.07.08.09.001500300045006000μ=4.3

Correlation Matrix (all 16 metrics)

Pearson correlation between all rating dimensions across 64,498 papers (pairwise-complete). Green = positive, red = negative. Metrics are ordered by their average correlation with all other metrics (descending). Strong correlations (>0.8) suggest redundancy; weak correlations confirm independent signal.

ImpactSignif.Found.NoveltyRigorGeneral.Evid.Surpris.Diffic.ClarityRefut.Interdis.Resour.Transl.Replic.Reprod.
Impact
0.96
0.87
0.84
0.73
0.75
0.63
0.65
0.61
0.60
0.36
0.34
0.33
0.36
0.17
0.17
Signif.
0.96
0.88
0.84
0.63
0.75
0.58
0.63
0.59
0.55
0.36
0.33
0.36
0.38
0.17
0.12
Found.
0.87
0.88
0.76
0.56
0.73
0.52
0.52
0.56
0.38
0.27
0.38
0.29
0.35
0.14
0.17
Novelty
0.84
0.84
0.76
0.64
0.59
0.49
0.70
0.65
0.47
0.32
0.33
0.17
0.18
0.02
0.11
Rigor
0.73
0.63
0.56
0.64
0.55
0.92
0.44
0.65
0.59
0.24
0.05
0.02
-0.03
0.21
0.48
General.
0.75
0.75
0.73
0.59
0.55
0.55
0.42
0.50
0.33
0.22
0.26
0.15
0.24
0.09
0.20
Evid.
0.63
0.58
0.52
0.49
0.92
0.55
0.41
0.60
0.38
0.21
0.02
-0.01
-0.04
0.20
0.48
Surpris.
0.65
0.63
0.52
0.70
0.44
0.42
0.41
0.44
0.20
0.57
0.21
0.14
0.12
0.14
0.12
Diffic.
0.61
0.59
0.56
0.65
0.65
0.50
0.60
0.44
0.11
0.14
0.10
0.15
-0.16
0.13
0.15
Clarity
0.60
0.55
0.38
0.47
0.59
0.33
0.38
0.20
0.11
0.10
0.17
0.12
0.18
0.10
0.25
Refut.
0.36
0.36
0.27
0.32
0.24
0.22
0.21
0.57
0.14
0.10
0.10
0.12
0.12
0.28
0.10
Interdis.
0.34
0.33
0.38
0.33
0.05
0.26
0.02
0.21
0.10
0.17
0.10
0.14
0.32
-0.01
-0.04
Resour.
0.33
0.36
0.29
0.17
0.02
0.15
-0.01
0.14
0.15
0.12
0.12
0.14
0.44
0.18
-0.37
Transl.
0.36
0.38
0.35
0.18
-0.03
0.24
-0.04
0.12
-0.16
0.18
0.12
0.32
0.44
-0.07
-0.23
Replic.
0.17
0.17
0.14
0.02
0.21
0.09
0.20
0.14
0.13
0.10
0.28
-0.01
0.18
-0.07
0.13
Reprod.
0.17
0.12
0.17
0.11
0.48
0.20
0.48
0.12
0.15
0.25
0.10
-0.04
-0.37
-0.23
0.13

Methodology: Each dimension uses a 1.0–10.0 scale extracted by the extended summarization prompt, with null for non-applicable cases (e.g., reproducibility for purely theoretical papers). Scores shown are averaged across models where multiple assessments exist. Histograms show the live score distribution at 0.5-point granularity with the mean indicated by a dashed line. Correlations are Pearson, computed pairwise-complete over all papers that have both metrics.