Statistics
Live distributions and correlations for all 16 rating dimensions across the full heatmap dataset. Statistics recompute automatically whenever new papers are rated.
64,498 papers · data updated 9/9/2026, 10:05:04 AM · auto-refreshes every minute
Impact
Composite overall predicted scientific impact.
Significance
Likely influence on future research, practice or applications.
Rigor
Soundness of the paper's design — proofs, baselines, controls — independent of the results obtained.
Novelty
Originality of the idea, method or framing.
Clarity
Writing quality and logical organization.
Difficulty
How much specialist background knowledge is needed to understand this paper.
Surprising
How unexpected the results are vs. current understanding.
Reproducible
Could an independent researcher replicate the main results from this paper alone?
Translational
How close this work is to real-world, commercial, or economic application.
Evidence
How well experiments, proofs and ablations support the claims.
Generalisable
How broadly the findings apply beyond the tested conditions.
Interdisciplinary
How many distinct research communities would benefit from this work.
Refutation
Whether this paper meaningfully challenges or overturns a prior finding.
Replication
Whether this paper independently corroborates a contested prior finding.
Resources
The scale of resources needed to produce or extend this work.
Foundational
Whether this is a reusable building block vs. a complete-in-itself result (empirically correlates strongly with significance).
Correlation Matrix (all 16 metrics)
Pearson correlation between all rating dimensions across 64,498 papers (pairwise-complete). Green = positive, red = negative. Metrics are ordered by their average correlation with all other metrics (descending). Strong correlations (>0.8) suggest redundancy; weak correlations confirm independent signal.
| Impact | Signif. | Found. | Novelty | Rigor | General. | Evid. | Surpris. | Diffic. | Clarity | Refut. | Interdis. | Resour. | Transl. | Replic. | Reprod. | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Impact | 0.96 | 0.87 | 0.84 | 0.73 | 0.75 | 0.63 | 0.65 | 0.61 | 0.60 | 0.36 | 0.34 | 0.33 | 0.36 | 0.17 | 0.17 | |
| Signif. | 0.96 | 0.88 | 0.84 | 0.63 | 0.75 | 0.58 | 0.63 | 0.59 | 0.55 | 0.36 | 0.33 | 0.36 | 0.38 | 0.17 | 0.12 | |
| Found. | 0.87 | 0.88 | 0.76 | 0.56 | 0.73 | 0.52 | 0.52 | 0.56 | 0.38 | 0.27 | 0.38 | 0.29 | 0.35 | 0.14 | 0.17 | |
| Novelty | 0.84 | 0.84 | 0.76 | 0.64 | 0.59 | 0.49 | 0.70 | 0.65 | 0.47 | 0.32 | 0.33 | 0.17 | 0.18 | 0.02 | 0.11 | |
| Rigor | 0.73 | 0.63 | 0.56 | 0.64 | 0.55 | 0.92 | 0.44 | 0.65 | 0.59 | 0.24 | 0.05 | 0.02 | -0.03 | 0.21 | 0.48 | |
| General. | 0.75 | 0.75 | 0.73 | 0.59 | 0.55 | 0.55 | 0.42 | 0.50 | 0.33 | 0.22 | 0.26 | 0.15 | 0.24 | 0.09 | 0.20 | |
| Evid. | 0.63 | 0.58 | 0.52 | 0.49 | 0.92 | 0.55 | 0.41 | 0.60 | 0.38 | 0.21 | 0.02 | -0.01 | -0.04 | 0.20 | 0.48 | |
| Surpris. | 0.65 | 0.63 | 0.52 | 0.70 | 0.44 | 0.42 | 0.41 | 0.44 | 0.20 | 0.57 | 0.21 | 0.14 | 0.12 | 0.14 | 0.12 | |
| Diffic. | 0.61 | 0.59 | 0.56 | 0.65 | 0.65 | 0.50 | 0.60 | 0.44 | 0.11 | 0.14 | 0.10 | 0.15 | -0.16 | 0.13 | 0.15 | |
| Clarity | 0.60 | 0.55 | 0.38 | 0.47 | 0.59 | 0.33 | 0.38 | 0.20 | 0.11 | 0.10 | 0.17 | 0.12 | 0.18 | 0.10 | 0.25 | |
| Refut. | 0.36 | 0.36 | 0.27 | 0.32 | 0.24 | 0.22 | 0.21 | 0.57 | 0.14 | 0.10 | 0.10 | 0.12 | 0.12 | 0.28 | 0.10 | |
| Interdis. | 0.34 | 0.33 | 0.38 | 0.33 | 0.05 | 0.26 | 0.02 | 0.21 | 0.10 | 0.17 | 0.10 | 0.14 | 0.32 | -0.01 | -0.04 | |
| Resour. | 0.33 | 0.36 | 0.29 | 0.17 | 0.02 | 0.15 | -0.01 | 0.14 | 0.15 | 0.12 | 0.12 | 0.14 | 0.44 | 0.18 | -0.37 | |
| Transl. | 0.36 | 0.38 | 0.35 | 0.18 | -0.03 | 0.24 | -0.04 | 0.12 | -0.16 | 0.18 | 0.12 | 0.32 | 0.44 | -0.07 | -0.23 | |
| Replic. | 0.17 | 0.17 | 0.14 | 0.02 | 0.21 | 0.09 | 0.20 | 0.14 | 0.13 | 0.10 | 0.28 | -0.01 | 0.18 | -0.07 | 0.13 | |
| Reprod. | 0.17 | 0.12 | 0.17 | 0.11 | 0.48 | 0.20 | 0.48 | 0.12 | 0.15 | 0.25 | 0.10 | -0.04 | -0.37 | -0.23 | 0.13 |
Methodology: Each dimension uses a 1.0–10.0 scale extracted by the extended summarization prompt, with null for non-applicable cases (e.g., reproducibility for purely theoretical papers). Scores shown are averaged across models where multiple assessments exist. Histograms show the live score distribution at 0.5-point granularity with the mean indicated by a dashed line. Correlations are Pearson, computed pairwise-complete over all papers that have both metrics.
