How Kurate.org assesses arXiv preprints on 16 dimensions across 47 categories, and how those assessments become rankings.
The system fetches the latest preprints from the arXiv API across Artificial Intelligence, Machine Learning, Quantum Physics, Applied Physics, Cosmology, Probability, Cryptography and Security, Distributed Computing, Game Theory, Information Theory, Robotics, Quantum Gases, Materials Science, Mathematical Physics, Computational Physics, General Physics, Optics, High Energy Astrophysics, Instrumentation and Methods, Inorganic Chemistry, Discrete Mathematics, Information Retrieval, Programming Languages, Social and Information Networks, Disordered Systems and Neural Networks, Mesoscale and Nanoscale Physics, Strongly Correlated Electrons, Superconductivity, Cryptographic Protocols, Public-key Cryptography, Economics, Optimization and Control, Chaotic Dynamics, General Relativity, High Energy Physics - Lattice, High Energy Physics - Phenomenology, High Energy Physics - Theory, Pattern Formation and Solitons, Chemical Physics, Plasma Physics, Space Physics, Biomolecules, Genomics, Neurons and Cognition, Populations and Evolution, Quantitative Methods, Machine Learning (Statistics) and downloads the full PDF for each. When a new category is added, the system fetches all papers published in the last 30 days; after that, it fetches incrementally from the most recent paper onward.
Each paper receives a single Claude Opus 4.8 assessment (adaptive extended thinking) generated from the full PDF text. One prompt produces a written impact assessment plus 16 scores on a 1–10 scale, each with its own short justification. View the assessment prompt →
Five core dimensions — the headline quality signals:
Eleven extended dimensions — the properties that decide whether a paper is worth your time:
Dimensions that genuinely do not apply to a paper are scored null rather than guessed — a survey has no reproducibility, a purely theoretical paper has no evidence strength. Definitions are deliberately anchored (what a 1, a 5 and a 10 mean) so scores stay comparable across papers, models and time.
All 16 metrics for every paper are laid out in one color-coded grid on the heatmap: one row per paper, one column per dimension, cell color from red (low) to green (high). Hovering a cell reveals the model's justification for that specific score, so a number is never presented without its reasoning.
Rows can be sorted by any dimension, filtered by field, date and per-metric thresholds — for example, high novelty and high reproducibility, or high significance among papers that are cheap to build on.
The statistics page shows the live distribution of each metric and the full 16×16 correlation matrix across the corpus. It makes the scoring behaviour auditable: which dimensions move together (significance and foundationality do), which are genuinely independent, and where the model is reluctant to use the full range (refutation and replication value are rare by construction).
Papers carry every arXiv tag they were submitted under, so they can be filtered by primary field or by tag combinations across fields (AND/OR) — clicking any tag on a heatmap row filters straight to it.
The system runs autonomously — fetching new papers on a schedule, assessing each one on all 16 dimensions, then idling until new work appears. The heatmap refreshes itself as new assessments land, so no manual step sits between a paper being scored and it appearing.
Limitations
AI-based evaluation is an approximation of scientific impact, not a replacement for human peer review. The 16 metrics reflect what a frontier model can infer from a paper's own text — they cannot verify that experiments were actually run as described. Scores are comparable within a field far more reliably than across unrelated fields, and papers assessed before the 16-metric prompt was introduced show only the five core dimensions until they are re-assessed.