A research system, not a black box

Scientific discovery should be explainable, testable, and honest about uncertainty.

PaperMetrix is a research platform for retrieving and recommending scientific papers. It combines established information-retrieval methods with modern scholarly embeddings, while preserving the evidence needed to understand and evaluate every ranking.

RRFTransparent fusion
BM25Lexical
S2Semantic
GraphCitation
Papers
0
Authors
0
Scholarly sources
0
Open-access records
0
Our philosophy

Useful recommendations require visible evidence.

A ranking can be helpful without pretending to be an objective measure of scientific quality.

01

Traceable by design

Component ranks, matched terms, retrieval methods, latency, and experiment identifiers are recorded instead of hidden.

02

Time-aware evaluation

Temporal benchmarks prevent a system from recommending papers that did not exist at the simulated query date.

03

Humans remain the judge

Offline metrics and interaction logs support analysis, but independent relevance judgments remain essential for publishable claims.

How it works

From fragmented metadata to an inspectable ranking.

  1. 01

    Acquire

    OpenAlex, Semantic Scholar, Crossref, arXiv, and Unpaywall contribute complementary metadata and access signals.

  2. 02

    Normalize

    DOIs and source identifiers are reconciled into a canonical scholarly record with provenance and missing-data reasons.

  3. 03

    Retrieve

    BM25 captures exact terminology; SPECTER2 captures scholarly meaning; graph signals can contribute citation structure.

  4. 04

    Fuse and explain

    Reciprocal Rank Fusion combines positions without pretending that unrelated score scales are directly comparable.

  5. 05

    Evaluate

    Frozen splits, request-level telemetry, independent judgments, and reproducible artifacts connect results to defensible claims.

What is new here

LLMs are an auditable channel—not the foundation of truth.

PaperMetrix introduces generative models only where they can be isolated, disabled, priced, and compared against a fixed baseline.

Independent LLM expansion

Query expansion is versioned by model and prompt, protected by budgets and circuit breakers, and currently restricted to staff experiments.

Hybrid retrieval with provenance

Lexical, semantic, and graph channels remain independently observable before fusion.

Publication-ready evaluation

Temporal integrity, frozen development/test partitions, pooling, agreement, and cost/latency reporting are first-class parts of the platform.

Online and offline connection

Impressions, clicks, saves, and relevance feedback can be linked to the exact retrieval request without turning behavior into automatic truth.

A clear boundary

What PaperMetrix does—and does not do.

It helps you

  • find related and topically relevant literature
  • inspect why a result was retrieved
  • organize papers into reusable research collections
  • run reproducible retrieval experiments

It does not claim to

  • measure scientific quality from rank alone
  • replace systematic-review screening
  • treat citations or clicks as unbiased relevance
  • let an LLM silently rewrite the benchmark
Start with a question

Follow the evidence trail from query to paper.

Search the current corpus and inspect the retrieval signals behind every live result.

Open scholarly search →