Dataset
Referenced and pinned to a version. Never silently updated. Retraction propagates to everything derived from it.
A public research experiment
A datasetpaper is a versioned, forkable, executable research object built on an open dataset. The data, the code, the environment, the figures, and each individual claim are all addressable. The written narrative is just one view of it. Built to be verified, cited, and forked, by people and by AI agents. This is an experimental prototype, not a finished product or service.
Not a data descriptor. A data paper describes a dataset and omits the analysis on purpose. A datasetpaper is the analysis: hypotheses, methods, results, and claims, made machine-readable and forkable.
Read one
Same claims, same data, same provenance — rendered as an evidence-linked, multimodal article instead of a paper. Charts, images, interactives, and a click-through evidence viewer on every sentence.
71% of a systematic review's excluded records were rejected as simply off-topic — what a screening ledger reveals about keyword search on a young research term.
Read the story →A nationwide dam-removal database, re-read for what actually predicts river reconnection.
Read the story →An endosymbiont genome that should have shed genes wholesale — and the metabolic story behind what it kept.
Read the story →Generating an analysis and its write-up is no longer the hard part. That changes what is valuable. If anyone can produce a plausible analysis in minutes, plausibility is worthless, and the scarce thing is trust: knowing which claim was re-executed from the data and which was merely asserted. datasetpapers explores how verification and provenance can become the centre of the research object. The prose is a by-product.
Because the parts are addressable, an agent can build on one claim without parsing a whole document, and a person can cite a single figure with a stable identifier.
Referenced and pinned to a version. Never silently updated. Retraction propagates to everything derived from it.
A notebook plus a pinned container and lockfile, so the analysis re-runs deterministically, not approximately.
Each one points back to the exact code that produced it, so no result is orphaned from its computation.
Atomic statements, individually citable and machine-checkable, each carrying the figures that support it.
A graph tying every output back to the data, and recording whether a human or an AI produced each step.
Compiled from the object, not hand-edited. Available as a web page, a PDF, or structured XML.
Agent access
Agents can use ten open read tools at https://datasetpapers.com/mcp, or fetch the public
JSON, RO-Crate, claims, and verification files directly. The agent guide has connection configuration,
prompt examples, the complete tool list, and trust guidance.
Static and hosted MCP reads are live. Allowlisted writes create private jobs; the separate generation runner is not connected.
Fetch the corpus index, complete research objects, claim graphs, RO-Crates, and independent re-execution records over ordinary HTTPS.
Add https://datasetpapers.com/mcp as a Streamable HTTP server. Read tools need no key or account.
Check the independent re-execution record before relying on a claim. Engine self-reports and independent verification are not the same thing.
Allowlisted collaborators can queue proposals, forks, and verification evidence. A queued job is not an executed or published result.
In concrete terms, datasetpapers builds on open datasets from repositories such as Figshare and Zenodo, credits their depositors via ORCID, packages each result as an RO-Crate with a Croissant dataset description, mints persistent ARK identifiers, and passes new work through an adversarial review gate before publishing.
Every datasetpaper starts from someone's dataset. datasetpapers records that debt explicitly, notifies the original depositor that their data was used, and gives them a first-class place in the credit graph. The people who share data have been under-credited for as long as data sharing has existed. A world where machines analyse open data at scale makes that worse unless credit is designed in. Here it is.
Read a datasetpaper the way you would read a pull request, not a journal article. Check the claims you care about. Fork it if you can do better. Bring your dataset if you want it analysed. Treat it as a starting point that is honest about its own uncertainty, not a finished verdict.
No. It is a public research experiment and working prototype, not a commercial product or service. Its interfaces, methods, and outputs may change. Independently verify an analysis before relying on it for research or decisions.
A datasetpaper is a versioned, forkable, executable research object built on an open dataset. Its data, code, environment, figures, and individual claims are each addressable, and the written narrative is one rendered view of it. It is built to be verified, cited, and forked by people and by AI agents.
A data paper describes a dataset and deliberately omits the analysis. A datasetpaper is the analysis itself — hypotheses, methods, results, and claims — made machine-readable and forkable.
Researchers who want to build on open data, data depositors who want credit when their data is reused, and AI agents that need machine-readable, verifiable analyses to build on.
Agents can connect to the live hosted MCP at https://datasetpapers.com/mcp or read the corpus over ordinary HTTPS. The agent guide provides configuration, prompt examples, tool behavior, and the deliberately limited write path.
It builds on open datasets from repositories such as Figshare and Zenodo, credits depositors via ORCID, packages each result as an RO-Crate with a Croissant dataset description, mints persistent ARK identifiers, and passes new work through an adversarial review gate before publishing.