About Paper Lineage
Every idea has ancestors. Paper Lineage traces 197 AI papers back from BLIP-2 (2023), through the papers they built on, to the ideas those papers introduced.
You can ask where an idea came from, what a paper built on, or how two papers are connected. Every answer points to the papers and the sentences behind it.
How links are verified
Each link between two papers holds two different kinds of claim, and they look different on every page:
- Cites · not yet reviewed
A citation fact: “BLIP-2 cites Flamingo 5 times”, with the sentence and section. Mined automatically from the later paper’s full text, and always shown.
- Extends
A lineage claim: “BLIP-2 extends BLIP”. A judgement, so it appears only after a human curator has read the evidence and accepted it (0 of 1,349 so far).
Summaries and concept tags written by AI are labelled “Drafted by AI · not yet reviewed”.
Where the data comes from
Starting from BLIP-2, a script read each paper’s full text (arXiv’s HTML renders, or the PDF), parsed its bibliography, and kept every in-text citation with its sentence and section. A reference counted as influence when it was cited twice or more, or inside a Method-like section. That walk went three generations back: 197 papers and 1,349 links.180 concepts in 25 themes were then written from the papers’ abstracts and credited to the paper that introduced each one.
Limits: only papers on arXiv are included, the dataset is traced backwards from one paper, and citation sentences are parsed automatically, so a few may be cut or misplaced.
Built with Sanity
The papers, links and concepts live in Sanity’s Content Lake. Curators work in a customised Sanity Studio and the Lineage Desk (an App SDK app). The Ask page uses Sanity Context, with one endpoint over the live dataset and one over a Knowledge Base of the papers’ full text. Live updates come from Sanity’s Live Content API.
Privacy
When you ask a question we store its text, the time, whether it was answered, and which papers were cited, under private IDs that public queries can’t read. We don’t store IP addresses (the rate limiter keeps only a salted hash, for at most a day), set no cookies, and run no analytics. Your conversations stay in your own browser.
Credits
Papers and metadata from arXiv and ar5iv; title lookups via OpenAlex. Suggest a paper that belongs here.