use cases

A PhD Student's Literature Review in One File: Citation Links, Abstract Search and SQL Together

Take a second-year PhD student in public health with 1,400 papers in Zotero, all about loneliness in older adults. Her supervisor points her to a 2024 review they both trust and asks for a reading list built around it. The list should hold the papers near that review in the citation graph from 2018 on, with the ones closest to her research question first.

Each half of that request is easy on its own. Citation tools follow references, and search engines rank by meaning. What she doesn’t have is one place that does both over her own library, and lets her rephrase the question next week and run it again.

A seed review with two hops of citations. Papers from before 2018 are marked red and left out, and the rest are ranked into a reading list by closeness to the research question.

A Table She Designs Herself

She makes the file and the table with plain SQL, choosing the columns herself, then hands the table to HyperCrux:

hypercrux init lit.db
hypercrux sql lit.db "CREATE TABLE paper (key TEXT PRIMARY KEY, title TEXT, year INTEGER, journal TEXT, vec BLOB)"
hypercrux adopt lit.db paper

adopt installs the triggers that make paper a record table. From then on every row needs a key such as paper:10.1000/rev.2024.118, and a paper’s links go when it does.

Loading It From Python

Her script reads the Zotero export and adds each paper’s reference list from a citation database such as OpenAlex or Crossref, along with an embedding of its abstract. To write the file it uses hcfile.py from the repository’s examples folder, which needs nothing beyond Python’s own sqlite3 module:

import sqlite3
from hcfile import put, link

for p in papers:            # from a Zotero or BibTeX export
    put("lit.db", "paper:" + p["doi"], {"title": p["title"], "year": p["year"],
                                         "journal": p["journal"], "vec": embed(p["abstract"])})
for p in papers:
    for ref in p["references"]:
        try:
            link("lit.db", "paper:" + p["doi"], "cites", "paper:" + ref)
        except sqlite3.IntegrityError:
            pass            # a reference that isn't in the library

embed stands for her embedding model. Most papers cite plenty of work she hasn’t saved, and the file refuses a link to a record that doesn’t exist, so those references are skipped. That refusal comes from a trigger in the file, which means it holds for any program that writes there. The script is safe to run again next month, too. put updates papers that are already in the file, and a link that exists is left alone.

The Whole Request as One Query

hypercrux sql lit.db "
  SELECT p.year, p.title
  FROM json_each(walk('paper:10.1000/rev.2024.118', 2)) w
  JOIN paper p ON p.key = w.value
  WHERE p.year >= 2018 AND p.vec IS NOT NULL
  ORDER BY distance(p.vec, ?)
  LIMIT 20" "$(embed 'loneliness interventions for older adults')"

walk follows links two steps out, and every link in this file is a citation, so that’s the papers the review cites and then the papers those cite. json_each turns the list into rows. The WHERE clause keeps the recent ones that have a vector, and distance puts the closest to her question first. Here embed is a small command that prints the question’s vector as a JSON array, which distance accepts as readily as a stored vector.

Walking outwards finds older work, since a paper can only cite what came before it. To see what in her library cites the review, she walks the other way:

hypercrux walk lit.db paper:10.1000/rev.2024.118 1 --in --type cites

People who write systematic reviews call these two moves backward and forward snowballing. Here each one is a single command.

Why One File Suits a Thesis

Everything travels together in one file. She can back it up with the rest of her thesis or send a copy to her supervisor, who can open it in any SQLite browser and look around. When she rewrites her research question in month eight, she runs the query again with the new wording and has a new order in moments. Walks are cheap: two links out among 100,000 records took 0.12 milliseconds in the recorded benchmarks, and her graph is a small fraction of that size.

Where It Stops

The ranking is only as good as the embedding model and the abstracts behind it. It won’t judge a study’s quality, and it can’t see papers that aren’t in her library. What it gives her is a sensible order to read in, and every paper on the list meets the conditions she set, which she can check by hand whenever she likes.

The same pattern fits any collection where items point at each other and also carry text, such as patents and the patents they cite, or notes in a personal wiki and their backlinks. The quick start builds a small one in a few minutes.