use cases

At 3 a.m., Find the Incident That Looks Like This One: On-Call Notes With Links and Similarity

Think of a platform team of six that keeps an online shop’s checkout running. Every incident gets a short write-up afterwards, saying what broke and what fixed it. After a few years that’s a few hundred pages in the team wiki, full of hard-won knowledge that nobody can find at 3 a.m.

The question at 3 a.m. is “has this happened before?” Wiki search matches words, and the words in an alert rarely match the words in a write-up. The alert says connection refused. The write-up from March says checkout errors after a database failover. They describe the same problem from opposite ends. And on a truly bad night the wiki may be down too, if it runs on the same infrastructure that’s failing.

An alert’s text finds the closest past incident on the payments service, and that incident’s fixed_by link leads to the runbook. A side note shows payments depending on the main database.

Incidents, Services and Runbooks

The file holds a record for each incident, with its date, severity, title and a vector made from the write-up. Services and runbooks are records too. Links carry what the team already knows: an incident affected a service and was fixed_by a runbook, and a service depends_on other services. A nightly job exports the wiki into the file and copies it onto every on-call laptop.

Has This Happened Before?

The engineer pastes the alert text into one query:

hypercrux sql oncall.db "
  SELECT i.key, i.title, i.date
  FROM json_each(walk('service:payments', 1, 'affected', 'in')) w
  JOIN incident i ON i.key = w.value
  WHERE i.vec IS NOT NULL
  ORDER BY distance(i.vec, ?)
  LIMIT 5" "$(embed checkout failing with connection refused)"

The walk goes one link backwards from the payments service to every incident that affected it, and the five closest to the alert come first. embed stands for a small command that runs the team’s embedding model and prints a JSON array. For this to work offline the model has to run on the laptop, and small open models run fine on a laptop’s processor.

From the Incident to the Fix

Say the March failover comes out on top. Its links lead straight to what fixed it:

hypercrux neighbours oncall.db incident:2026-03-14-db-failover --type fixed_by
incident:2026-03-14-db-failover -fixed_by-> runbook:db-failover

Then hypercrux get oncall.db runbook:db-failover prints the runbook itself, as JSON.

What Else Could Be Behind It

If the database looks healthy, the next question is what else payments leans on. The dependency links answer that, up to three steps out:

hypercrux walk oncall.db service:payments 3 --type depends_on
1  service:fraud-check
1  service:postgres-main
2  service:redis-cache

The number is how many links away each service is. Payments uses the fraud check directly, and the fraud check uses the cache, so a cache problem can show up as a payments alert two steps removed.

A Copy on Every Laptop

Everything is in one file, so getting it onto laptops is simple. It needs copying the safe way, though, because a plain copy of a file that’s being written can catch it halfway. SQLite’s VACUUM INTO writes a complete, consistent copy:

hypercrux sql oncall.db "VACUUM INTO 'oncall-copy.db'"

It won’t overwrite an existing file, so the job writes to a fresh name each night and ships that. The copy is a full HyperCrux file, and hypercrux check on it passes. On a laptop it works without a network, which at 3 a.m. is the point.

Limits Worth Knowing

A similar past incident is a lead. The engineer still has to check that today’s problem really is the same one. The search is also only as good as the write-ups behind it, and a two-line write-up gives it little to work with.

Speed won’t be the issue. The recorded benchmarks put a search among 1,000 vectors of 384 values at about 4 milliseconds on a two-core cloud machine, and a few hundred incidents is less than that.

Two Kinds of Knowing

The links hold what the team knows for certain, such as which runbook fixed which incident. The vectors catch what they only half remember. Kept together in one file, they turn a wiki full of write-ups into answers for whoever is holding the pager. The reference covers each command used here.