Model or dataset
mattprusak/autoresearch-genealogy avatar
mattprusak/autoresearch-genealogy

autoresearch-genealogy: 13 prompts with built-in verification for family history

Structured prompts, vault templates, and archive guides for AI-assisted genealogy research. Built for Claude Code.

1,182 stars120 forksRubyMIT

At a glance

What is it?
A prompt library, an Obsidian vault template and 24 archive guides, extracted from a real nine-generation genealogy project. Useful, and unusually careful about evidence.
Who is it for?
This repository is a good fit if you are doing genealogy with an AI assistant and care more about not inventing ancestors than about moving fast. The 13 prompts each carry a Verify condition and guard rails, the reference directory documents confidence tiers and source hierarchy, and the worked examples show what a session actually looks like.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 99 days ago.
What is it written in?
Mainly Ruby, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A prompt library, not a piece of software

The framing in the first paragraph is the honest one: structured prompts, vault templates and research workflows for AI-assisted genealogy research, built for Claude Code and adaptable to any AI tool or manual workflow.

That last clause matters more than it looks. Nothing here is a library you import, and the repository's own language field says Ruby on a codebase that is almost entirely markdown files. There is a `scripts/` directory, a `spec/` directory and a `CHANGELOG.md`, so some automation exists, but the deliverable is text you paste into a tool, not a package with a versioned API.

So the unit of this project is a prompt file. Thirteen of them, each a self-contained instruction set. The value is in what those instructions contain, not in anything the repository executes for you.

Every prompt carries a Metric and a Verify condition

The prompts are designed for Claude Code's `/autoresearch` command, and the README describes what that means: they run autonomously, searching the web, browsing image archives, updating your vault and verifying their own work. What makes them different from a list of clever queries is their structure. Each prompt defines inputs to replace, then a Goal, a Metric, a Direction, a Verify condition, Guard rails, Iterations and a Protocol.

A Metric turns a vague goal into something checkable. A Verify condition is what stops the loop from declaring victory on its own terms. In a domain where the failure mode is a confident, well-formatted, entirely wrong ancestor, that pairing is the whole idea.

The README's own philosophy section states it as structured autonomous research with mechanical verification, not AI guessing. The reasoning given is that genealogy has no compiler: sources disagree with each other, confidence is probabilistic rather than binary, and a name can appear as Sakkarias in one record and Zacharias in another. That is an accurate description of why a generic research agent fails at this task.

The thirteen prompts, and the order they are meant to run in

The catalogue covers a lot of ground. 01-tree-expansion reviews source-backed candidate relationships for deceased ancestors. 02-cross-reference-audit finds and fixes discrepancies between your tree and your documents. 03-findagrave-sweep locates memorials for every deceased ancestor. 04-gedcom-completeness makes sure your GEDCOM file matches your vault data. 05-source-citation-audit verifies that every person file cites at least two independent sources.

Then come the research-shaped ones: 06-unresolved-persons for unnamed people mentioned in your documents, 07-timeline-gap-analysis for life events where records should exist but have not been found, 08-open-question-resolution to attack every open question systematically, and 09-bygdebok-extraction for digitised local history books in any country.

The last four are the most specialised. 10-colonial-records-search targets colonial American ancestors in pre-1800 records, 11-immigration-search looks for passenger manifests and naturalisation records, 12-dna-chromosome-analysis maps genetic segments from per-chromosome ancestry data, and 13-image-archive-deep-dive browses image-only archives and saves cropped evidence.

The Quick Start is explicit about sequence, and it is the opposite of what most people would guess: run one of the two audits before the expansion prompt.

The vault template is 19 files that work outside Obsidian

The vault template is described as a complete Obsidian vault starter kit with YAML frontmatter and plain markdown, readable anywhere. Nineteen files split into core files and templates.

The core files are the working surface: family tree, research log, open questions, data inventory, timeline, genetic profile, chromosome painting, witness network, unresolved persons and research strategy. The templates are person, transcription, certificate, postcard, region, surname, hypothesis and draft letter.

Two design choices are worth calling out. A witness network file is not a stock Obsidian feature, and it suggests the vault is organised around provenance, who vouches for what, rather than around a tree view. A hypothesis template sitting alongside person files is the other: the vault has somewhere to put a claim that is not yet supported, which is the part most family history tooling leaves out.

If you do not use Obsidian, there is a dedicated No-Obsidian Setup guide, and the README is clear that the template works as a normal folder of markdown files. That is worth believing only after you check, since a vault full of wikilinks and YAML frontmatter usually depends on its editor.

Archive guides for 24 countries, and what they do not cover

The archives directory holds 24 country and region guides, each covering where to find records, what is free versus paid, and which parts an AI tool can reach directly versus which need a browser.

Europe is the deepest: Ireland, England and Wales, Scotland, France, Italy, Spain and Portugal, Germany, the Netherlands, Austria, Hungary, Norway, Sweden, Poland, and Russia and Ukraine. The Americas cover the USA across colonial, immigration, census and vital records, plus African American records, Canada, and Mexico and Latin America. Oceania covers Australia and New Zealand. One guide is cross-national: Jewish genealogy.

The split between what a tool can access directly and what needs a browser is the useful part, because it is the difference between an agent that can finish a task and one that stalls at a login. It also implies a limitation the README does not dwell on: much of this material sits behind institutional subscriptions, and the guides tell you where to look rather than providing access.

Alongside sit 11 reference documents on confidence tiers, source hierarchy, vault file manifest, DNA interpretation guard rails, naming conventions including patronymics, farm names and przydomki, GEDCOM format, common pitfalls, a glossary and an AI capabilities assessment. There are also 8 workflow guides, from getting started and OCR pipelines to oral history protocol and discrepancy resolution.

Privacy rules for living people are built in, not bolted on

Step three of the Quick Start instructs you to mark living or possibly living people clearly and to avoid exact birth dates for them. Later it points at a Privacy Mode guide before you use public AI tools or share exports, and a share-safely checklist sits with the other printable ones.

For a repository whose whole mechanism is uploading family documents to a model provider, that is the correct default to build in rather than recommend. The DNA prompt deserves the same attention: prompt 12 analyses per-chromosome ancestry data, and living people's genetic data is sensitive in a way that nineteenth-century census records are not.

The privacy posture is one of the stronger arguments for using this project as written. For a dry run without touching your own family, the First Run Walkthrough in the walkthroughs directory uses a synthetic fixture, which is the sensible way to see what a session looks like before pointing it at living relatives.

Editorial conclusion

This repository is a good fit if you are doing genealogy with an AI assistant and care more about not inventing ancestors than about moving fast. The 13 prompts each carry a Verify condition and guard rails, the reference directory documents confidence tiers and source hierarchy, and the worked examples show what a session actually looks like. Two things to settle before you commit an afternoon to it. The tree reports Ruby as the primary language on a repository that is almost entirely markdown, which suggests the `scripts/` directory holds whatever automation exists and that the rest is content, not code you can depend on as a library. And the README describes 24 archive guides without naming paywalled databases, so expect to supply your own subscription for several countries. Start with `START_HERE.md`, then run one of the two audit prompts before any expansion prompt.

Frequently asked questions

Can ChatGPT be used for genealogy research?

The project says it is built for Claude Code's autoresearch command but adaptable to any AI tool or a manual workflow. Its own position is that genealogy has no compiler, sources disagree and confidence is probabilistic, so each of its 13 prompts pairs a Goal with a Metric and a Verify condition.

Do I need Obsidian to use autoresearch-genealogy?

No. The vault template is described as YAML frontmatter and plain markdown that is readable anywhere, and the repository includes a No-Obsidian Setup guide for people who do not use that editor.

Which countries do the archive guides cover?

There are 24 guides. Europe is covered in most depth, including Ireland, England and Wales, Scotland, France, Italy, Germany, the Netherlands, Poland and others. The Americas, Canada, Mexico and Latin America, Australia and New Zealand, and Jewish genealogy across national borders are also covered.

Which prompt should I run first in autoresearch-genealogy?

The Quick Start says to run either 02-cross-reference-audit or 05-source-citation-audit before 01-tree-expansion. It also sends you to START_HERE.md first, which routes you by what you already have: names, documents, DNA results, a tree or a finding to verify.

Official sources

  1. Issues
  2. License: MIT
  3. mattprusak/autoresearch-genealogy on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mattprusak-autoresearch-genealogy.svg)](https://hysenlabs.com/projects/mattprusak-autoresearch-genealogy)