Do it for your family
This is how we rebuilt this story — from the youngest brother's lost name to the ship of arrival. The steps, with the traps we stepped on ourselves.
Gather what the family already has
Recorded testimonies (the Shoah Foundation holds ~55,000 — relatives can request a copy, and many don't know), drawer photos and papers, grandchildren's school projects, and the memory of the living.
The trap: Postponing the questions to the living — living memory is the most fragile archive.
Our example: Everything here began with one 1997 tape, one great-grandson's school PDF, and the family's memory.
I'll tell you everything my family knows about [name], a Shoah survivor. Organize it into a "starting dossier": (1) canonical data — names with EVERY spelling variant I mention, dates, places; (2) what is memory vs. document; (3) the 10 most important unanswered questions. Invent nothing; whatever I don't say, mark as "unknown".
Transcribe the testimony (local AI, strict rules)
Whisper runs on your own computer — the audio never leaves home. The rules matter more than the tool: verbatim is sacred (the survivor's "broken" grammar is a document), never guess names (doubt becomes [?]), two tracks (faithful + readable), timestamps on everything. Raw audio (processed audio is worse); reprocess critical passages in isolation.
The trap: Trusting the machine's first pass.
Our example: The machine wrote "they lived there"; it was "they died there" — the tape's most important sentence.
Transcribe this audio segment under STRICT rules: (1) absolute verbatim — keep the survivor's grammar slips, repetitions and hesitations, they are the document; (2) NEVER guess proper names: when in doubt write [?]; (3) [mm:ss] timestamps per turn; (4) "improve" nothing. Then, in a separate file, produce a readable edited track — same timestamps.
Structure before you search
One page of canonical data (names with spelling variants, dates, places) and a timeline where EVERY line declares its source: INTERVIEW hh:mm · DOCUMENT · HISTORY · INFERENCE. Preserve the survivor's contradictions.
The trap: "Correcting" the teller's memory.
Our example: His Stalingrad doesn't square with his unit — both versions stand side by side to this day.
From this transcribed testimony, build a timeline table where EVERY row declares its source: "INTERVIEW hh:mm", "DOCUMENT", "HISTORY" or "INFERENCE". If the survivor contradicts himself, record BOTH versions side by side with timestamps — don't choose for him.
The world's free archives
Arolsen Archives (no signup: DPs, camps) · Yad Vashem (names, deportations, nominal lists) · USHMM (survivors & victims database + free remote document requests) · JDC Archives · USC VHA's public record (it indexes names the tape doesn't say) · in Poland: Geneteka, JRI-Poland, Szukaj w Archiwach, state archives (the residents' kartoteka — our central find), address books, Yizkor books · in each country of the diaspora, its own. AI helps with spelling variants and with reading Polish, German and Hebrew.
The trap: Thinking Google is all there is.
Our example: Great-grandmother's maiden name sat on a 1946 Arolsen card; the youngest brother's name, in USC's indexing.
Generate every plausible spelling variant of [Mojsie Stobiecki] for archive searches: Polish, transliterated Yiddish, German, Russian, localized forms and common indexing errors (s/z/c, i/j/y, w/v swaps). Format as a plain list to paste into search fields. Then: read this Polish archive record [paste text/OCR] and translate it field by field, without interpreting — whatever is illegible, say it's illegible.
The method that guards against errors
A certainty grade on everything (confirmed / probable / hypothesis), written next to the find. Homonyms are the rule: verify with your own eyes — AI reads indexes, the proof is the scan. A negative is also a find: record where the search ends. And AI hallucinates: prompts with "candidate vocabulary" contaminate transcripts — run with and without, compare.
The trap: Closing the beautiful story before the proof.
Our example: TWO Ajzyk Stobieckis born in the same small town in 1901: we pinned a death on the uncle and corrected it to the cousin within 24h. We tell this error on purpose.
I found a record of [name] born [year] in [town]. Before we accept it's ours: (1) list what other bearers of the same name could exist in that region and era; (2) say which record fields would tell homonyms apart (parentage, exact date, address); (3) which document I would need to see WITH MY OWN EYES to confirm. Play devil's advocate: your job is to stop me from pinning the beautiful story on the wrong person.
Write to the archives (there are humans on the other side)
Identify yourself as family, cite exact references (that's what the online research is for), one clear question per request. And the act any family can do — the most important: submitting Pages of Testimony at Yad Vashem for those with no name anywhere. What AI doesn't do: accounts, captchas, forms — those clicks are yours, and rightly so.
The trap: Asking for "everything about the family" — vague requests die in the queue.
Our example: Ten formal requests await the family's signature — including the one that may settle the father's fate.
Draft an email in [English/Polish] to the [archive name], from a family member requesting [a copy of document X]. Rules: identify me as [name]'s [grandchild], cite the exact references I already found online [paste call numbers], ONE clear question, respectful and short. Later, translate their reply for me when it arrives.
Ethics and limits
Privacy of the living; don't impulsively contact those who filed memorial pages (that's a family decision); source rights when republishing; respect for what the survivor chose not to tell. And honesty above closure: "we don't know" is a worthy answer.
The trap: Turning someone's story into content.
Our example: Our "What we don't know" page exists precisely for this.
Preserve and tell
A catalog with stable IDs and provenance; source files in plain text (durability); backup in three places; and finally the telling — a site, a book, a family film.
The trap: Leaving it all on a single computer.
Our example: This entire site.
The one-page checklist
- Record the living NOW — a phone audio is enough
- Request the testimony copy from the Shoah Foundation (relatives can)
- Transcribe verbatim: nothing corrected, nothing guessed, [?] when in doubt
- One page of canonical data with spelling variants
- A timeline with the source declared on every row
- Arolsen · Yad Vashem · USHMM · JDC · state archives of the home country
- A certainty grade written next to every find (confirmed/probable/hypothesis)
- A homonym until proven otherwise: confirm with your eyes on the scan
- Write to the archives: family, exact reference, one question
- Three copies of everything; a catalog with IDs; and tell the story
Print this page: the checklist and the prompts were made to live next to the keyboard.
We began with a 1997 tape and a name the grandfather himself no longer remembered. We ended with the name — Majer —, with a sister no one knew had existed — Frania —, and with the ship, the street, the passenger list. The archives kept what memory could not. Go find yours.
Tools: we name categories on purpose. What we used here: local whisper (large) to transcribe, an AI assistant to organize/translate/draft, and the family's eyes to confirm every scan.