Digital ExtensionsEvidence software
Practice corpus · Free

A records production
you can practice on.

Reviewing records is a skill that cannot be taught on three sample pages. The problems that make the work hard — a treatment gap you only see once you sort by date, two evaluators who examined the same person and disagree, a document that arrived on a later disc and never got read — do not appear at small scale.

So here is one at full scale. 1,050 documents, 2,664 pages, one fictional case, free to use and free to hand out.

Everything in it is fictional.

No real medical record, de-identified or otherwise, is in this pack or was used to make it. Every patient, provider, facility, employer, carrier, county, address, claim number, and medical record number was invented for it. The clinical content was written to be internally consistent, not to describe anyone.

Nothing here is medical or legal advice, and nothing here is an official medical record or court document. Do not file it, submit it, or represent any part of it as a genuine record.

License

CC BY 4.0

Copy it, put it on a thumb drive, hand it out at a chapter meeting, excerpt documents into a workshop packet, build a course around it, charge for that course. Keep the attribution line with it.

One thing the license does not require, and we ask anyway: keep the word synthetic next to it wherever it travels.


The whole case, as it actually arrives. One archive, one email address, one link that will still work next year.

Full production

The case as it actually arrives: two deliveries, sixteen producing parties, Bates-stamped, filenames that tell you nothing.

  • 1,050 documents
  • 2,664 pages
  • 255 MB zipped

It is 255 MB, so we send a link rather than a file.

Verify what you downloaded

SHA-256 of keller-production-v1.2.zip:

1b0545d16c912b9fc3b4ff0c73773dbea4f0b2ee78e8493f491647473abec0d5

Every file inside is listed with its own hash in manifest.json.

These are the exercises the pack was built for. All three run in a word processor and a PDF reader — no special software. That is deliberate: the point of running them by hand is to find out exactly where the method costs you something, before anyone tries to sell you a tool that fixes it. Give a room the corpus, a laptop each, and two hours.

  1. 01

    The one-motion test

    Can you get from a claim back to its proof in one motion?

    1. Build a short chronology in a table — date, event, source. Twelve to fifteen rows is plenty.
    2. Put it away for ten minutes and do something else.
    3. Come back, pick any row, and get to the exact page the entry came from. Not the document — the page.

    Time the first one. Then time the fifth.

    Almost everyone writes a source column that names a document and discovers on the way back that the document is nine pages and the sentence they meant is on one of them. The fix is obvious once you have felt the problem: cite the page, every time, while the document is still open in front of you. It costs about four seconds then and about four minutes later.

    What it catches. The cost of an imprecise citation, which is invisible while you are writing and unavoidable when you come back. Against the full production it gets much worse, because there the document is not called anything — it is called KEL-CIR-000412.pdf.

  2. 02

    The scavenger-hunt test

    Can someone else check your work?

    1. Swap chronologies with the person next to you.
    2. Each of you picks three entries from the other’s table and verifies them against the sources — same date, same fact, same page.
    3. Note two numbers: how long the three took, and how many you could not confirm at all.

    An entry that cannot be checked is not evidence of anything; it is an assertion with a date on it. And the person checking is not always friendly — sometimes it is opposing counsel, and sometimes it is you in fourteen months with no memory of what you meant.

    What it catches. A chronology that fails even when every single entry in it is correct. That is the part worth sitting with: accuracy and verifiability are different properties, and only one of them survives being handed to someone else.

  3. 03

    The conflict test

    Does the record agree with itself?

    1. Read the pack for one thing only: statements about what this man can and cannot do. Lifting limits, stairs, bathing and dressing, walking distance, work capacity.
    2. Write down each one with its date, its author, and where they were when they made it.
    3. There are four formal accounts of his function in the pack. Find all four before you go further.

    Then answer three questions. Which pairs actually disagree — reading the setting each assessment was made in before you decide, because two accounts written in different environments can differ completely and both be accurate. For the pair that genuinely conflicts, what would you need to resolve it, and is it in the pack? And if function changed between them, what changed it?

    Expect the room to split. That is the intended outcome, and the disagreement is the content.

    What it catches. The chronology that quietly picks a side. Both accounts are sourced and both are signed, and a summary that lists one and omits the other is not a summary — it is an argument, and one that will not survive being read next to the file.

The full instructions, including notes for whoever is running the session, ship as the README inside the archive.

The formats are
the real ones.

  • Born-digital PDFs, clean scans, faxes, Word files, spreadsheets, and a photograph. Records arrive the way the provider sent them, and a practice set that is all tidy PDFs teaches a skill nobody needs.
  • 412 files have no text layer at all. Search finds nothing in them until someone runs OCR. Discovering that a document was in the pack the whole time is one of the things the pack is for.
  • Filenames carry no information. In the full production a file is named for its first Bates number and nothing else. Building the index yourself is the first real step of the work — and the manifest is there for when you have made the point.
  • Two deliveries, three weeks apart. The supplemental disc arrives after the initial production, which is how a case actually gets a document nobody has read.

Cite it. The paths will hold.

Within version 1, documents may be added but are never renamed, moved, or removed. A reorganization ships as version 2 at its own URL, and version 1 stays live.

That is a promise worth having in writing, because it is what lets you build your own material on top of this one. A worksheet that says “open document 25, page 2” will still send people to that page next year. So will a video, a blog post, or a slide.

Enforced upstream rather than promised here: the build that packs these archives fails if a published path has been renamed, removed, or had its contents change. A reorganization cannot ship as 1.x — it becomes version 2 at its own URL, and this one stays up.

Ran the tests? That is the pitch.

The Keller Practice Corpus is free and stays free, whatever you use it with. If running the exercises by hand made the cost of the method obvious, Corpus Chronology is what we built to carry it — and it opens this exact production.