Original research draft · October 2026
15. Use Data and Automation Responsibly
15.1 Ask what a tool actually checked
A citation extractor recognizes patterns. A database lookup may match a reporter citation to a case. A link checker tests a destination. A citation manager formats stored metadata. Each tool can help, but none of these operations alone proves that the source supports the legal proposition.
Read a tool's result as a statement about its actual function. “Matched” should not become “legally valid.” “URL returned a page” should not become “the cited document is present.” Some websites return an ordinary successful HTTP response for an error page or access challenge. Inspect the returned content.
15.2 Distinguish cases, clusters, opinions, and citations
CourtListener's data separates court records, dockets, opinion clusters, individual opinions, and citation records. A cluster can group related opinions in one decision. Several citation strings can identify the same cluster. A single docket can contain more than one decision.
Counting citation strings therefore does not count distinct cases, and counting opinion records does not necessarily count distinct decisions. Before reporting a number, name the unit. Keep the database IDs needed to join records without confusing them with reader-facing legal citations.
For the manuscript's examples, preserve both the cluster identity and the specific opinion used. Deduplicate parallel citations at the cluster level when studying decisions, but retain each citation string when studying publication identifiers. The analytical question determines the correct unit.
15.3 Keep bulk data's limits visible
A bulk snapshot describes the provider's data at a stated time. It can contain missing fields, historical courts, duplicate representations, inconsistent source strings, or records added from different collections. A large dataset increases opportunities for research; it does not eliminate the need for source review.
The CourtListener court snapshot used in this project contains 3,361 records. That is a database-row count, not a count of currently operating United States courts. The live MCP court endpoint reported a different total during acquisition. The bulk and live sources were preserved separately rather than silently made to agree.
Citation-frequency statistics also need context. A high citation count can reflect many kinds of treatment. It does not itself measure approval, precedential weight, current validity, or pedagogical quality. If using frequency to select examples, describe it as a selection method and read the selected material.
15.4 Design a check that can fail
A useful verification process can produce “not found,” “ambiguous,” “wrong name,” “wrong version,” or “needs review.” If every input receives a polished citation, the system may be hiding uncertainty. Preserve failure states in the research record and show the writer what needs attention.
Check a known-good reference and a deliberately broken reference when evaluating a tool. A system that accepts both has not demonstrated useful validation. Keep the test focused on the claimed capability: a citation-identity test cannot establish substantive legal analysis.
For automated drafting, require a source packet for every real authority and a visible label for fictional examples. If a generated citation cannot be independently located, remove it from the authority list until it is resolved. Do not replace it with a plausible-looking reporter number.
15.5 Preserve provenance through transformations
When converting CSV or XML into JSON, retain the source URL, snapshot or publication date, retrieval date, raw file, and checksum. Document any normalization. Converting null and empty text to the same value may be harmless for a display but unacceptable for an analysis of missing data.
Keep source text separate from editorial explanations. This allows an editor to see which statements came from the issuing institution, which metadata came from a provider, and which instructions are Citation Code's recommendations. It also makes a future correction easier: the team can identify affected examples without reconstructing the entire research process.
Sources and convention notes
CourtListener bulk-data documentation; CourtListener case-law API documentation; GPO bulk-data repository. Counts describe the archived research files, not an independent audit of the providers' complete collections.