Skip to reference
The Citation CodeDownload PDF
Reference

Source forms / F077 · Digital and data

Dataset

Original research draft · October 2026

Working pattern

[Creator], [Dataset title] ([version/snapshot]), [repository/identifier], [table/variable].

Capture from the source

Creator; version; unit; date; variables; filters; stable identifier

Check before you use it

State the denominator and transformations when reporting a derived result.

Common error: A dataset citation alone does not disclose filtering or missing-data choices.

Completed example

Lab, dataset, identifiers, variables, and records are fictional. The arithmetic follows supplied training counts; it is not an empirical finding about actual courts, agencies, or citation systems.

Northmere Research Lab, Register Match Study (snapshot TRAIN-DATA-77-v2, May 1, 2025), tasks table, fields task_id and matched, https://training.example/data/TRAIN-DATA-77-v2.

Read the supplied source facts

Fictional packet TRAIN-W77 supplies the Northmere Research Lab's Register Match Study dataset, snapshot TRAIN-DATA-77-v2 dated May 1, 2025. Its tasks table has 120 rows. Twenty rows lack the outcome field, and eight additional rows are duplicate task identifiers; these excluded groups do not overlap. The remaining 92 unique completed tasks include 69 marked matched. The assignment asks for the match rate among unique completed tasks. A dashboard elsewhere reports 69 of 120 without explaining its denominator. The packet defines each variable and supplies the exclusion log and row identifiers.

Why this works

The citation identifies the dataset and snapshot; the prose must also disclose the calculation. Under the supplied exclusions, 120 minus 20 minus 8 leaves 92 tasks, and 69 divided by 92 equals 75%. State that this is a derived rate among unique completed tasks, not a rate among all rows or a result directly reported by the source. Preserve the exclusion criteria, missing-data treatment, and transformation log so another researcher can reproduce the denominator. Dataset identity alone does not explain analytic choices, and the broad dashboard figure answers a different question.

When the facts change

If missing outcomes are treated as nonmatches, explicitly describe that alternative and recalculate the denominator rather than silently changing the percentage. If duplicates overlap missing rows, the simple subtraction used here would no longer be justified without a row-level reconciliation. If a newer snapshot adds tasks, cite it separately and avoid comparing rates until the inclusion rules and unit of analysis are aligned.

Apply the receiving court or publication’s requirements. These working forms do not establish citation permission, precedential weight, or support for a legal proposition.