Research and Source Management
Dataset Documentation Card
Describe a dataset's origin, fields, collection method, exclusions, transformations, limitations, license, privacy conditions, and exact version.
Free editable Markdown · Data teams, researchers, and editors ·
Accessible HTML preview
Blank template
The downloaded file contains the same fields in editable Markdown.
Identity and provenance
- Dataset title
- [Exact name]
- Version or release
- [Stable identifier]
- Responsible owner
- [Organization and contact role]
- Original purpose
- [Why the data was collected]
- Source location
- [Authorized URL or repository]
- Collection period
- [Start and end]
- Geographic or system coverage
- [Included scope]
- Unit of observation
- [Meaning of one record]
- Collection method
- [Observed, reported, inferred, generated, or joined]
- Refresh schedule
- [Frequency and lag]
Fields and processing
- Field name
- [Name] — **Meaning/unit:** [Definition] — **Missing code:** [Value]
- Derived field
- [Name] — **Formula:** [Calculation and inputs]
- Record count received
- [Number]
- Deduplication
- [Key, rule, and removed count]
- Exclusions
- [Rule, reason, and count]
- Joins
- [Source, key, unmatched records]
- Manual changes
- [Who, why, and audit location]
- Analysis-ready version
- [Checksum, tag, or dated filename]
Governance and fitness checks
- Collection purpose and record-entry process are documented.
- Population coverage and missing groups are stated.
- Units, categories, dates, and time zones are explicit.
- Transformations can be reproduced from stored inputs.
- Personal or sensitive fields have approved handling.
- License, attribution, retention, and sharing terms are recorded.
- Known bias, error, lag, and unsuitable uses are visible.
- Version changes and update ownership are documented.
How to use this template
- Identify the exact dataset, owner, purpose, collection process, coverage, unit, period, and version.
- Document fields, units, allowed values, missingness, and derivations used in the analysis.
- Record every cleaning, join, exclusion, and transformation with before-and-after counts.
- Assess methodological limits, privacy, permission, license, retention, and unsuitable uses.
- Review the card with data and domain owners, then publish or share only the approved version.
Describe provenance and collection
Name the organization and team responsible for collecting or assembling the data, the original purpose, the collection period, and the process by which a record enters the dataset. State whether data is observed, reported, inferred, generated, or joined from other sources. Identify coverage and the unit represented by one row. A dataset created for administration may omit events outside that system, while a voluntary survey may overrepresent people willing or able to respond. These conditions shape every later claim.
Make fields and transformations interpretable
Provide a compact data dictionary for fields used in analysis, including type, unit, allowed values, missing-value code, and derivation. Record cleaning, deduplication, normalization, joins, exclusions, and manual decisions as reproducible steps. Explain whether dates use event time or entry time, which time zone applies, and how currencies or categories were harmonized. Do not interpret blank, zero, “unknown,” and “not applicable” as interchangeable. Track the number of records before and after material transformations.
State limits, rights, and safe uses
Document known gaps, measurement error, population limits, update lag, possible bias, and uses the data cannot support. Record the license or permission, attribution requirement, retention rule, and whether personal or sensitive information is present. Publish only the documentation and fields approved for that audience. Version the card with the dataset, assign an owner, and list changes since the previous release. A reader should be able to understand why a seemingly available analysis may still be irresponsible or invalid.
See the fields in context
Fictional example: bicycle stand observations
The Port Lark stand survey and every value are invented; this is documentation structure, not a real dataset.
- Unit: One fictional manual observation of a designated bicycle stand during a scheduled visit.
- Coverage: Twelve invented stands observed on four dry weekdays; weekends and severe weather are excluded.
- Field: `occupied_spaces` is a nonnegative count recorded at visit time, with blank reserved for unreadable forms.
- Transformation: Stand IDs are standardized; one duplicate fictional form is excluded with its record ID logged.
- Limitation: Observations describe scheduled moments and cannot estimate every day's peak demand.
Frequently asked questions
Is a data dictionary enough documentation?
No. Field definitions are essential, but users also need provenance, collection method, processing, limitations, permissions, and version information.
Should excluded records be deleted?
Preserve the approved raw source separately and document exclusions in the working pipeline. Do not erase auditability or retain data contrary to applicable rules.
What is an unsuitable use?
It is a decision or claim the dataset cannot responsibly support because of coverage, method, granularity, rights, or risk—for example, inferring individuals from aggregate observations.
When should the card receive a new version?
Update it whenever the dataset, schema, collection method, transformation, license, known limitation, or approved use changes materially.