Skip to content

Research and Source Management

Dataset Documentation Card

Describe a dataset's origin, fields, collection method, exclusions, transformations, limitations, license, privacy conditions, and exact version.

Free editable Markdown · Data teams, researchers, and editors ·

Download Markdown

Accessible HTML preview

Blank template

The downloaded file contains the same fields in editable Markdown.

Identity and provenance

Dataset title
[Exact name]
Version or release
[Stable identifier]
Responsible owner
[Organization and contact role]
Original purpose
[Why the data was collected]
Source location
[Authorized URL or repository]
Collection period
[Start and end]
Geographic or system coverage
[Included scope]
Unit of observation
[Meaning of one record]
Collection method
[Observed, reported, inferred, generated, or joined]
Refresh schedule
[Frequency and lag]

Fields and processing

Field name
[Name] — **Meaning/unit:** [Definition] — **Missing code:** [Value]
Derived field
[Name] — **Formula:** [Calculation and inputs]
Record count received
[Number]
Deduplication
[Key, rule, and removed count]
Exclusions
[Rule, reason, and count]
Joins
[Source, key, unmatched records]
Manual changes
[Who, why, and audit location]
Analysis-ready version
[Checksum, tag, or dated filename]

Governance and fitness checks

  • Collection purpose and record-entry process are documented.
  • Population coverage and missing groups are stated.
  • Units, categories, dates, and time zones are explicit.
  • Transformations can be reproduced from stored inputs.
  • Personal or sensitive fields have approved handling.
  • License, attribution, retention, and sharing terms are recorded.
  • Known bias, error, lag, and unsuitable uses are visible.
  • Version changes and update ownership are documented.

How to use this template

  1. Identify the exact dataset, owner, purpose, collection process, coverage, unit, period, and version.
  2. Document fields, units, allowed values, missingness, and derivations used in the analysis.
  3. Record every cleaning, join, exclusion, and transformation with before-and-after counts.
  4. Assess methodological limits, privacy, permission, license, retention, and unsuitable uses.
  5. Review the card with data and domain owners, then publish or share only the approved version.

Describe provenance and collection

Name the organization and team responsible for collecting or assembling the data, the original purpose, the collection period, and the process by which a record enters the dataset. State whether data is observed, reported, inferred, generated, or joined from other sources. Identify coverage and the unit represented by one row. A dataset created for administration may omit events outside that system, while a voluntary survey may overrepresent people willing or able to respond. These conditions shape every later claim.

Make fields and transformations interpretable

Provide a compact data dictionary for fields used in analysis, including type, unit, allowed values, missing-value code, and derivation. Record cleaning, deduplication, normalization, joins, exclusions, and manual decisions as reproducible steps. Explain whether dates use event time or entry time, which time zone applies, and how currencies or categories were harmonized. Do not interpret blank, zero, “unknown,” and “not applicable” as interchangeable. Track the number of records before and after material transformations.

State limits, rights, and safe uses

Document known gaps, measurement error, population limits, update lag, possible bias, and uses the data cannot support. Record the license or permission, attribution requirement, retention rule, and whether personal or sensitive information is present. Publish only the documentation and fields approved for that audience. Version the card with the dataset, assign an owner, and list changes since the previous release. A reader should be able to understand why a seemingly available analysis may still be irresponsible or invalid.

See the fields in context

Fictional example: bicycle stand observations

The Port Lark stand survey and every value are invented; this is documentation structure, not a real dataset.

  • Unit: One fictional manual observation of a designated bicycle stand during a scheduled visit.
  • Coverage: Twelve invented stands observed on four dry weekdays; weekends and severe weather are excluded.
  • Field: `occupied_spaces` is a nonnegative count recorded at visit time, with blank reserved for unreadable forms.
  • Transformation: Stand IDs are standardized; one duplicate fictional form is excluded with its record ID logged.
  • Limitation: Observations describe scheduled moments and cannot estimate every day's peak demand.

Frequently asked questions

Is a data dictionary enough documentation?

No. Field definitions are essential, but users also need provenance, collection method, processing, limitations, permissions, and version information.

Should excluded records be deleted?

Preserve the approved raw source separately and document exclusions in the working pipeline. Do not erase auditability or retain data contrary to applicable rules.

What is an unsuitable use?

It is a decision or claim the dataset cannot responsibly support because of coverage, method, granularity, rights, or risk—for example, inferring individuals from aggregate observations.

When should the card receive a new version?

Update it whenever the dataset, schema, collection method, transformation, license, known limitation, or approved use changes materially.

File details

File name
dataset-documentation-card.md
Format
Markdown (.md)
Size
3 KB
Designed for
Data teams, researchers, and editors

Usage note: Create one card for the exact dataset version used in analysis, not merely for a project with a similar name. Link to detailed technical documentation when it exists and keep the card beside the data and analysis files. This editorial aid does not replace a data protection assessment, consent process, security review, license analysis, or formal model card.