Data Governance & Risk Checklist for Learning Use Cases
A practical, action-oriented checklist to evaluate privacy, consent, retention, access, bias, provenance, and operational risks when datasets are repurposed for organizational learning, analytics, or AI. Includes guidance on how to use the checklist, suggested mitigations, and a simple action template for assignment and follow-up.
Purpose
This checklist helps teams assess the legal, ethical, and operational risks when using datasets for learning, analytics, or AI. Use it before approving a dataset for training, research, or internal learning projects. For each item, record the current state, assign an owner, link supporting evidence, and list any required mitigation steps.
How to use this checklist
- Review each check and mark its status: Yes (Compliant), No (Problem), or Partial / Needs review.
- For any response other than Yes, capture a short remediation plan, an owner, and a target date.
- Estimate an overall risk score and decide whether to proceed, proceed with controls, or block use until resolved.
- Keep evidence links (policy documents, consent records, logs, datasets samples) with the record.
Core checklist items
-
PII presence
Is personally identifiable information present in the dataset? If yes, what categories (name, email, SSN, device ID, location, health data)? Action: where possible remove or pseudonymize; document legal basis for retention/processing.
-
Consent & purpose alignment
Is there documented consent or another lawful basis for the intended reuse? Does the purpose of reuse align with what users consented to? Action: obtain consent, narrow purpose, or anonymize data if consent cannot be aligned.
-
Retention policy and data lifecycle
Is there a defined retention period and deletion workflow for this dataset? Action: set retention, schedule secure deletion, and log deletions.
-
Access controls & segregation
Are access controls (least privilege, role-based access) applied? Are sensitive subsets segregated? Action: apply RBAC, just-in-time access, and approval gates before granting access.
-
Anonymization / de-identification thresholds
Are anonymization or de-identification techniques applied and tested against re-identification risk? Action: document technique, tested thresholds, and residual risk.
-
Training data provenance
Is the source of the data documented (origin system, collection method, date ranges, transformations)? Action: capture provenance metadata and a simple lineage diagram.
-
Labeling & data quality
Are labels accurate, consistent, and representative? Have labeler instructions and inter-rater agreement been recorded? Action: sample labels, compute quality metrics, retrain labelers if needed.
-
Known biases & fairness risks
Have known biases been identified (sampling bias, label bias, historical bias)? Is there an assessment of potential disparate impacts? Action: run bias scans, document mitigations and monitoring plans.
-
Redaction & sensitive fields
Are there fields that must be redacted (PHI, security tokens, proprietary secrets)? Action: apply redaction rules and validate outputs.
-
Logging, audit trail & monitoring
Are access, changes, and model training events logged and retained according to policy? Action: enable immutable logs and periodic review.
-
Escalation & incident response
Is there a documented escalation path for data incidents (contacts, SLA, regulatory notification triggers)? Action: confirm contacts and run tabletop exercises.
Action template (copy this into your project record)
Dataset name: __________ | Owner: __________
Overall risk rating (1–5): __________ | Decision: Proceed / Proceed with Controls / Block
Top 3 remediation actions:
- Action, owner, target date, evidence link
- Action, owner, target date, evidence link
- Action, owner, target date, evidence link
Next review cadence: Monthly / Quarterly / Annually / On significant change
Examples of common mitigations
- Pseudonymize identifiers, retain mapping keys only in a separate secure store.
- Limit training data to aggregated or sampled records where feasible.
- Apply synthetic data techniques when high-risk PII cannot be removed.
- Use differential privacy or noise injection for analytics where appropriate.
Notes on compliance & governance mapping
Map each checklist outcome to existing policies: privacy policy, data retention policy, acceptable use, and model governance. For regulated data classes (health, payment, children’s data), require legal review before reuse.
When to escalate
Escalate when the dataset contains unconsented PII, high-risk sensitive categories, or unresolved provenance gaps that prevent accountability. Escalation should involve data protection officers, legal, and security as appropriate.
Next steps for teams and librarians
- Embed this checklist into the project intake or dataset approval workflow.
- Keep an evidence folder (consent records, anonymization reports, labeler instructions) with the dataset record.
- Consider adding automated checks (DLP scans, metadata provenance checks) as the dataset enters the platform.
Discussion
Comments and conversation will live here.