Synthetic Data & Privacy-Preserving Methods Evaluation Checklist

Interactive checklist and evaluation form to determine whether synthetic or other privacy-preserving approaches are appropriate, document selected privacy metrics, capture empirical utility tests, and record governance, release controls, and known limitations.

Interactive Tool

Synthetic Data & Privacy-Preserving Methods Evaluation

Use this guided evaluation to decide whether a privacy-preserving approach (synthetic data, anonymization, differential privacy, or hybrid methods) is appropriate for a dataset and to record the trade-offs, tests, governance approvals, and next steps. Save the results to retain an auditable decision record and help teams validate real-world utility before relying on synthetic-only training or sharing.

YYYY-MM-DD or free-text date.
Describe the measurable purpose (e.g., model training, analytics testing, partner sharing) and success criteria. Be specific about which analyses or model behaviors must be preserved.
Choose the most appropriate classification for this dataset according to your policies.
Confirm classification, retention, and permitted uses are documented and approved.
Check all that apply. If hybrid, describe below.
Explain generator type, anonymization rules, DP mechanism, or hybrid design. Include links to implementation repo or config if available.
Select metrics you will measure to assess privacy guarantees and re-identification risk.
Enter numeric epsilon value used or planned. Leave blank if DP not used. Note: smaller epsilon indicates stronger privacy but lower utility.
1 = negligible, 5 = high. Attach details in Known limitations and test results.
1.0 10.0
Select the utility tests used to verify analytical value remains acceptable.
Summarize statistical differences, model metric deltas (AUC, accuracy, MAPE, etc.), and any observed failures or regime shifts. Include links to detailed test reports.
Enter percent change (positive or negative) between models trained on real vs synthetic data for key metrics. Use absolute percent where possible.
Which tools, environments, or pipelines will this synthetic/anonymized dataset need to work with?
Includes Data Protection Officer, Legal, Information Security, and business data owner approvals as required.
List people who approved and where approvals are recorded (tickets, signed forms, policy records).
Select controls that will accompany data release.
Describe how the data will be shared, with whom, and any contractual or technical safeguards.
Be explicit about any analytical weaknesses, rare-event failures, potential bias amplification, or untested slices.
Scripts or notebooks that reproduce the utility/privacy checks should be provided where practical.
Provide a link to CI, repo, or artifact store where tests live.
Select the team recommendation based on the evidence collected.
Explain why the recommendation was chosen and what acceptance criteria must be met to change it.
Team or person accountable for follow-through.
Person responsible for monitoring re-identification risk and utility drift.
1 = low priority, 5 = urgent. Use to plan governance and testing resource allocation.
1.0 10.0
Name and role of the person completing this evaluation.
Optional notes that auditors or governance reviewers will find useful.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.