Synthetic Data

I have spent a lot of time thinking and writing about synthetic data. It is the subject of my PhD.

These posts cover how synthetic data is generated, how to measure whether it is any good, and the trade-offs between privacy and utility that nobody escapes. I write from health research but the same problems show up in finance, pharma, and anywhere the real data is sensitive.

Synthetic Data
graph_2 orange

Representativeness in Synthetic Data: What It Means and How to Measure It

Understanding the concept of representativeness in synthetic data and the methods used to measure it.

Read More
padlock blue

How Synthetic Data Is Used in Healthcare, Research and Beyond

Explore real-world use cases for synthetic data in healthcare, clinical trials, finance and more.

Read More
table green

Multiple Imputation and Perturbation: Why They're Not Built for Synthetic Data

This blog explores why multiple imputation and perturbation are not suitable for generating synthetic data.

Read More
dag_2 orange

What are GANs and how can they generate synthetic data?

This blog explores Generative Adversarial Networks (GANs) and how they can be used to generate synthetic healthcare data.

Read More
padlock green

What is Synthetic Data and Why Does it Matter?

This blog is the first in a series exploring synthetic data, its benefits, and its applications in various fields.

Read More