Open Science Lunch – Can Synthetic Data Make Research Data More Open?

Opportunities and Privacy Risks in the Age of AI. With Qinyi Liu, UiB.

Access to research data is essential for transparency, reproducibility, and scientific collaboration, yet sharing data can be difficult when datasets contain sensitive or personal information. Synthetic data have increasingly been proposed as a way to address this tension: instead of simply releasing the original records or anonymization, researchers can use statistical and AI-based methods to generate artificial data that reproduce useful characteristics of the original dataset.

In this talk, I will introduce what synthetic data are, how recent advances in artificial intelligence have changed the way they are generated, and where they may support more open and reusable research data. Using concrete research examples, I will illustrate how synthetic data can be incorporated into research workflows, for instance to enable data exploration, method development, collaboration, teaching, and selected forms of data sharing without providing direct access to the original sensitive records. I will also discuss how researchers can assess whether synthetic data are suitable for these purposes by evaluating their statistical fidelity, analytical utility, privacy risk, and fairness. At the same time, synthetic data have important limitations: they are not automatically anonymous, may reproduce biases or sensitive patterns present in the source data, and may not preserve all of the information required for valid scientific conclusions. The talk will therefore consider both the opportunities synthetic data offer for privacy-conscious open science and the practical questions researchers should ask before generating, sharing, or reusing them.

Når: 24.09.26 kl 12.00–13.00
Hvor: www.ub.uio.no/english/courses-events/events/open-science-lunch/
Sted: Digitalt
Målgruppe: Ansatte, Studenter
Ansvarlig: Bror-Magnus S. Strand
Legg i kalender