In recent years, the debate on synthetic data, synthetic personas and digital twins has taken on a central role in the world of market research.
Researchers, institutes and clients now find themselves at a crossroads similar to the one experienced with the advent of the first online surveys: on one side enthusiasm and high expectations, on the other doubts, skepticism and questions about the quality of artificially generated results.
In his recent ESOMAR webinar, Ray Poynter — together with contributions from Jon Puleston and Rupert van Hullen in the Global Research Software 2025 — outlined a clear and pragmatic picture of the opportunities, limits and risks of new AI-based data generation technologies.

 

What Are Synthetic Data According to ESOMAR

In June 2025, the ICC/ESOMAR Code introduced an official definition of synthetic data: data artificially generated to replace what would normally be collected directly from people. These are datasets that replicate distributions, patterns and relationships typical of real data, not simple random simulations. According to ESOMAR, today there are two main approaches:

  1. Model-based tabular synthesis
    Based on statistical algorithms, machine learning, GANs, Bayesian models or diffusion techniques: it creates synthetic records that preserve the structure of the real dataset.
  2. LLM-generated data
    Includes:
    synthetic personas, conversational agents that respond like human beings; digital twins, dynamic digital versions of the behavior of real individuals.

 
These latter approaches, although widely discussed, are still in their early stages and require more solid evaluation criteria.
 

Why Are Synthetic Data Such a Hot Topic Today?

The supply is growing rapidly: in addition to established players such as Ipsos, Kantar and Toluna, specialized providers like Fairgen and LivePanel are emerging, offering solutions based on generative AI and advanced modeling.
Companies’ attitudes toward synthetic data today fall into three categories:

  • those who already actively use synthetic data;
  • those who reject them outright, fearing loss of quality;
  • those who wait for more solid evidence before adopting them.

 

As Poynter and Puleston point out, this phase resembles the “wild west” of new technologies: great opportunities, but also hype risk, not always replicable results and excessive promises.
 

Main Types of Synthetic Data

In the context of market research, we can distinguish four categories of synthetic data:

  1. Boosting / Augmenting / Imputation
    Artificially increasing parts of the sample to rebalance underrepresented quotas or compensate for missing data.
  2. Fully synthetic data
    Datasets built from scratch using models that mimic the real population.
  3. Anonymized synthetic data
    Real data transformed to protect privacy while maintaining valid statistical properties.
  4. Randomized synthetic data
    Data generated with controlled logic, useful for technical tests and validations.

 
Today, the most mature area is boosting, thanks to its concrete application in managing hard-to-reach samples.
 

The Most Common Case: Sample Boosting

Let’s look at a typical boosting example: a study requires 500 interviews for each segment—young men, young women, adult men, adult women. If only 350 young men are collected, how should we proceed?

The options are well known: accept underrepresentation, weight the data, or generate synthetic data using statistical or AI techniques.

Traditional Statistical Methods

Before generative AI, approaches such as:

  • random sampling
  • nearest neighbours
  • SMOTE (Synthetic Minority Oversampling Technique)

 
These methods create new “hybrid” cases based on the most similar subjects and still work very well for numerical and structured datasets.

Boosting with Advanced AI

Today, the integration of machine learning and large historical panels enables much more refined solutions. Some case studies presented at ESOMAR LATAM show how these boosting techniques have performed with average errors between 1% and 3%.
 

How to Evaluate the Quality of Synthetic Data

When evaluating the use of synthetic data, we need to ask two fundamental questions:

Does it work in general with your type of data?
To answer, past studies can be used: remove part of the collected interviews, generate the boost, and compare the new result with the previously collected real data.

Does it work for this specific project?
A holdout sample can be used: a real subsample not used by the model, useful for comparing the quality of synthetic boosting. These techniques apply to boosting but become more complex for fully synthetic datasets, because there is no direct reference to a real population.
 

Limits and Considerations When Using Synthetic Data

Despite promising applications, Poynter urges caution.

  1. Temporal effects
    Models depend on input data. If the data is outdated, results will be distorted.
    Main risks: seasonality, long-term trends, disruptive events (e.g., pandemic, launch of a new device, AI).
  2. Impossible to estimate margin of error
    With fully synthetic data, sampling error cannot be discussed. Bootstrapping measures only model stability, not its accuracy toward the real population.
  3. Excessive promises
    Vendors claiming “triple accuracy” should be treated cautiously. Recent history shows that even seemingly solid models can contain bias (as seen in distorted facial reconstruction models).
  4. Data cannibalization
    Using synthetic datasets as the basis for generating new synthetic datasets leads to model drift, model collapse and loss of realism.

 

The World of Synthetic Personas

If synthetic data represent the quantitative side, synthetic personas are the qualitative component of the revolution.
They are intelligent agents that answer questionnaires, participate in qualitative discussions, generate insights, and allow testing messages and creativity.
They are increasingly designed as digital twins, dynamic digital replicas of real consumers. Digital twins do not simply represent a target: they replicate decision-making processes, preferences and attitudes, enabling much more sophisticated predictive simulations.
 

Ipsos Case in Japan: What Digital Twins Can (and Cannot) Do

An Ipsos study on 150 Japanese women made it possible to build individual digital twins and compare their responses with real ones in a follow-up study on topics related to the menstrual cycle.
Key findings:

  1. Main themes were consistent between real individuals and twins.
  2. Emotional depth was lower in twins.
  3. In creative phases, twins generated many ideas, but not always relevant.
  4. In final evaluations, twins favored rational benefits, while real people often choose based on emotion.

 
Conclusion:
Twins replicate structure and direction well, but less so human sensitivity.
 

Risks and Critical Issues of Digital Twins (according to van Hullen)

The power of digital twins comes with a number of risks that should not be underestimated, as highlighted by Rupert van Hullen in the ESOMAR Global Research Software 2025.
The need to handle large amounts of personal information increases the risk of data leakage, meaning sensitive details may be inadvertently exposed due to system vulnerabilities or data management errors. This is compounded by the risk of adversarial attacks, where malicious actors attempt to extract confidential information or manipulate models.

Another critical issue is the presence of special categories data such as political opinions, religious beliefs or health information. These are highly sensitive and require a higher level of protection and strict handling procedures.
Open-ended responses also represent a weak point: they often contain more personal information than expected, leading to forms of unintended disclosure that are difficult to fully control.

Finally, the very complexity of the pipelines needed to build digital twins multiplies potential vulnerability points: each new processing or storage step expands the risk surface.

To address these challenges, ESOMAR highlights the importance of a structured mitigation approach.
Advanced anonymization techniques, such as adversarial anonymisation, significantly reduce traceability. Architectures with strict access controls and data segregation help limit exposure to sensitive information. All of this must be accompanied by a privacy-by-design approach and continuous ethical review, essential to monitor risk evolution and ensure responsible use of digital twins.
 

Opportunities of Synthetic Data for Market Research

There are several areas where synthetic data and personas are already proving useful, such as:

  • rapid testing of concepts, messages and creativity,
  • questionnaire optimization before fieldwork,
  • integration of hard-to-reach or costly samples,
  • predictive simulations and what-if analysis,
  • qualitative pre-screening through personas or digital twins.

 
Synthetic data do not replace the entire research pipeline: they are complementary tools that accelerate work and reduce costs in the early phases of projects.
 

A Revolution Already Underway

The central message is clear: synthetic data are already here, they work, and they will grow rapidly.

Boosting and tabular synthesis are currently the strongest areas.
Synthetic personas and digital twins represent the most innovative frontier — rich in opportunities but also limitations, especially in terms of emotional and human sensitivity.

The right approach is to experiment, test, validate, avoid promises of absolute realism, and not fully replace real data.

Just as happened with online surveys twenty years ago, the adoption of synthetic data will require time, standards and greater familiarity.
But the direction is already set: not a replacement, but a new tool capable of expanding the possibilities of market research.

Read also AI and market research: quality remains human.