The Handoff Problem: Where AI Stops and Human Research Starts in Retail
The research team of a Fortune 500 skincare brand tested the same question two ways. First, a panel of synthetic respondents answered what could be improved about facial masks. The list was detailed, including new ingredients, packaging changes, and alternative application methods. Then the team asked real consumers, and most of them had no comments to add; 90 percent saying nothing. What's important to note is that neither answer is wrong. The synthetic panel shows what an attentive, articulate consumer might say. The real panel shows what consumers actually say. Both insights matter. Thus, the question facing retail teams isn’t whether to use artificial intelligence or human research; it’s where one stops being the right answer and the other starts.
Synthetic data is no longer experimentation. A Qualtrics survey of more than 3,000 market researchers found that 69 percent had used synthetic responses in the past year. Today, it's used in two primary domains.
First, it's used before fieldwork, where synthetic respondents, AI-generated simulations trained on real behavior and conversations, pressure-test how a category is framed, surface hypotheses, and check concept directions. This doesn't replace discovery with actual people, but enables faster, cheaper iteration.
The second use case is across large-scale evaluation, where synthetic responses are used for ranking, screening and sorting against set criteria, enabling users to test hundreds of concepts in days and rank claims across segments in hours, not weeks. The compression is real, as pioneer brands now use synthetic consumer frameworks to shorten research cycles from weeks to under 24 hours, enabling rapid "what-if" testing of pricing, packaging or messaging.
For both, there’s a current trend in which personas and segmentation frameworks are continuously refined against actual data or augmented with behavioral data. What links these uses is that the consumer base is known, the work is upstream of a final decision, and the next stage will catch what this stage misses.
Synthetic data falls short in unstructured situations. Some categories depend on lived experience, in-the-moment behavior, and sensory interaction. How shoppers move through a store. Why they pick one product over another. The elements that trigger an unplanned purchase at checkout. This tracks with what recent academic research and practitioners report: synthetic respondents reproduce population averages reasonably well, but show less variance than real surveys and break down on the regression coefficients and intervention effects that matter for downstream decisions. Other recent academic work finds that language models can identify sensory stimuli accurately but fail to replicate the associative meanings humans attach to them, exactly the territory retail packaging, shelf, and store environment research lives in.
Now comes the question of how to choose. Most consumer research teams use two questions to set the method. Is the response space structured, with defined options and bounded comparison, or is it experiential, lived, contextual? Do you have grounded behavioral or historical data on this consumer in this context, or are you in new territory? If the answer is structured and grounded, synthetic data wins. If it’s experiential and novel, then real data wins.
One method is rarely the whole answer. Synthetic-led doesn't mean synthetic-only. Most founders building these tools draw the same line: while startups such as Yabble, Aaru, Keplar, ElectricTwin and Experial.ai report 90 percent-plus similarity to human research results on structured tasks, they still recommend final human validation after sometimes up to 10-plus synthetic studies. Where the handoff sits depends on what the decision actually commits you to. The earlier you are in the research funnel, the more likely the answers are to get reviewed at the next stage, so synthetic can lead. Once the decision is the final output and there is no next stage to catch possible hallucination, humans have to come back in.
Ultimately, the insight function's primary job is no longer to run studies but to understand handoffs. The more the technology matures, the competitive edge won't be coming from using synthetic, but from knowing when not to. Teams that win the next cycle will not be the ones with the most AI in their stack. They will be the ones that know which questions synthetic should not answer.
Clemens Pfefferkorn is a venture associate at Silicon Foundry, Kearney’s corporate venturing unit, where he advises Fortune 500 retail and consumer-goods clients on innovation strategy, venture engagement, and ecosystem development.
Related story: Retail's AI Isn't Failing Because it's Too Slow. It's Failing Because it's Not Listening
- Categories:
- Marketing
Clemens Pfefferkorn is a venture associate at Silicon Foundry, Kearney’s corporate venturing unit, where he advises Fortune 500 retail and consumer-goods clients on innovation strategy, venture engagement, and ecosystem development. His current work focuses on the AI-driven consumer research landscape. Previously at Kearney’s IMP³ROVE Innovation Competence Center, he led R&D portfolio optimization projects and supported the design of corporate-startup innovation structures.





