Back to articles
Evaluation & Benchmarks

LLM Survey Answers May Look Human Without Being Psychometrically Valid

3 min read

Introduction

Using large language models as synthetic survey respondents is becoming attractive for market research, product studies, and social science experiments. Yet the usual test is often superficial: does an answer sound like something a person might say? The paper Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents argues that this is the wrong standard. The more important question is whether generated responses preserve the statistical and psychological structure of a real population.

What was tested

The researchers used a Lithuanian organizational-psychology dataset containing 263 employees. The survey had 68 items grouped into 12 subscales, covering attitudes toward organizational change, work engagement, and individual work performance. The evaluation covered 37 models from OpenAI, Anthropic, Google, and twelve open-weight model families.

Rather than relying on a single prompt, the study varied how much information was disclosed about each persona through a five-level ladder. It also tested presentation choices and reasoning effort. Additional experiments swapped demographic attributes such as gender, role, and education, included a cross-language check, and used a verbatim-recall probe to examine whether memorized material could explain the results.

The central evaluation was the Psychometric Similarity Score, or PSS. Instead of judging isolated answers, the score examines properties such as the joint distribution of responses, latent structure, reliability, mediation pathways, and demographic effects. The authors compared the models with five non-LLM statistical baselines and a held-out human-versus-human ceiling. Respondent bootstrapping and an item-permutation null for Tucker’s phi were used to quantify uncertainty and provide a stricter comparison.

Main findings

  • LLMs recover the broad direction of human psychometric relationships, but directional agreement is not the same as structural validity.
  • A Gaussian-copula baseline beats every tested LLM on the sample-driven PSS components. This suggests that a relatively simple statistical model can reproduce population structure more effectively than unconstrained persona simulation.
  • The LLM crowd is more similar to itself than to humans: mean inter-LLM PSS is 0.73. Agreement across models therefore should not automatically be treated as evidence of realism.
  • Counterfactual demographic swaps show a pronounced education-related effect, with a mean absolute effect size of 0.56. The result indicates that demographic cues can systematically change model outputs, without demonstrating that the models have recovered the real-world mechanism behind those differences.
  • The correlation between verbatim-recall performance and PSS rank is 0.00, providing no indication that memorization drove the leaderboard.

Why it matters

The study moves synthetic-respondent evaluation away from surface plausibility and toward population-level validity. A response can be fluent, coherent, and individually believable while still producing unreliable scales, distorted correlations, or an artificial latent structure. For researchers using generated survey data, matching a mean or generating convincing personas is therefore not enough.

This does not amount to a blanket claim that LLMs have no role in survey research. It is a warning about the conditions for responsible use. Synthetic respondents should be compared with statistical baselines, human-data ceilings, and counterfactual tests, rather than judged by model agreement alone. Studies should also report which psychometric properties are preserved and which are not.

The practical lesson is straightforward: if the research question concerns relationships among constructs, subgroup differences, or mediation, validation must occur at the joint-distribution and measurement levels. “Sounds human” and “measures like humans” are different claims, and this audit shows why they should not be conflated.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles