Known limitations

What it still gets wrong — with a status.

A living registry of the ways a synthetic person can fail when interrogated like a real one. Six failure families, eighteen named traps, each with its external evidence, what we do about it, and a status that only changes by measurement — never by optimism. Registry v1 dated 2026-08-20; last status change 2026-09-02.

How to read the status.

covered a defense is live and its rate is measured or the incident class is closed by mechanism · partial a defense is live but the rate is not yet measured, or the root cause is only mitigated · open measured or evidenced, not yet mitigated. Where a trap is "our measurement" or "our incident", the number lives in the backtest ledger or in the engine's own records — this page does not restate what is not published.

Family A

Pleasing the interviewer.

Language models tend to agree with whoever asks — even when the evidence doesn't. How we're beating it: an anti-acquiescence clause at the person's core, with the agreement rate measured by the adversarial battery — a defense only counts once its rate is measured.

TrapExternal evidenceWhat we doStatus
1 · Acquiescence / sycophancySharma et al., ICLR 2024 (rooted in preference training); Kantar 2023 (synthetic panels systematically more positive than 5,000 humans)Anti-acquiescence clause and knowledge licence live in the dialogue prompt. The clause alone is insufficient against a trained-in bias, so the rate is what counts: reverse-pressure and induced-false-fact probes in the battery.partial
18 · Nonsense complianceOur contribution, motivated by public demos of chatbots agreeing with word saladSanity-check clause in the dialogue prompt. Measured on the production model, n=100: 38 of 100 failed without it, 13 with it — see the ledger. A clause reduces; it does not eliminate.covered
Family B

Filling the void.

Where a real person wouldn't know, the model might invent a plausible answer. How we're beating it: a per-person knowledge licence: what they plausibly know is declared, "I don't know" is a valid answer — and what falls outside the licence is blocked.

TrapExternal evidenceWhat we doStatus
2 · Knowledge vacuum → confabulation by suggestionOur contributionThe instrument itself (question, options, answer) is always injected; the study's evidence digest travels with it; the knowledge licence names the boundary. The vacuum-probe rate is measured by the battery.covered
3 · Training-cutoff anachronismUnderwood et al., 2025 (models import contemporary assumptions even when fine-tuned on period prose)Post-cutoff facts enter only by injection; the temporal-horizon rule is published in the prompt. Measured, n=100: decidable approval 6.7% → 87.0%, inconclusive verdicts 127 → 2 — see the ledger.covered
Family C

Population flattening.

Variance collapse and demographic caricature — the hardest family, and the one we say openly is not solved by prompting alone. How we're beating it: fidelity measured per subgroup and style dosing against caricature; where prompting was exhausted by measurement, the work moves to who answers — model, card and population genome.

TrapExternal evidenceWhat we doStatus
4 · Extremity / variance collapseWu 2025 (variance collapse in digital twins); Bisbee et al., Political Analysis 2024 (mean agreement hides intra-group variance loss); Maier 2025Post-calibration on the reading side (Platt, leave-one-block-out) is live. On the origin side, alternative elicitation is implemented but not yet measured against the ruler.partial
5 · Demographic caricature / identity flatteningWang et al., Nature Machine Intelligence (misportray and flatten, worst for minoritized groups); Xiao 2026 ("Chameleon's limit": higher individual fidelity can mean more stereotyped populations)Per-subgroup error tables in every backtest; a style-dosing rule fixed opener repetition. The error gradient by generation, region and class survives prompt-level fixes — so the root work is on who answers: model, card and population genome.open
12 · Subgroup fidelity heterogeneityMa et al., ACL 2025 (fidelity uneven across groups); Choi 2026 (structural, marginal and individual fidelity are separate axes)Per-segment tables published with every run, never the aggregate alone — the ledger declares where the ruler is ok, under attention, restricted or insufficient.partial
13 · Unreal non-responseOur measurement: the neutral option is almost never chosen by synthetic respondents where real respondents choose itA generic anti-acquiescence instruction was tested and rejected — it worsens material facts. Needs a targeted design.open
Family D

Card-vs-speech drift.

The person's material facts live on a card; the speech must not contradict it. How we're beating it: speech is checked against the canonical card, conversation by conversation — a contradiction becomes a measured alert, never silence.

TrapExternal evidenceWhat we doStatus
6 · Material fact vs. declared dispositionOur measurement: facts derivable from the card land close to the real share; facts outside it drift farEnrich the card with material-life axes (possessions, routine) so facts are carried, not guessed.open
16 · Card-consistency driftAmirova 2024 (silicon interviews structurally uniform — "confident, optimistic" — regardless of persona)Battery probe re-asks three card dimensions mid-dialogue and compares against the card, tolerance ±2. Measured first; mitigated after.open
Family E

Instrument artifacts.

Social desirability, question order, anchoring — the same traps that distort human surveys. How we're beating it: option-order rotation across respondents, blind tests against real research, and the biases measured inside the instrument itself.

TrapExternal evidenceWhat we doStatus
7 · Social desirability biasArgyle 2023 (divergence on sensitive topics); Chapala 2026 (third-person reframing best of four strategies)Battery sensitive panel records the first-person vs. third-person divergence on the same construct. Candidate mitigation: third-person elicitation on sensitive items.open
9 · Order / anchoring effectsHuang 2026 (anchoring at shallow processing; not eliminable by convention)Latin-square rotation of scales is live for instruments. Anchoring in free dialogue is measured by the battery before it is mitigated.partial
11 · Intervention drift (pseudo-experiments)"Illusion of Intervention", arXiv 2026 (a treatment can shift latent persona attributes, so A/B on synthetics may compare different latent populations)Negative-control outcomes are required next to every treatment effect in comparative studies.partial
17 · Option-referent ambiguityOur incident: a generic polar label read one way by the rubric and the opposite way by the userA deterministic label-vocabulary gate; a never-guess dialog when the binding is ambiguous; the binding published verbatim in both prompts and declared in the result.covered
Family F

The engine as a system.

Cross-context contamination and provider effects — failures of the machinery, not of the persona. How we're beating it: contamination monitored, vocabulary published in the prompt with preserve-and-mark (nothing silently discarded), and provider effects attributed to the provider.

TrapExternal evidenceWhat we doStatus
8 · WEIRD / cultural pullTao et al., PNAS Nexus 2024 (models pull toward Anglo-Protestant self-expression values)Structural antidote by design: census-calibrated Brazilian population, Portuguese prompts, Brazilian instruments. Calibration is the mitigation; per-subgroup measurement keeps it honest.partial
10 · Cross-context contaminationOur incident: a stimulus from one screen leaked into a dialogue on anotherFull snapshot isolation between contexts; the dispatched payload is kept durable for forensics.covered
14 · Closed-enum silent discardOur incidents: a required vocabulary not published in the prompt made per-turn losses invisible for monthsVocabulary published verbatim in the prompt; unknown values preserved with a mark instead of discarded; a generic guard fails when a new model call site is added unclassified.covered
15 · Provider / quantization effectsOur measurements: the same arm can differ across providers, while the calibration parameters stay stableA regression sensor on the calibration parameter band; the provider pin is chosen by a smoke test before any flip.covered

What moves a status.

A new trap class observed in production enters the registry in the same arc. Any engine, provider or prompt-block change runs the battery before it goes live. The registry is revisited at every backtest campaign — and a status is updated by measurement, never by optimism. Ceilings that are not failures — geographic floor, vintage and drift, heterogeneity above the stratum — live on the methodology page.

Backtest updates

Get every new backtest by email.

When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.

We use your details only for this — never sold or shared — and handle them under Brazil's LGPD. Ask us to delete them anytime.

Every limitation overcome polishes the lens.

app.syntheticperson.ai