What it still gets wrong — with a status.
A living registry of the ways a synthetic person can fail when interrogated like a real one. Six failure families, eighteen named traps, each with its external evidence, what we do about it, and a status that only changes by measurement — never by optimism. Registry v1 dated 2026-08-20; last status change 2026-09-02.
How to read the status.
covered a defense is live and its rate is measured or the incident class is closed by mechanism · partial a defense is live but the rate is not yet measured, or the root cause is only mitigated · open measured or evidenced, not yet mitigated. Where a trap is "our measurement" or "our incident", the number lives in the backtest ledger or in the engine's own records — this page does not restate what is not published.
Pleasing the interviewer.
Language models tend to agree with whoever asks — even when the evidence doesn't. How we're beating it: an anti-acquiescence clause at the person's core, with the agreement rate measured by the adversarial battery — a defense only counts once its rate is measured.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 1 · Acquiescence / sycophancy | Sharma et al., ICLR 2024 (rooted in preference training); Kantar 2023 (synthetic panels systematically more positive than 5,000 humans) | Anti-acquiescence clause and knowledge licence live in the dialogue prompt. The clause alone is insufficient against a trained-in bias, so the rate is what counts: reverse-pressure and induced-false-fact probes in the battery. | partial |
| 18 · Nonsense compliance | Our contribution, motivated by public demos of chatbots agreeing with word salad | Sanity-check clause in the dialogue prompt. Measured on the production model, n=100: 38 of 100 failed without it, 13 with it — see the ledger. A clause reduces; it does not eliminate. | covered |
Filling the void.
Where a real person wouldn't know, the model might invent a plausible answer. How we're beating it: a per-person knowledge licence: what they plausibly know is declared, "I don't know" is a valid answer — and what falls outside the licence is blocked.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 2 · Knowledge vacuum → confabulation by suggestion | Our contribution | The instrument itself (question, options, answer) is always injected; the study's evidence digest travels with it; the knowledge licence names the boundary. The vacuum-probe rate is measured by the battery. | covered |
| 3 · Training-cutoff anachronism | Underwood et al., 2025 (models import contemporary assumptions even when fine-tuned on period prose) | Post-cutoff facts enter only by injection; the temporal-horizon rule is published in the prompt. Measured, n=100: decidable approval 6.7% → 87.0%, inconclusive verdicts 127 → 2 — see the ledger. | covered |
Population flattening.
Variance collapse and demographic caricature — the hardest family, and the one we say openly is not solved by prompting alone. How we're beating it: fidelity measured per subgroup and style dosing against caricature; where prompting was exhausted by measurement, the work moves to who answers — model, card and population genome.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 4 · Extremity / variance collapse | Wu 2025 (variance collapse in digital twins); Bisbee et al., Political Analysis 2024 (mean agreement hides intra-group variance loss); Maier 2025 | Post-calibration on the reading side (Platt, leave-one-block-out) is live. On the origin side, alternative elicitation is implemented but not yet measured against the ruler. | partial |
| 5 · Demographic caricature / identity flattening | Wang et al., Nature Machine Intelligence (misportray and flatten, worst for minoritized groups); Xiao 2026 ("Chameleon's limit": higher individual fidelity can mean more stereotyped populations) | Per-subgroup error tables in every backtest; a style-dosing rule fixed opener repetition. The error gradient by generation, region and class survives prompt-level fixes — so the root work is on who answers: model, card and population genome. | open |
| 12 · Subgroup fidelity heterogeneity | Ma et al., ACL 2025 (fidelity uneven across groups); Choi 2026 (structural, marginal and individual fidelity are separate axes) | Per-segment tables published with every run, never the aggregate alone — the ledger declares where the ruler is ok, under attention, restricted or insufficient. | partial |
| 13 · Unreal non-response | Our measurement: the neutral option is almost never chosen by synthetic respondents where real respondents choose it | A generic anti-acquiescence instruction was tested and rejected — it worsens material facts. Needs a targeted design. | open |
Card-vs-speech drift.
The person's material facts live on a card; the speech must not contradict it. How we're beating it: speech is checked against the canonical card, conversation by conversation — a contradiction becomes a measured alert, never silence.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 6 · Material fact vs. declared disposition | Our measurement: facts derivable from the card land close to the real share; facts outside it drift far | Enrich the card with material-life axes (possessions, routine) so facts are carried, not guessed. | open |
| 16 · Card-consistency drift | Amirova 2024 (silicon interviews structurally uniform — "confident, optimistic" — regardless of persona) | Battery probe re-asks three card dimensions mid-dialogue and compares against the card, tolerance ±2. Measured first; mitigated after. | open |
Instrument artifacts.
Social desirability, question order, anchoring — the same traps that distort human surveys. How we're beating it: option-order rotation across respondents, blind tests against real research, and the biases measured inside the instrument itself.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 7 · Social desirability bias | Argyle 2023 (divergence on sensitive topics); Chapala 2026 (third-person reframing best of four strategies) | Battery sensitive panel records the first-person vs. third-person divergence on the same construct. Candidate mitigation: third-person elicitation on sensitive items. | open |
| 9 · Order / anchoring effects | Huang 2026 (anchoring at shallow processing; not eliminable by convention) | Latin-square rotation of scales is live for instruments. Anchoring in free dialogue is measured by the battery before it is mitigated. | partial |
| 11 · Intervention drift (pseudo-experiments) | "Illusion of Intervention", arXiv 2026 (a treatment can shift latent persona attributes, so A/B on synthetics may compare different latent populations) | Negative-control outcomes are required next to every treatment effect in comparative studies. | partial |
| 17 · Option-referent ambiguity | Our incident: a generic polar label read one way by the rubric and the opposite way by the user | A deterministic label-vocabulary gate; a never-guess dialog when the binding is ambiguous; the binding published verbatim in both prompts and declared in the result. | covered |
The engine as a system.
Cross-context contamination and provider effects — failures of the machinery, not of the persona. How we're beating it: contamination monitored, vocabulary published in the prompt with preserve-and-mark (nothing silently discarded), and provider effects attributed to the provider.
| Trap | External evidence | What we do | Status |
|---|---|---|---|
| 8 · WEIRD / cultural pull | Tao et al., PNAS Nexus 2024 (models pull toward Anglo-Protestant self-expression values) | Structural antidote by design: census-calibrated Brazilian population, Portuguese prompts, Brazilian instruments. Calibration is the mitigation; per-subgroup measurement keeps it honest. | partial |
| 10 · Cross-context contamination | Our incident: a stimulus from one screen leaked into a dialogue on another | Full snapshot isolation between contexts; the dispatched payload is kept durable for forensics. | covered |
| 14 · Closed-enum silent discard | Our incidents: a required vocabulary not published in the prompt made per-turn losses invisible for months | Vocabulary published verbatim in the prompt; unknown values preserved with a mark instead of discarded; a generic guard fails when a new model call site is added unclassified. | covered |
| 15 · Provider / quantization effects | Our measurements: the same arm can differ across providers, while the calibration parameters stay stable | A regression sensor on the calibration parameter band; the provider pin is chosen by a smoke test before any flip. | covered |
What moves a status.
A new trap class observed in production enters the registry in the same arc. Any engine, provider or prompt-block change runs the battery before it goes live. The registry is revisited at every backtest campaign — and a status is updated by measurement, never by optimism. Ceilings that are not failures — geographic floor, vintage and drift, heterogeneity above the stratum — live on the methodology page.
Get every new backtest by email.
When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.