Bands, not points. Never green without a backtest.
A synthetic population is an instrument. Like any instrument, it is only as good as its calibration — so we measure it against reality, publish the rule, and say out loud what it still gets wrong.
Every answer tells you how much to trust it.
You can decide with this. The margin of error is small and was measured against reality.
Points the right way, with a wider margin. Good for prioritizing, not for pinning the number.
The population exists and answers, but hasn't been backtested on this topic yet. Read it as a hypothesis.
The seal only turns green with ρ ≥ 0.6 and MAE < 15pp against a real outcome. Below ρ 0.4 it turns red — and the decision goes back to a human.
Accuracy ruler: a published Brazilian consumer/opinion survey, Feb/2026 (N=1,557). For domains far from it, the ruler is declared extrapolation — not direct measurement. Synthetic simulation is not prediction.
Seven layers, not a profile.
Each synthetic person is built in seven layers, anchored in official public data — so that the same person answers the same stimulus differently at different moments, the way humans do.
Calibration sources: IBGE Census 2022 (incl. Nupcialidade e Família, Nov/2025 — SIDRA tables 9879, 9881 and 9882, household composition) · PNAD Contínua · FGV Social · TSE · public statistical series — vintages and declared biases documented in the app's methodology page.
How far a population can be resolved.
Public statistics set the ceiling. A synthetic population can only be as fine-grained, and as current, as the data that calibrates it — so the ceiling is declared before you run a study, not discovered inside the results.
Geographic floor
Populations are calibrated at country, region and state level — the levels at which the continuous series that keep them current stay representative.
What this means for youCensus tract data exists, but it ages a decade between censuses: a population aimed at one neighbourhood would be precise about 2022 and quiet about today. A study at that resolution is declared extrapolation, and the seal degrades with it.
Vintage and drift
A census is decennial; what happens in between comes from complementary surveys with their own scope and cadence.
What this means for youEvery population pins the vintage of the source that calibrated it, and a study is always scored against the ruler of its own vintage. Distance from that vintage is drift — declared, never smoothed over.
Heterogeneity above the strata
A continental country with wide social distance, layers of immigration and Indigenous peoples present in every state does not compress into a national average.
What this means for youMarginals reproduce the strata that are published. They do not, on their own, resolve a minority population sitting inside a stratum — the group is represented in the totals without being legible as itself.
When a slice is not enough
Where the published strata absorb the group you want to hear, filtering the general library is the wrong instrument.
What this means for youThat question needs a dedicated population, with its own sources and its own seal. Where we cannot build one honestly, the study goes to the refusal ledger instead of coming back with a confident number.
This is the same constraint in every country — which is why the seal is never transferred from one to another. Each one is measured where it stands.
What it still gets wrong — said out loud.
We keep a living registry of the ways a synthetic person can fail when interrogated like a real one. Six failure families, each with a detection mechanism, a built defense — and an honest status that only changes by measurement, never by optimism. The full registry — eighteen named traps, each with evidence, defense, status and date — is public at /limitations.
Pleasing the interviewer
Language models tend to agree with whoever asks — even when the evidence doesn't.
How we're beating itAn anti-acquiescence clause at the person's core, with the agreement rate measured by an adversarial battery — a defense only counts once its rate is measured.
Filling the void
Where a real person wouldn't know, the model might invent a plausible answer.
How we're beating itA per-person knowledge licence: what they plausibly know is declared, "I don't know" is a valid answer — and what falls outside the licence is blocked.
Population flattening
Variance collapse and demographic caricature — the hardest family, and the one we say openly is not solved by prompting alone.
How we're beating itFidelity measured per subgroup and style dosing against caricature; where prompting was exhausted by measurement, the work moves to who answers — model, card and population genome.
Card-vs-speech drift
The person's material facts live on a card; the speech must not contradict it.
How we're beating itSpeech is checked against the canonical card, conversation by conversation — a contradiction becomes a measured alert, never silence.
Instrument artifacts
Social desirability, question order, anchoring — the same traps that distort human surveys.
How we're beating itOption-order rotation across respondents, blind tests against real research, and the biases measured inside the instrument itself.
The engine as a system
Cross-context contamination and provider effects — failures of the machinery, not of the persona.
How we're beating itContamination monitored, vocabulary published in the prompt with preserve-and-mark (nothing silently discarded), and provider effects attributed to the provider.
No instrument is born ready — neither was the telescope. Every limitation overcome polishes the lens: more eyes, more honest, for every decision about people.
Get every new backtest by email.
When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.