The words we use — and what the market calls them.
Every term below is defined the way it is used on this site, condensed from the page where it lives. The last section maps our vocabulary to the names the market and the literature use for the same thing.
What the instrument is made of.
The unit, the layers, the population and the ruler — in the order you meet them.
Synthetic person
A cognitive architecture calibrated against population strata and the behavioral literature — never an identifiable individual. The unit of analysis is the cohort, not the person.
Synthetic population
A set of synthetic people whose marginals reproduce official public statistics — for Brazil, the 2022 Census, PNAD Contínua, FGV Social classes and TSE municipal context. Each population carries its own seal.
Seven layers
Every synthetic person is built in seven layers — identity, family, economic, cultural, cognitive, emotional, social — so that the same person answers the same stimulus differently at different moments, as humans do.
Calibration
Anchoring a population's marginals in official public data, and pinning the vintage of the source that did it. Distance from that vintage is drift — declared, never smoothed.
Accuracy ruler
The real, reported outcome a population is measured against. Today: a published Brazilian consumer and opinion survey, February 2026, N=1,557. For domains far from it, the ruler is declared extrapolation, not direct measurement.
What every answer carries.
No bare numbers. Each estimate ships with its band, its seal and its provenance.
Uncertainty band
The interval around an estimate, not a single point. Every behavioral estimate comes with its band and its calibration sources.
Confidence seal
Three states. Calibrated: you can decide with this; the margin is small and was measured against reality. Directional: points the right way, with a wider margin — good for prioritizing, not for pinning the number. Seed: the population exists and answers, but hasn't been backtested on this topic — read it as a hypothesis. A seal only turns green with ρ ≥ 0.6 and MAE < 15 pp against a real outcome; below ρ 0.4 it turns red and the decision goes back to a human.
Provenance card
The list of calibration sources behind an output — what the estimate was anchored to, so a number is never separated from where it came from.
Backtest
A study run against an outcome that already happened and was reported, with the error measured per question and per segment. The only thing that moves a seal. Open-book designs are labeled open-book; blind designs are pre-registered before any real number is opened.
Segment fidelity
Where a ruler is valid and where it is not: per-segment error with a declared reading — ok, attention, restricted, or insufficient when the cell is too small to judge.
Formats and instruments.
Six ways to listen, and three instruments for the question "what will happen?".
1:1 interview · Focus group · Quantitative survey · Concept & creative test · Open question · Conversations
The six research formats, all against one calibrated population: in-depth conversation with any person; a moderated discussion where participants influence each other; a battery of scale questions answered by thousands (an existing instrument can be imported and run blind); the reaction to a product, claim, pack or piece before media spend; a doubt the AI assembles into a study you approve; and free-form chat with a segment.
Field
Announce a stimulus and watch seed individuals form and change positions over rounds of influence — recorded on a scrubbable timeline, with a 1:1 chat to ask anyone why they moved.
Observatory
A top-down scenario becomes a prediction with a margin, an honest confidence grade and a navigable three-level audit trail. Attached evidence moves the prediction, never the confidence grade — which comes from the measured ruler.
Human Portrait
The emotional read of one respondent: 13 behavioral dimensions, a per-sentence replay, and what was left unsaid.
What we keep in the open.
Two public registers and one battery.
Refusal ledger
The public list of what we refused to build: predictive policing, partisan electoral targeting, pricing that violates LGPD, modeling workers without consent, addiction optimization. "No" is part of the product.
Limitations registry
A living registry of the ways a synthetic person can fail when interrogated like a real one — six failure families, eighteen named traps, each with external evidence, a defense and a status that only changes by measurement.
Adversarial battery
Scripted hostile interviews run as a backtest — eight probes, one per failure family — producing published rates recorded in the calibration ledger.
What the market calls this.
Same object, different names. We say "synthetic person" because the unit is a calibrated cohort member, not a persona sketch and not a user proxy.
Synthetic respondents
The market-research industry's term (NIQ and others) for artificial personas that stand in for survey respondents. Closest to what a synthetic population does in a quantitative survey here.
Synthetic personas
In the industry, often a single representative consumer profile generated by AI. Here, the population is the object and the person is one calibrated member of it — which is why we avoid the word on its own.
Synthetic users
The product and UX term, and a competitor's brand name. Used for interviews, surveys and usability studies with LLM-based participants.
Silicon sampling
The academic name for generating survey-style responses from a language model conditioned on profiles that represent people or segments (Sarstedt et al., Psychology & Marketing, 2024). Every commercial variant above rests on it.
Synthetic simulation is not prediction
Our own rule. Simulation comes before field research, not instead of it: it ranks hypotheses and sends the survivors to the field, where the confirming measurement happens.
Get every new backtest by email.
When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.