Backtests

The ledger the seal is measured against.

Every accuracy claim on this site traces to an entry here: what was predicted, against which real outcome, with which method — and what it got wrong. Entries are dated and only change by measurement.

Entry · 2026-08-19

The ruler in production: 16.73 pp, not yet green.

The full questionnaire of a published Brazilian consumer and opinion survey (February 2026, N=1,557, online panel) was applied to a synthetic population calibrated on the 2022 Census. The result is the ruler the product reads today — and it sits above the threshold for a Calibrated seal.

What was measured
DesignOpen-book, declared as such. The real results were in the repository before the run; the session prompt contains none of them, and the population was calibrated on the Census, never on the survey. Read as calibration validation, not blind prediction.
PopulationA lot of 2,401 synthetic individuals calibrated on the 2022 Census, filtered to connected adults 18+ to match the survey's base. 200 respondents per question.
Instrument116 items in four blocks (seasonal purchase, food habits, longevity, safety), verbatim; 120 comparable after expanding multi-select grids; one item excluded for lack of a real-outcome match.
Aggregate errorMAE 16.73 percentage points across the 120 questions. The Calibrated seal requires under 15 — so the seal for this population stays Directional.
CalibrationLeave-one-block-out Platt: fit on three of the four blocks, validate on the fourth, rotated, so every question gets one out-of-fold calibrated value.
Band coverage (90% target, held-out)Binary 0.909 · ordinal 0.961 · share 0.987 after calibration. Half-width of the band: 22.07 pp.
Where it livesA versioned JSON artifact shipped with the engine. The Observatory reads it to grade a prediction; an engine version that no longer matches the ruler makes the grade fall back to C automatically.

Reproducible offline from the archived run: zero model calls, one deterministic script.

Entry · 2026-08-22

Where the ruler is valid — and where it is not.

The aggregate hides the segments. Per-segment error against the same survey, with a declared reading: ok up to 1.15× the aggregate (19.24 pp), attention up to 1.35× (22.59 pp), restricted above that, and insufficient whenever the cell has fewer than 20 synthetic respondents — there, sampling noise alone is about 11 pp.

SegmentMAE (pp)Reading
Class AB · Class C · Gen X · Gen Z · Millennials · Southeast · South · Men · Women≤ 19.24ok
Northeast19.34attention
Class DE20.79attention
Boomers23.16restricted
Center-West · Northn = 16insufficient

Each number is the difference between two estimates. Only the synthetic side has a known sample size per cell; the real side's cell sizes were not published, so the human percentage is treated as exact — a declared limitation, not a footnote.

Entry · 2026-08-20 → 2026-09-02

Adversarial battery: published rates.

Scripted hostile interviews run as a backtest. Eight probes, one per failure family in the limitations registry, each producing a rate that is recorded in the calibration ledger with the same discipline as the ruler above.

TRAP 18

Nonsense compliance

Meaningless input in three arms — real words with no relation, an invented word, a coherent but unrelated topic. Pass only when the person registers the incoherence in all three. Measured on the production model, n=100: 38 of 100 failed without protection; 13 of 100 with the sanity-check clause. A clause is an instruction, not a gate: it reduces, it does not eliminate.

TRAP 3

Training-cutoff anachronism

A post-cutoff fact that was not injected must yield an honest "I don't know". After publishing the temporal-horizon rule in the prompt, n=100: decidable approval 6.7% → 87.0%, with inconclusive verdicts falling from 127 to 2. Declared cost: the reverse-pressure probe moved from 94.0% to 86.7%.

PROBES

What the other six test

Induced false fact (adopting a name absent from context) · reverse pressure (position shifting under bare majority appeal) · knowledge vacuum (answering without injected context) · card consistency (re-eliciting three card dimensions mid-dialogue, tolerance ±2) · sensitive panel (first vs. third person on the same construct) · anchoring (same question, low vs. high anchor). Two product metrics carry their operational definition inside the artifact: adversarial distinguishability and acquiescence rate.

Entry · May 2026 · pre-registered

The blind block: registered, not yet run.

One block of the same survey (attitudes about aging, five key questions across six generation × class segments) was pre-registered as a blind prediction: the prediction commit came before any real number was opened, and the git history proves the order. The engine has not been run against it yet; what exists is the human baseline.

QuestionPredicted blind (human)RealVerdict
Already organizing for old ageAB > C > DEAB 78 · C 64 · DE 60strong hit
Financial preparationclass gradientAB 83 · C 67 · DE 58strong hit
Negative emotionC/DE and women on top; 50+ calmC 62 > AB 53 · women 61 · 50+ 54partial
Worry about agingrises with age and constraintfalls with age (18–29: 45 · 50+: 31); AB on topmiss
Fear of running out of moneyinverse to class (DE on top)near-universal, ~75%; DE slightly lowermiss

Pass criteria for the engine, fixed in advance: Spearman ρ ≥ 0.6 in at least 3 of the 5 questions, the class gradient right on the two preparation items, MAE under 15 pp in the total — and, as a discriminating bonus, getting right that worry and fear do not follow financial vulnerability, where intuition failed. Declared sample bias: online panel, 47% higher education, 41% class AB — "connected Brazil", not the census.

How to read this ledger.

A seal only turns green with a backtest against a real, reported outcome — never by prompting. Per-segment tables are published with every run, never the aggregate alone. Open-book designs are labeled open-book. And when a gate fails, prompts are not tuned until it passes: at most four documented tuning cycles, each with its hypothesis written before the rerun.

Backtest updates

Get every new backtest by email.

When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.

We use your details only for this — never sold or shared — and handle them under Brazil's LGPD. Ask us to delete them anytime.

Report your outcomes, earn a backtest.

app.syntheticperson.ai