The ledger the seal is measured against.
Every accuracy claim on this site traces to an entry here: what was predicted, against which real outcome, with which method — and what it got wrong. Entries are dated and only change by measurement.
The ruler in production: 16.73 pp, not yet green.
The full questionnaire of a published Brazilian consumer and opinion survey (February 2026, N=1,557, online panel) was applied to a synthetic population calibrated on the 2022 Census. The result is the ruler the product reads today — and it sits above the threshold for a Calibrated seal.
| What was measured | |
|---|---|
| Design | Open-book, declared as such. The real results were in the repository before the run; the session prompt contains none of them, and the population was calibrated on the Census, never on the survey. Read as calibration validation, not blind prediction. |
| Population | A lot of 2,401 synthetic individuals calibrated on the 2022 Census, filtered to connected adults 18+ to match the survey's base. 200 respondents per question. |
| Instrument | 116 items in four blocks (seasonal purchase, food habits, longevity, safety), verbatim; 120 comparable after expanding multi-select grids; one item excluded for lack of a real-outcome match. |
| Aggregate error | MAE 16.73 percentage points across the 120 questions. The Calibrated seal requires under 15 — so the seal for this population stays Directional. |
| Calibration | Leave-one-block-out Platt: fit on three of the four blocks, validate on the fourth, rotated, so every question gets one out-of-fold calibrated value. |
| Band coverage (90% target, held-out) | Binary 0.909 · ordinal 0.961 · share 0.987 after calibration. Half-width of the band: 22.07 pp. |
| Where it lives | A versioned JSON artifact shipped with the engine. The Observatory reads it to grade a prediction; an engine version that no longer matches the ruler makes the grade fall back to C automatically. |
Reproducible offline from the archived run: zero model calls, one deterministic script.
Where the ruler is valid — and where it is not.
The aggregate hides the segments. Per-segment error against the same survey, with a declared reading: ok up to 1.15× the aggregate (19.24 pp), attention up to 1.35× (22.59 pp), restricted above that, and insufficient whenever the cell has fewer than 20 synthetic respondents — there, sampling noise alone is about 11 pp.
| Segment | MAE (pp) | Reading |
|---|---|---|
| Class AB · Class C · Gen X · Gen Z · Millennials · Southeast · South · Men · Women | ≤ 19.24 | ok |
| Northeast | 19.34 | attention |
| Class DE | 20.79 | attention |
| Boomers | 23.16 | restricted |
| Center-West · North | n = 16 | insufficient |
Each number is the difference between two estimates. Only the synthetic side has a known sample size per cell; the real side's cell sizes were not published, so the human percentage is treated as exact — a declared limitation, not a footnote.
Adversarial battery: published rates.
Scripted hostile interviews run as a backtest. Eight probes, one per failure family in the limitations registry, each producing a rate that is recorded in the calibration ledger with the same discipline as the ruler above.
Nonsense compliance
Meaningless input in three arms — real words with no relation, an invented word, a coherent but unrelated topic. Pass only when the person registers the incoherence in all three. Measured on the production model, n=100: 38 of 100 failed without protection; 13 of 100 with the sanity-check clause. A clause is an instruction, not a gate: it reduces, it does not eliminate.
Training-cutoff anachronism
A post-cutoff fact that was not injected must yield an honest "I don't know". After publishing the temporal-horizon rule in the prompt, n=100: decidable approval 6.7% → 87.0%, with inconclusive verdicts falling from 127 to 2. Declared cost: the reverse-pressure probe moved from 94.0% to 86.7%.
What the other six test
Induced false fact (adopting a name absent from context) · reverse pressure (position shifting under bare majority appeal) · knowledge vacuum (answering without injected context) · card consistency (re-eliciting three card dimensions mid-dialogue, tolerance ±2) · sensitive panel (first vs. third person on the same construct) · anchoring (same question, low vs. high anchor). Two product metrics carry their operational definition inside the artifact: adversarial distinguishability and acquiescence rate.
The blind block: registered, not yet run.
One block of the same survey (attitudes about aging, five key questions across six generation × class segments) was pre-registered as a blind prediction: the prediction commit came before any real number was opened, and the git history proves the order. The engine has not been run against it yet; what exists is the human baseline.
| Question | Predicted blind (human) | Real | Verdict |
|---|---|---|---|
| Already organizing for old age | AB > C > DE | AB 78 · C 64 · DE 60 | strong hit |
| Financial preparation | class gradient | AB 83 · C 67 · DE 58 | strong hit |
| Negative emotion | C/DE and women on top; 50+ calm | C 62 > AB 53 · women 61 · 50+ 54 | partial |
| Worry about aging | rises with age and constraint | falls with age (18–29: 45 · 50+: 31); AB on top | miss |
| Fear of running out of money | inverse to class (DE on top) | near-universal, ~75%; DE slightly lower | miss |
Pass criteria for the engine, fixed in advance: Spearman ρ ≥ 0.6 in at least 3 of the 5 questions, the class gradient right on the two preparation items, MAE under 15 pp in the total — and, as a discriminating bonus, getting right that worry and fear do not follow financial vulnerability, where intuition failed. Declared sample bias: online panel, 47% higher education, 41% class AB — "connected Brazil", not the census.
How to read this ledger.
A seal only turns green with a backtest against a real, reported outcome — never by prompting. Per-segment tables are published with every run, never the aggregate alone. Open-book designs are labeled open-book. And when a gate fails, prompts are not tuned until it passes: at most four documented tuning cycles, each with its hypothesis written before the rerun.
Get every new backtest by email.
When a population earns or loses a seal, or a new real outcome is measured against the engine, we write once. No newsletter cadence — only new entries in the ledger.