What it measures
Most emotion evaluations ask what a person feels. This benchmark asks whether a model represents why - the causal structure appraisal theory says produces the emotion. Every stimulus is scored as a 17-dimensional appraisal vector (each dimension -3.0 to +3.0), not an emotion label. A model's answer is compared to human gold ratings dimension-by-dimension.
The dimensions come from appraisal theory (Smith & Ellsworth, Scherer's CPM, OCC): pleasantness, goal relevance, goal congruence, certainty, control, responsibility, fairness, effort, expectation, novelty, urgency, intensity, coping potential, social desirability, moral worth, attribution, and anticipated emotion.
The stimuli
| Set | Size | Form | Used for |
|---|---|---|---|
| Human-gold vignette holdout | 84 rows / 23 unique scenarios | Test-split vignettes, human consensus (3–4 calibrated raters) | Headline rank (v3.0.0) |
| Unique-scenario gold | 94 distinct texts | Name-masked collapse of 560 rated rows | One vector per scenario; not the public rank |
| Template vignettes | 560 rows, 94 texts | Generator priors; intensity never reached the text | Deprecated eval set. Smoke only. See QC_FINDINGS.md |
| crowd-enVENT | 1,200 | Reader consensus mapped onto our 17 dims | Adjacent gold; mapping not yet validated against this rubric |
| Audio clips | 1,440 RAVDESS refs | Acted speech, wavs not redistributed | Roadmap: no public audio score in v2 |
Keyword-free matters: vignettes contain no emotion words. The model must infer appraisal structure, not pattern-match a sentiment lexicon. Quote 84 rows and 23 unique scenarios. Do not quote 560.
The protocol, step by step
- Prompt. Each test item is sent zero-shot with the canonical rubric: the 17 dimension definitions, the scale, the scenario, and "Return JSON only, keys exactly those 17 names."
- Parse. The first
{...}JSON object in the response is extracted. A response counts as parsed only if it yields ≥8 of the 17 dimensions. - Score per dimension. For each of the 17 dimensions, the Pearson correlation between the model's ratings and the human gold across the holdout items: when the situation changes, does the model's rating move the way the humans' does? A dimension the model never varies scores 0.
- Aggregate. Headline appraisal tracking = mean of those 17 r values, with a 95% CI from resampling the 23 unique scenarios (the 84 rows are name-swapped copies, so rows are not independent). Second public metric = discriminant validity (mean diagonal minus mean |off-diagonal| of the multitrait matrix). Roadmap slots stay null. Nothing unimplemented is averaged in as zero.
- Record. Everything lands in
results/<model>.json: the score, the per-item r values, the per-dim stats, the raw predictions, the gold source, the timestamp, and the status.
Reading the numbers
| Metric | Plain meaning | Technical |
|---|---|---|
| Tracking | Does each rating move with the humans' across situations? This is the rank. A constant answer scores 0. | mean over dims of per-dim Pearson r across items; scenario-bootstrap 95% CI |
| Calibration (diagnostic) | How closely each 17-dim profile matches gold. Not ranked: a constant average profile scores 0.864, above the human ceiling, because every situation shares a typical profile. | mean per-item Pearson r (the v2 headline) |
| Human mimicry | Can a simple detector tell the model's ratings from a real rater's? Always printed with the human band. Inside the band is necessary, not sufficient: random noise lands there too. | diagonal-Gaussian discriminant AUC on consensus-centered vectors |
| No-model baselines | A constant average profile, and that profile plus noise at the raters' spread. Every score reads against these. | scripts/make_baselines.py, train-split raters only |
| Discriminant validity | Did it recover 17 dimensions, or smear one valence signal across all of them? | mean(diag) − mean(|off-diag|) of r(pred[i], gold[j]) |
| Human ceiling | How well a held-out rater matches the others. Models sit next to this, they do not "beat" it. | leave-one-rater-out Pearson r |
| Direction % (diagnostic) | How often the model rated a dimension in the same sign as gold. Not ranked: the constant average profile scores 93%, because most dimensions lean the same way in most situations. | % of (item, dim) pairs with matching sign; exact zeros skipped |
The current finding (v3.0.0): frontier models track at 0.42 to 0.59 against a human ceiling of 0.61 (after the attribution gold correction), with discriminant validity near 0.25. They recover the direction of human appraisal, but they do not yet separate unfair from uncontrollable from urgent. The v2 headline hid most of this: on calibration a constant answer looked better than every model.
Status badges - read these first
| Badge | Means |
|---|---|
| smoke | Scored against generator priors, not humans. Hidden from the main board unless you toggle provisional. Archived under results/archive/. |
| measured | Scored against human gold meeting the agreement gate (quadratic-weighted κ ≥ 0.6 or Krippendorff α ≥ 0.8). |
Gold provenance is declared per file in gold_status.json. As of v3.0.0 the headline set is
vignettes_test_human.jsonl. crowd-enVENT remains adjacent gold whose 17-dim mapping is not yet
validated against this rubric. Per-rater vectors live in annotation/rater_vectors.jsonl.
The leaderboard pipeline
scoring/run_eval.py- the runner. One command per model; supports OpenAI, Together, Fireworks (any OpenAI-compatible endpoint). Writes a result JSON.models.json- the registry the filters read: display name, org, origin (ISO country of the publishing org), open/closed weights, family, params.site/build_site.py- merges results + registry intodocs/data.json, the single payload this site fetches.docs/index.html- the leaderboard: sortable columns, origin/weights/org chips, checkbox compare, per-item sparklines, the bar chart and 17-axis radar.
Adding a model
# 1. registry entry in models.json (name, org, origin, weights)
# 2. run against the frozen holdout (this is the default):
python scoring/run_eval.py --provider together --model org/model-id --name "Display Name"
# 3. rebuild:
python site/build_site.py
See CONTRIBUTING.md. Smoke results and custom prompts are rejected.
Auditing a score
Every leaderboard number is reproducible from the repo alone:
results/<model>.json→per_itemhas every item's r, prediction, and gold vectorper_dimhas per-dimension r + mean absolute errorscoring/score.pyis the reference implementation of every metric (numpy+scipy only)- The prompt text is generated deterministically from
schema/rating_schema.json
Versioning
VERSION + git tags pin the artifact; the leaderboard and all results record which version scored
them. Stimuli are append-only within a major version - corrections bump minor (1.0→1.1), schema changes bump
major (2.0). Full rules and the sub-score definitions are in
SPEC.md - the normative document.