Methodology v3.0.0

How the Appraisal-EI Benchmark works. Pin a tag. Never evaluate main.

What it measures

Most emotion evaluations ask what a person feels. This benchmark asks whether a model represents why - the causal structure appraisal theory says produces the emotion. Every stimulus is scored as a 17-dimensional appraisal vector (each dimension -3.0 to +3.0), not an emotion label. A model's answer is compared to human gold ratings dimension-by-dimension.

The dimensions come from appraisal theory (Smith & Ellsworth, Scherer's CPM, OCC): pleasantness, goal relevance, goal congruence, certainty, control, responsibility, fairness, effort, expectation, novelty, urgency, intensity, coping potential, social desirability, moral worth, attribution, and anticipated emotion.

The stimuli

SetSizeFormUsed for
Human-gold vignette holdout84 rows / 23 unique scenariosTest-split vignettes, human consensus (3–4 calibrated raters)Headline rank (v3.0.0)
Unique-scenario gold94 distinct textsName-masked collapse of 560 rated rowsOne vector per scenario; not the public rank
Template vignettes560 rows, 94 textsGenerator priors; intensity never reached the textDeprecated eval set. Smoke only. See QC_FINDINGS.md
crowd-enVENT1,200Reader consensus mapped onto our 17 dimsAdjacent gold; mapping not yet validated against this rubric
Audio clips1,440 RAVDESS refsActed speech, wavs not redistributedRoadmap: no public audio score in v2

Keyword-free matters: vignettes contain no emotion words. The model must infer appraisal structure, not pattern-match a sentiment lexicon. Quote 84 rows and 23 unique scenarios. Do not quote 560.

The protocol, step by step

  1. Prompt. Each test item is sent zero-shot with the canonical rubric: the 17 dimension definitions, the scale, the scenario, and "Return JSON only, keys exactly those 17 names."
  2. Parse. The first {...} JSON object in the response is extracted. A response counts as parsed only if it yields ≥8 of the 17 dimensions.
  3. Score per dimension. For each of the 17 dimensions, the Pearson correlation between the model's ratings and the human gold across the holdout items: when the situation changes, does the model's rating move the way the humans' does? A dimension the model never varies scores 0.
  4. Aggregate. Headline appraisal tracking = mean of those 17 r values, with a 95% CI from resampling the 23 unique scenarios (the 84 rows are name-swapped copies, so rows are not independent). Second public metric = discriminant validity (mean diagonal minus mean |off-diagonal| of the multitrait matrix). Roadmap slots stay null. Nothing unimplemented is averaged in as zero.
  5. Record. Everything lands in results/<model>.json: the score, the per-item r values, the per-dim stats, the raw predictions, the gold source, the timestamp, and the status.
Same prompt, every model. The rubric is fixed by SPEC, not tuned per model. Zero-shot (no examples) keeps the test about the model's representation, not prompt engineering.

Reading the numbers

MetricPlain meaningTechnical
TrackingDoes each rating move with the humans' across situations? This is the rank. A constant answer scores 0.mean over dims of per-dim Pearson r across items; scenario-bootstrap 95% CI
Calibration (diagnostic)How closely each 17-dim profile matches gold. Not ranked: a constant average profile scores 0.864, above the human ceiling, because every situation shares a typical profile.mean per-item Pearson r (the v2 headline)
Human mimicryCan a simple detector tell the model's ratings from a real rater's? Always printed with the human band. Inside the band is necessary, not sufficient: random noise lands there too.diagonal-Gaussian discriminant AUC on consensus-centered vectors
No-model baselinesA constant average profile, and that profile plus noise at the raters' spread. Every score reads against these.scripts/make_baselines.py, train-split raters only
Discriminant validityDid it recover 17 dimensions, or smear one valence signal across all of them?mean(diag) − mean(|off-diag|) of r(pred[i], gold[j])
Human ceilingHow well a held-out rater matches the others. Models sit next to this, they do not "beat" it.leave-one-rater-out Pearson r
Direction % (diagnostic)How often the model rated a dimension in the same sign as gold. Not ranked: the constant average profile scores 93%, because most dimensions lean the same way in most situations.% of (item, dim) pairs with matching sign; exact zeros skipped

The current finding (v3.0.0): frontier models track at 0.42 to 0.59 against a human ceiling of 0.61 (after the attribution gold correction), with discriminant validity near 0.25. They recover the direction of human appraisal, but they do not yet separate unfair from uncontrollable from urgent. The v2 headline hid most of this: on calibration a constant answer looked better than every model.

Status badges - read these first

BadgeMeans
smokeScored against generator priors, not humans. Hidden from the main board unless you toggle provisional. Archived under results/archive/.
measuredScored against human gold meeting the agreement gate (quadratic-weighted κ ≥ 0.6 or Krippendorff α ≥ 0.8).

Gold provenance is declared per file in gold_status.json. As of v3.0.0 the headline set is vignettes_test_human.jsonl. crowd-enVENT remains adjacent gold whose 17-dim mapping is not yet validated against this rubric. Per-rater vectors live in annotation/rater_vectors.jsonl.

The leaderboard pipeline

Adding a model

# 1. registry entry in models.json (name, org, origin, weights)
# 2. run against the frozen holdout (this is the default):
python scoring/run_eval.py --provider together --model org/model-id --name "Display Name"
# 3. rebuild:
python site/build_site.py

See CONTRIBUTING.md. Smoke results and custom prompts are rejected.

Auditing a score

Every leaderboard number is reproducible from the repo alone:

Honest caveat: a model's vignette text is public, so frontier training corpora may contain it. This is the standard public-benchmark contamination tradeoff - mitigated by keyword-free stimuli, private gold labels (planned), and versioned releases. Report anything that looks gamed.

Versioning

VERSION + git tags pin the artifact; the leaderboard and all results record which version scored them. Stimuli are append-only within a major version - corrections bump minor (1.0→1.1), schema changes bump major (2.0). Full rules and the sub-score definitions are in SPEC.md - the normative document.