Affective Research · Leaderboard

Appraisal-EI Benchmark

17-dimension appraisal structure vs human gold. Headline = tracking: per dimension, does the model's rating move with the humans' across situations (Pearson r across items, mean over dims), with a 95% CI that resamples the 23 unique scenarios. Second metric = discriminant validity (dimension-specific vs valence smear). The human baseline is leave-one-rater-out on the same 84-item holdout. Two no-model rows (a constant average profile, and that profile plus noise) sit on the board so every score reads against "no model at all". Pin tag v3.0.0, never main.
Tracking (mean per-dim Pearson r) by model