Per-dimension Pearson r between model and gold ratings across the test split. Green = tracks the dimension, red = anti-correlated, "–" = not measured (too few paired observations or gold lacks the dimension).
v3.0.0 public metrics are tracking and discriminant validity only. Calibration (the v2 headline) is shown as a diagnostic: a constant average profile scores 0.864 on it, above the human ceiling. Roadmap slots (value–action, persistence, acoustic risk, steering) are named and null: they are not zeros and they are not averaged into a headline. Human mimicry is measured (reported, not averaged): can a discriminator tell the model's vectors from rater vectors? 1.0 = indistinguishable from the rater pool. The human reference band, computed from the rater vectors (…), marks what an actual human rater scores: every rater has a fingerprint the discriminator learns. The band is printed next to every score. Inside the band is necessary, not sufficient: the no-model "average + noise" row lands inside it too. Discriminant validity near 0.25 means models recover "this is bad" better than they separate unfair vs uncontrollable vs urgent.