Methodology

Gold Medal Page

Olympic-style medal tally across the v5 dataset cards. Each (dataset, target) is one competition. Only datasets with at least 5 criterion-model families competing count toward the medal table — uncontested entries (PsyProxy alone) are excluded. PsyProxy competes as one family (best of 4 lenses × strict/permissive + Sinhala translation arms). OpenAI competes as one family (best of Rathje-construct regression on GPT-4o-mini, GPT-4.1-nano, GPT-5-nano). Lexicon-based and topic-model baselines compete as their own singletons. Primary metric per task type: binary =FVE․Binomial, ordinal =Quadratic Kappa, regression =R², multilabel/multiclass =Macro F1. In the graph below, we represented these all as R²-equivalent scores.

source · methodology/gold_medal_page.html
View the chart full size

Two things worth knowing before reading the graph. On the two leftmost tasks the ChatGPT wrapper makes almost no mistakes — AUC 0.995 on IMDb and 0.994 on Reddit suicidality. Both datasets come from public data-science competitions, so the relationship between those texts and those targets was most likely in the model’s training data. We read those two columns as a memorisation ceiling rather than a measurement of reading. Separately, the topic-model hyperparameters were tuned on the training folds: left at their defaults they returned R² near zero on almost every task.

Metric policy

binary
Binary = FVE, with AUC/F1 as secondary checks
ordinal
Ordinal = quadratic κ, with within-one/MAE as secondary checks
regression
Regression = R², with RMSE/MAE as secondary checks
multiclass
Multiclass = macro-F1, with accuracy/macro-AUC as secondary checks

Per-task podium

The same ten evaluations the animation walks through, in the same order, scored on the same strict holdout results. Each family enters its best member, so PsyProxy competes once — as whichever lens is strongest on that task. It takes first place on 7 of the ten.

IMDb · Movie reviews: positive or negative sentiment
binary · FVE · strict · 6 families
GLLM wrappers0.898
SPsyProxy · Technology v4.00.668
BDictionary & sentiment lexicons0.286
Redt · Reddit posts: suicidality risk labeling
binary · FVE · strict · 6 families
GLLM wrappers0.879
SPsyProxy · Health v4.00.807
BTopic models0.678
AmzR · Game reviews: five-star rating prediction
regression · · strict · 6 families
GPsyProxy · Technology v4.00.691
SConcept mining0.458
BDictionary & sentiment lexicons0.339
DisR · Theme park reviews: star rating
regression · · strict · 6 families
GPsyProxy · Technology v4.00.631
SConcept mining0.380
BTopic models0.320
DisB · Disney reviews: Hong Kong vs. California
binary · FVE · strict · 6 families
GPsyProxy · Psychology v4.00.541
STopic models0.456
BConcept mining0.261
DrugB · Medication benefit ratings
regression · · strict · 6 families
GPsyProxy · Health v4.00.274
STopic models0.156
BLinguistic feature tools0.138
AmzV · Verified purchase classification
binary · FVE · strict · 6 families
GPsyProxy · Technology v4.00.215
SLinguistic feature tools0.177
BTopic models0.171
DrugT · Treatment effectiveness
ordinal · Quad. κ · strict · 6 families
GPsyProxy · Technology v4.00.537
SConcept mining0.342
BTopic models0.302
LIAR · Political truthfulness
ordinal · Quad. κ · strict · 6 families
GPsyProxy · Psychology v4.00.272
SDictionary & sentiment lexicons0.208
BLinguistic feature tools0.183
Cmpl · Financial complaint resolution
binary · FVE · strict · 6 families
GTopic models0.066
SPsyProxy · Psychology v4.00.058
BDictionary & sentiment lexicons0.043