Datasets·Social science·sold_sinhala

SOLD Sinhala (4-arm translator comparison)

Translator-comparison study (4 arms): Gemma / Google / GPT-4o-mini / GPT-5.4 against the SOLD Sinhala offensive-language detection task. See eval14_sold_sinhala_* runs for per-arm cards.

Distribution of SOLD offensive vs not (Sinhala -> English translated)
1
2
5,809 at floor4,191 at ceiling
10,000
items
2,000
holdout n
SOLD offensive vs not (Sinhala -> English translated)
target
Binary
kind
4
systems compared
Criterion validity

Reported holdout systems from the verified card

Binary classification uses FVE as the task-primary metric. Secondary columns keep the companion metrics visible so binary, ordinal, regression, and multiclass cards are not compared through one flattened score.

Model-family mix
OpenAI / LLM · 2Baseline · 2
SystemFamilyVariantFVEAUCF1Primary scale
llmHealth Lens via GPT-5.4 translation
OpenAI / LLMpermissive0.1850.7840.646
baselineHealth Lens via Google Translate
Baselinepermissive0.0980.6990.552
baselineHealth Lens via Gemma translation
Baselinepermissive0.0830.6850.521
llmHealth Lens via GPT-4o-mini translation
OpenAI / LLMpermissive0.0680.6750.502