Worked example

Worked example: Amazon video-game reviews

Predicting star ratings from review text alone — with an equation you can actually read.

The task: take 299,990 verified-purchase Amazon video-game reviews (a public Kaggle corpus) and predict each review’s 1–5 star rating from its text alone.

Review 1 of 10real review · verbatim
I bought this and let it charge up, put it in my Balance Board and nothing happened. I also bought another to give as a gift. The joke was on me cause neither one of them worked!
1 of 5
the reviewer’s actual star rating
Ten real reviews from the corpus, chosen to span the full 1–5 range and to carry strong signal on the proxies in the equation below. Scroll with the arrows — the rating follows the text.

PsyProxy scores every review on a fixed set of psychological proxy dimensions — roughly a thousand named, defined measures of psychological content that existed before this dataset was ever seen.

A deeper-structure topic model

Before any outcome is modeled, this scoring stage behaves like a topic model — but one that works a level deeper. A topic model re-derives surface word-co-occurrence clusters from scratch on every new corpus; PsyProxy instead arrives with a fixed factor-analytic structure distilled from decades of psychological measurement, and simply measures how strongly each dimension stirs in each text. Point it at 299,990 video-game reviews and the dimensions with the most signal spread read like a table of contents for the corpus:

Excitement index · top 15 proxies on this dataset · Psychology lens · N = 299,990
Gaming Addiction4.58
Product Evaluation2.91
Wearability2.78
Entertainment Attitude2.78
Tool Usability2.55
Gift-Giving2.52
Enjoyment2.42
Importance2.35
Usability2.29
Product Return Orientationin the equation2.29
Network Engagement2.25
System Reliability2.14
Price Evaluation2.10
Brand Identity2.06
Software Piracy Intention2.00
The excitement index is a proxy’s signal spread — its standard deviation across all 299,990 reviews. High spread means the corpus keeps expressing different amounts of that dimension: gaming addiction, product evaluation, gift-giving, price, piracy. Nobody told the lens this was a gaming corpus.

The differences from a topic model are what make the scores usable downstream: these dimensions are not re-fit to each corpus, so results are comparable across datasets; each one is a psychological measure with a name, a definition, and score bands rather than an unlabeled word cluster; and the structure captures what the text expresses psychologically, not just which words travel together. Note also that excitement measures spread, not predictive power — only one of the fifteen most-excited proxies was selected into the rating equation below.

A plain linear regression then maps those proxy scores to the star rating. No fine-tuning, no black box — PsyProxy is the best predictor on this dataset overall, and the model it produces is a single readable equation.

For interpretation we use the Psychology lens, because it best illustrates concept-linked consumer evaluation — even though another PsyProxy lens is slightly stronger for raw prediction.

The equation

The compact model selects fifteen proxies. Every proxy in the equation is a documented measure: it has a name, a definition of what it measures, and score bands that describe what low, moderate, and high activations look like in real text. The coefficient is its weight; 3.887 is the intercept.

Target = 3.887
  - 0.132 * Frustration
  - 0.173 * Perceived Time Waste
  - 0.004 * Product Return Orientation
  + 0.124 * Perceived Usefulness
  - 0.098 * Dominance Orientation
  - 0.135 * Negative Person Evaluation
  - 0.112 * Ingratiating Praise
  - 0.088 * Disruption Impact
  - 0.144 * Social Ease
  + 0.041 * Inspiration
  - 0.177 * Moral Wrongness Appraisal
  - 0.146 * Work Valuation
  + 0.156 * Global Life Satisfaction
  - 0.119 * Interdependent Happiness
  - 0.120 * Achievement Drive
  [PsyProxy.ai/Lens/PL99]

Fitted on the Psychology lens v0.99 as published in the PsyProxy paper; a refresh on the current Psychology v4.0 lens is queued.

Here is what one full dictionary entry looks like — verbatim from the current Psychology lens (v4.0), which names every proxy through a governed, evidence-only pipeline:

anger experiencePsychology lens v4.0 · proxy 158 of 500

anger experience measures the presence and salience of anger in expressions, from little clear anger content to explicit anger that becomes a prominent personal state. It also encompasses anger described as affecting attention, judgment, interests, or wellbeing.

minimal anger experienceExpressions activating this band contain little or no clear anger content. Anger is absent, too weakly expressed, or not diagnostically central to the expression.
explicit anger experienceExpressions activating this band directly state feeling angry or having anger, often in reaction to a person, action, or situation. Anger is clearly present as an emotional response, but the expression does not center on broader disruption to functioning.
disruptive anger experienceExpressions activating this band describe anger as a strong, recurring, or pervasive personal state that can interfere with interest in others, decision-making, health, or wellbeing. The expression presents anger as prominent enough to affect functioning or self-regulation.
A proxy’s dictionary entry: the name, the definition of what it measures, and its score bands from lowest to highest activation. When a review scores high on a proxy in the equation, these bands say what that score means in plain language.

Reading the equation

Each coefficient tells you what pushes a predicted rating up or down. The signs behave exactly the way anyone who has read product reviews would expect.

What raises the rating

  • Global Life Satisfaction (+0.156)— the largest positive weight. An overall evaluative judgment that things are satisfying, pleasing, or going well, here applied to a specific customer experience.
  • Perceived Usefulness (+0.124)— reviews that describe the product as genuinely useful predict higher stars.
  • Inspiration (+0.041)— a smaller but positive lift when the product sparks enthusiasm.

What lowers the rating

  • Moral Wrongness Appraisal (−0.177)— the most negative weight. The reviews scoring highest on this proxy accuse sellers of scams, false advertising, misdescribed products, and wasted money. Those moral-transgression complaints are overwhelmingly one- and two-star reviews, so the proxy correlates negatively with the rating (r = −0.38).
  • Perceived Time Waste (−0.173)— feeling the product wasted the customer’s time is nearly as damaging.
  • Frustration (−0.132)— expressed frustration drags the predicted rating down.
  • Ingratiating Praise (−0.112)— flattery-style praise reads as insincere and, counterintuitively, predicts lower ratings.

Selected descriptive statistics

A few of the proxies from the fitted model, scored over all 299,990 reviews. The target (star rating) is on its native 1–5 scale; proxy scores are on the lens’s projection scale.

VariableMeanSD
Star rating (1–5) — the target4.021.41
Frustration1.871.78
Perceived Usefulness−1.321.27
Global Life Satisfaction0.200.81
Moral Wrongness Appraisal−1.181.21
Perceived Time Waste−0.831.65

N = 299,990 for every row.

The reviews again — through the equation’s eyes

Here are the same ten reviews, this time with the equation’s proxies color-coded at the sentence-part level. Each highlighted part is where a proxy fires when the text is measured — the human-readable trace of how the prediction is assembled.

Review 1 of 10proxy-coded
I bought this and let it charge up, put it in my Balance Board and nothing happened. I also bought another to give as a gift. The joke was on me cause neither one of them worked!
FrustrationMoral Wrongness AppraisalDisruption Impact
1 of 5
the reviewer’s actual star rating
The same ten reviews, now coded: each highlighted sentence part is where one of the equation’s proxies fires (its strongest proxy, at least 1.5 standard deviations above the corpus norm, measured live through the Psychology lens). ▲ = pushes the predicted rating up, ▼ = pulls it down. Hover a highlight for the exact score.

Where to go next

See how this result stacks up against other systems on the benchmarks page, or explore the proxy dimensions behind these measurements on the lenses page.