What we did, in plain English
We made up four teen athlete profiles — one per country, each based on published research on how stress shows up there. Each teen got ten realistic pressure moments: the night before a big match, the minute before a serve, the ride home after a loss. We asked four AI assistants — ChatGPT, Claude, Gemini, and Perplexity — for advice on all forty. Once in April 2026 (v1), and again with the exact same prompts in July 2026 (v2), after all four shipped new models.
That's 160 answers per version — 320 graded so far. Every answer was graded twice — two separate passes, in two different orders, that couldn't see each other's scores — then averaged. When the two passes disagreed a lot, the answer got flagged for a person to settle: 32 answers were flagged in April (that review never got finished), and zero in July — the July passes agreed within one point 98% of the time.
One thing to know: the grading AI was Claude, which is also one of the four being graded. That's disclosed on the Methods page along with what we did to keep the grading fair — separate blind passes and written-down grading rules. Claude did not come out on top.
No trick questions. Just everyday wording a 15-year-old might type before practice.
Why bother? Because teens ask AI for advice all the time — before games, during class, late at night — and when that advice comes from one culture only, a lot of teens quietly learn the help was not made for them. That is not a reason to stop asking AIs for help. It is a reason to know what kind of help you are getting — and for the people building these tools to widen whose pressure they understand.
4
synthetic teen profiles
320
graded answers — 160 per version, two versions so far