The machines sat the actual rationality exam. Here is what still fools them.
The distribution
Total CART points out of 123 across all 16 scored subtests, using the authors' own item keys, thresholds and raw-to-CART conversion tables. No human norms are drawn on this chart: the published norms are for the 148-point original form and the authors caution against treating the CART as a standardized test.
Leaderboard
Tap a model for its 16-subtest breakdown and every item it missed. Models marked CLAUDE CODE took the test inside an agent session rather than through the API.
What trips models up
Left: average share of available points per subtest across API-tested models. Right: the individual items missed by the largest share of models. The preference-anomaly and framing subtests are the last places where language models still behave like people.