RQ Bench · the real CART ·

The machines sat the actual rationality exam. Here is what still fools them.

The distribution

Total CART points out of 123 across all 16 scored subtests, using the authors' own item keys, thresholds and raw-to-CART conversion tables. No human norms are drawn on this chart: the published norms are for the 148-point original form and the authors caution against treating the CART as a standardized test.

Leaderboard

Tap a model for its 16-subtest breakdown and every item it missed. Models marked CLAUDE CODE took the test inside an agent session rather than through the API.

What trips models up

Left: average share of available points per subtest across API-tested models. Right: the individual items missed by the largest share of models. The preference-anomaly and framing subtests are the last places where language models still behave like people.

By subtest

Hardest items