The machines sat a human rationality exam. Almost all of them passed.
The distribution
RQ is scaled like IQ: the human average is 100 with a standard deviation of 15, mapped from the test's assumed human score distribution (mean 60%, SD 15%).
Leaderboard
Tap a model for its 17-module breakdown. Scored with the test's own open-source grading code — same thresholds, CART composite curves and all.
What trips models up
Average module score across the whole field. Language models turn out to be superhumanly immune to superstition and conspiracy — and just as frameable as we are.