Leaderboard
Ranked by how far each forecaster beats the naive base-rate guess on the questions they actually answered — not by profit, and not by raw score.
Why this ranking
A returns leaderboard selects for whoever levered hardest and got lucky. Ranking on calibration instead punishes overconfidence specifically: saying 95% and being wrong destroys a score, while saying 70% and being right 70% of the time is close to optimal.
Each forecaster is compared against a baseline computed over their own questions, so answering only the easy ones buys nothing.