Leaderboard
| # | Model | Type | Accuracy ↑ | Unans. Pre. ↑ | Unans. Rec. ↑ | Toks. (k) / Q ↓ | Tool Calls / Q ↓ |
|---|---|---|---|---|---|---|---|
| 1 | GPT-5.4 OpenAI | API | 44.8 | 82.1 | 82.4 | 24.3 | 7.6 |
| 2 | Claude Opus 4.6 Anthropic | API | 35.5 | 80.4 | 82.6 | 61.7 | 7.2 |
| 3 | GPT-OSS 20B OpenAI | Open | 29.1 | 79.8 | 39.0 | 80.8 | 2.4 |
| 4 | Qwen3 4B Instruct Alibaba | Open | 24.6 | 79.8 | 32.0 | 34.1 | 1.6 |
| 5 | Gemma 3 27B it Google | Open | 23.1 | 80.8 | 40.8 | 66.0 | 1.9 |
Open
Open weights
API
Proprietary / closed weights
↑ Higher is better
↓ Lower is better
Contribute
To add a model to the leaderboard, please open a GitHub issue including the information below, or email at oliver.bentham@utah.edu. Self-reported scores are accepted but highlighted as such until independently verified. We reserve the right to remove entries that cannot be reproduced.
Include: model name, organization, access type (open / API), parameter count (if known), a link to the model card or paper, and your reported scores. We may independently verify results before listing.