icon

InfiniteScienceGym

Evaluating large language models' data analysis capabilities on controlled, simulated scientific data repositories.

Leaderboard

# Model Type Accuracy ↑ Unans. Pre. ↑ Unans. Rec. ↑ Toks. (k) / Q ↓ Tool Calls / Q ↓
1 GPT-5.4 OpenAI API 44.8 82.1 82.4 24.3 7.6
2 Claude Opus 4.6 Anthropic API 35.5 80.4 82.6 61.7 7.2
3 GPT-OSS 20B OpenAI Open 29.1 79.8 39.0 80.8 2.4
4 Qwen3 4B Instruct Alibaba Open 24.6 79.8 32.0 34.1 1.6
5 Gemma 3 27B it Google Open 23.1 80.8 40.8 66.0 1.9
Open Open weights API Proprietary / closed weights
↑ Higher is better ↓ Lower is better

Contribute

To add a model to the leaderboard, please open a GitHub issue including the information below, or email at oliver.bentham@utah.edu. Self-reported scores are accepted but highlighted as such until independently verified. We reserve the right to remove entries that cannot be reproduced.

Include: model name, organization, access type (open / API), parameter count (if known), a link to the model card or paper, and your reported scores. We may independently verify results before listing.