IMO LLM Benchmark Suite
6.5
A modular software suite for rigorously benchmarking Large Language Models (LLMs) on International Mathematical Olympiad (IMO) problems, providing a standardized and reproducible evaluation framework.
120h
mvp estimate
6.5
viability grade
3
views
technology stack
Python
Medium
SQLite
inspired by
LLMs benchmarked on International Mathematical Olympiad problems