Evaluating LLM Reasoning in the Operations Research Domain with ORQA
AAAI Conference on Artificial Intelligence (AAAI-25)
Mahdi Mostajabdaveh · Timothy T. Yu · Samarendra Chandan Bindu Dash · Rindranirina Ramamonjison · Jabo Serge Byusa · Giuseppe Carenini · Zirui Zhou · Yong Zhang
Research introducing ORQA, an Operations Research Question Answering benchmark designed to evaluate the ability of large language models to reason through complex operations-research and mathematical-optimization problems.
DOI 10.1609/aaai.v39i23.34673
