Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLM Reasoning in the Operations Research Domain with ORQA

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2412.17874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17874 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T06:01:13.464975Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:04:55.047752Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:53:26.427900Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 68d6994c-3dc5-4957-968c-396688aaeb8a · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.249266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.213502Z digest=sha256:78000e9135bf494286bcb8a6e5acb8fdce39d0c0836318a77c3b2d78d43d55ab

Observation 12f0f0e2-2d3a-407f-969d-8376ec9c58a1 · outbound

This paper cites The Falcon Series of Open Language Models.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA The Falcon Series of Open Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.219136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.219136Z digest=sha256:1c8e44302913b66db6eca6c9337be755515a202d70dd7bdb8c4a871c1716d5f7

Observation dc6b0e4d-8ae7-4e3b-b0c6-2ed46a0dbf33 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.224508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.224508Z digest=sha256:d567556b55df1928c3588e219d8e61c4e24badf689065f9a0ac502dc24b30d15

Observation 22f4c639-14a5-40f2-ac9a-b50adc2ac2c2 · outbound

This paper cites GPT-4 Can't Reason.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA GPT-4 Can't Reason

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.229345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.229345Z digest=sha256:eb43626dbddf31710edc5002390c59ca301f0a7de3a91778c554c88cc47f8158

Observation 445d8fbe-c42f-4526-a14a-cd79a4b33107 · outbound

This paper cites Gpt-4: A Review on Advancements and Opportunities in Natural Language Processing.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Gpt-4: A Review on Advancements and Opportunities in Natural Language Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.234417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.234417Z digest=sha256:8ecd21431ce37d05fb28b5a39dbc622b9105dc42b4011755baab5af1f0f3446f

Observation 8d10d948-8641-484e-9d54-a332e500303a · outbound

This paper cites W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.239163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.239163Z digest=sha256:eef3b86e5c8ed50264a8473616d9fd489ddd56fda5fda75991a93791c410374f

Observation 2c6a37fd-fbfb-4db2-9ecc-5db8c75c4a51 · outbound

This paper cites H.; Choi, E.; Collins, M.; Garrette, D.; Kwiatkowski, T.; Nikolaev, V.; and Palomaki, J.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA H.; Choi, E.; Collins, M.; Garrette, D.; Kwiatkowski, T.; Nikolaev, V.; and Palomaki, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.225689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.244043Z digest=sha256:b1719ff56996df98bf591dbdd6661b0f7f6654ad726f9cdae8a0fe40a59e3bd3

Observation a632a6cd-fa12-471a-b467-fcec5e46d3f2 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.211088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.248227Z digest=sha256:6cb82907e064a4a480319033d16d415f2f3398bf720a8bc5983bddd892b86785

Observation 29851cbd-16e7-4c4a-8a40-82bae0bf2f07 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.252630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.252630Z digest=sha256:8dbb081bc04600fda71ea929935fd3f620f6934b00eb6e2353aab75bfa2171b6

Observation a54a65ce-f3ea-48d0-a5f4-99a9eadabeab · outbound

This paper cites L.; Jiang, L.; Lin, B.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA L.; Jiang, L.; Lin, B

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.196973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.261035Z digest=sha256:64afb73a703c77da0c1c4782afc38bd61fd7f61e01d15c00b2393a25c980561a

Observation f84d7a43-34f2-4587-bee8-14eea64c0160 · outbound

This paper cites Artificial Intelligence for Operations Research: Revolutionizing the Operations Research Process.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Artificial Intelligence for Operations Research: Revolutionizing the Operations Research Process

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.266814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.266814Z digest=sha256:d8a7b6414a6f73ae4f6e80e9f2e59f85326e0742b36fe91cd76254372a85d97f

Observation 6322410d-4b80-4556-aa8e-730d1c9c4351 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.271099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.271099Z digest=sha256:f218f4e4ceaeef2a997e76a001aa1b5a6cb297293e9ca156e63efa81a7ca8ed5

Observation 0a493bf4-5146-420c-91ff-d0e72e32099e · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.275868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.275868Z digest=sha256:fb2e7c2bbafd3036f8fbe2a40940150e26e098aad190f2ca699852ca8c8f4eb5

Observation bc6ae4a2-945b-4b15-b85e-7614a78c1903 · outbound

This paper cites S.; and Lieberman, G.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA S.; and Lieberman, G

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.183521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.280616Z digest=sha256:3323a056922e826ee134e413d377141d8fcab4982310daa3e2134066fb66cb0a

Observation fd90435d-9fff-4f73-a00c-001554724563 · outbound

This paper cites Mistral 7B.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.285048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.285048Z digest=sha256:9c0b1f8647daaf530ede107411749a15d260970d28e0fe613dd2db8ed66f80e8

Observation 6ceb40ee-cf92-43d1-a83e-d2965699681b · outbound

This paper cites Mixtral of Experts.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Mixtral of Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.289632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.289632Z digest=sha256:7764d61cf66f7414ddea246a0c153006d4403807acf4b706119d67ee58329b38

Observation 22ab3a8c-1f56-46d7-9c10-5991548ad834 · outbound

This paper cites K.; Cappart, Q.; Rousseau, L.-M.; and Laurent, T.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA K.; Cappart, Q.; Rousseau, L.-M.; and Laurent, T

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.294120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.294120Z digest=sha256:6dd6e123e7a236e0cbb6fb2e679e58f6671fb3217c6bd864fd5961991b16b5ca

Observation 02e260e8-575c-4937-8bb4-51bdb5ccf500 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.298871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.298871Z digest=sha256:c02b43a85fe6cbc83f6bfee6813e436768a29a13b510ffe4a3f6bf37782ac7bc

Observation 9b97d735-46c1-4aed-a19a-68178faf30b0 · outbound

This paper cites A Study on Large Language Models' Limitations in Multiple-Choice Question Answering.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA A Study on Large Language Models' Limitations in Multiple-Choice Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.308137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.308137Z digest=sha256:463b77e060e0a374348be6fa08aa1d26f7174713088cd177cd063376632aabaa

Observation f8f16f9f-9f07-423a-8cad-dc87cde89c62 · outbound

This paper cites S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.312704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.312704Z digest=sha256:604063be712e5875c2eaefbaceb878bbbb33dfd7798ee8e284e933e9f9a20957

Observation 8d2d1c1e-d381-429d-8875-495f8f4069d2 · outbound

This paper cites EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.316678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.316678Z digest=sha256:5d819d1e03889d010a5700536f1f26b711163e2f5f519b3d81999ee51eff1e3b

Observation df5805a9-b5cd-4fc3-92bc-c560ea5a3497 · outbound

This paper cites LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.320631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.320631Z digest=sha256:189e91173a91447afff3c621b26e1918e1d4dd1a59501cc7d9ec2689dffd7d1d

Observation fa20c8d6-e5c0-4a8a-8c66-111d6d2a58ff · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.324531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.324531Z digest=sha256:0d2ff7eba8caafead24ca265209f23703ff8d5bef0e6c40d7cc8a32f86bdd013

Observation c38f5b1e-0639-411a-8d06-a7e29e619112 · outbound

This paper cites E.; Motzfeldt, A.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA E.; Motzfeldt, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.140950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.328477Z digest=sha256:a12218044142ab55e3225e252ec4ce72d02e4801c8443a007f38c7fbf37dca05

Observation 67538198-086c-4acd-b18e-595df23c8c51 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.127909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.333600Z digest=sha256:32c9131d3f1f6fa32ddda08151234f6f7a6e2dc8afba6c25510cbe889ce92b6e

Observation 879dc9c7-7fda-4a30-8318-e8c7f7d3a988 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.113552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.338709Z digest=sha256:76a52265eb413bba465ee9067e9742094d901c93d2779b2841b344ed89a9f55f

Observation 5aaa81e2-16f5-451b-ab2f-4d0067d0f0c0 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.343137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.343137Z digest=sha256:42ff46d74248b8cfda7cefe26398e7a56e1602ead309643d6e9a685265f7b3a3

Observation 7d44e252-3626-4ea8-b077-37656a4cac12 · outbound

This paper cites S.; and Tahmasbi, N.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA S.; and Tahmasbi, N

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.098629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.347516Z digest=sha256:e2a96540fb5f0f40f02da31b0e23a15dd7b8f310e609f5805ec546ae104da7da

Observation 526960c3-40f1-48db-8458-ce8e01ee7723 · outbound

This paper cites T.; Ramamonjison, R.; Carenini, G.; Zhou, Z.; and Zhang, Y.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA T.; Ramamonjison, R.; Carenini, G.; Zhou, Z.; and Zhang, Y

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.084444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.351578Z digest=sha256:bcdd42ca9505dcb2b6c8c21271f531470f367792cfe42096e062e9ca3741cab9

Observation dfc987ae-847c-49c5-8dd6-83fc84fd472e · outbound

This paper cites A.; Archetti, C.; Ayhan, H.; Battarra, M.; Bennell, J.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA A.; Archetti, C.; Ayhan, H.; Battarra, M.; Bennell, J

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:14.069728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.355746Z digest=sha256:931cd6fe0578abd1e2327752341d40008bf132c25aec353291b9bccb72629db7

Observation 9aef081e-5ccb-4448-bb56-7c90277285b1 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.054526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.359562Z digest=sha256:b6994a25cbdcbe28ab2499eb5d00405c46e0445bb8eda59420e158a981b2ed7a

Observation d6bf7693-3dab-40b2-b25e-4213fa80522c · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.039284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.364144Z digest=sha256:5037a64f4729e36cffa6f61d2d9f091c91cf46267624904a10ab4f3522f731f4

Observation bb14e046-1d4d-426b-a103-e19b177cda99 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.022198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.368764Z digest=sha256:f7e9cd3824a458ffb87d8912b1bf423d952221944d873a6c6efa9e0305f6d012

Observation 16508159-1e93-42b4-b3ac-5f4e58a7ecac · outbound

This paper cites STREET: A Multi-Task Structured Reasoning and Explanation Benchmark.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-11T06:01:13.616408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.372619Z digest=sha256:32699ae53a1125874e56c2c1571c0d127e32353f141eab71299ae2a98a1bea7c

Observation f497bc25-89a4-4546-86bf-35861ca70fe7 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:14.007827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.376923Z digest=sha256:d0607c618cbc713c20200e0fb8c1a97e8df473145ff12a6bde8b1c7dda3bac85

Observation 0d2af395-c654-4ade-bf6c-c034c4342590 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:13.993883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.381085Z digest=sha256:40e8bbeaebad64e49e02b4b426ae17722d6f34798df72c515e20e8291dae91e5

Observation bdd9c261-5c0f-47b3-9e8b-1f8b9dd6a8cd · outbound

This paper cites ARB: Advanced Reasoning Benchmark for Large Language Models.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA ARB: Advanced Reasoning Benchmark for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.384978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.384978Z digest=sha256:c670db7169f0ecd4abdbc19cb93f6585fbc42a86f29edf16f86106804c9dabb1

Observation e8281403-1f21-4c4b-9c86-3f9fc1a176a2 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:13.980261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.389220Z digest=sha256:1fd69c578e3520f10896df9a2449bc1247a669b881ed3e71200128ad0789018d

Observation 8e6d57ac-5ea3-41b6-b2b0-3c2b8ff5ae7b · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.393464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.393464Z digest=sha256:51ec7bf6ce60ec97d6fbecf485b9d7b6b189462c4ac360521c6d044e837053ff

Observation dbb1a25c-22e4-476f-ad6b-0be011c76fc7 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.398115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.398115Z digest=sha256:792f898f82a44eacb011928bf68b4f6d7dbf72e70c07280b4f3f7bfd61e822a8

Observation f0250fb9-a761-4cee-b6dc-563a20532bd4 · outbound

This paper cites H.; Baldwin, T.; Verspoor, K.; and Cohn, T.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA H.; Baldwin, T.; Verspoor, K.; and Cohn, T

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:13.966992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.403477Z digest=sha256:2b8b069a545ded189ec5875825971cde722509750a5f30961188f4c0d96394b2

Observation fdac3854-9208-4f7e-b8a9-0ce2b26796fe · outbound

This paper cites PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.407994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.407994Z digest=sha256:976ec35dd5bdc5e8746a6edf7573ee394c5a4a02db7dfc68a62d9fa24e3e9d69

Observation a8462897-dfda-474e-ae4b-9c3520581c96 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.412714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.412714Z digest=sha256:d550d9d05109b6775a43c38cd9734e33eb916fbb34bbd982f5236fe79da5355b

Observation 960a3e2a-c468-41eb-b4f4-2c22536c5e4f · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:13.942681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.417335Z digest=sha256:e1ef009ee459b4821473c06510009eca2aefada3c5ab82dc8ef171dade142316

Observation e42cd0a4-f71f-42ed-aec8-ed3db2ce6619 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.422045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.422045Z digest=sha256:68e757a7c4b76e9e05271d4b039c91be94b3773a2deba9c75b9507e2127c9d5e

Observation af23395b-a856-40fa-a939-58714baa9e45 · outbound

This paper cites V.; Zhou, D.; et al.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA V.; Zhou, D.; et al

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.426697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.426697Z digest=sha256:d157bea45a4456a8ea0f07a7d1591734e8e6744b65eac658e894e4313562a5b1

Observation 025e930c-52f9-484c-941a-0987afd78e79 · outbound

This paper cites J.; Han, X.; Fu, X.; Zhong, T.; Zeng, J.; Song, M.; et al.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA J.; Han, X.; Fu, X.; Zhong, T.; Zeng, J.; Song, M.; et al

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T06:01:13.919078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.430657Z digest=sha256:54bf96bc0f9efecc3a39a92566c544edbd62a1a6e87cfd04f435c23c68482421

Observation c60de498-ef34-4493-a227-76fdf4730b67 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.434581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.434581Z digest=sha256:cb44221eb277aba2340e474005e9a8a16363b6efa5341a9e71e8c08fb4969adf

Observation ebd57184-84af-43c4-b0db-b0b26e5327e8 · outbound

This paper cites R.; and Cao, Y.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA R.; and Cao, Y

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.438814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.438814Z digest=sha256:ed006bf46061b855a4394291c8bdb98315cfda78bbd1ff5cac5189385421c380

Observation f59b197f-d882-4a78-bf39-e7ad611c1cd7 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:13.886557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.442771Z digest=sha256:d8802bf03fc06a8aceef3f06e547f1a33e709be59941f70c263aa27b6b42418a

Observation 69e2fd78-8702-4995-b7b9-e355741499c9 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:01:13.873261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T06:01:13.446897Z digest=sha256:1d48b0784f4bfb958502651f25289548634a4ed78b5d414a66fa7800f893f8a6

Observation 58e7c573-3579-4e93-b162-8016e5747c45 · outbound

This paper cites A Survey of Large Language Models in Medicine: Progress, Application, and Challenge.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA A Survey of Large Language Models in Medicine: Progress, Application, and Challenge

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.450801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.450801Z digest=sha256:2cfd965c84bde558e9828ee0ba4a95e85b4e96c3fca8212433c743331ec85e2d

Observation 0464d609-0546-4789-bff1-a2d653934511 · outbound

This paper cites Self-Discover: Large Language Models Self-Compose Reasoning Structures.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA Self-Discover: Large Language Models Self-Compose Reasoning Structures

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.455187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.455187Z digest=sha256:70993adb80e29848c9234357a9cebec6087bee7667e26a7ef66a79dc330a7987

Observation a3653acd-2a51-433b-9a5e-c1df42a0fb71 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA , " * write output.state after.block = add.period write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.459761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.459761Z digest=sha256:109f1770d7e3b05b1a3092fc966a6047b771e1eff417e621fa0c392b35b6679f

Observation 1025bf59-9919-42cc-b7ea-322f029ad3b4 · outbound

This paper cites write newline.

Evaluating LLM Reasoning in the Operations Research Domain with ORQA write newline

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T06:01:13.464975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:01:13.464975Z digest=sha256:32b1817108bc5feda7e77b2a974102de523858d58404b0f4f250079f718876ca

Pith citing papers

Observation 1bd31f04-0455-440b-8edd-2387edd8cb9a · inbound

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents cites this paper.

OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agents Evaluating LLM Reasoning in the Operations Research Domain with ORQA

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:53:26.429358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T12:52:20.911788Z digest=sha256:e46b352b4a43203f4f6d8fb5ca934eb3a7e2a54d7faf56bf9cf602bb9c0b72c4

Observation e5411e1e-52e1-4d6c-a80c-8aaaa983226d · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language Evaluating LLM Reasoning in the Operations Research Domain with ORQA

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:55.047752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:55.047752Z digest=sha256:f6b80ab50aa73b81f411c97b983f71b686c4a97fa8bd5bba93cafaab70360eb9