Pith. sign in

Paper Citation Record · LEDGER

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2505.24823.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24823 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:22:26.217604Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a48d0e-96b2-47ab-837b-9a9a873b4ed1 · outbound

This paper cites A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.263718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.263718Z digest=sha256:d302f71aa8c149b826d5e50470be6f4e32ba274261b7c0d0458aa75c20a4d4c6

Observation 808a5c32-b05e-48b9-99ca-4c0c0033234a · outbound

This paper cites Romera-Paredes, M.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Romera-Paredes, M

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.403429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.403429Z digest=sha256:9bd50cd88f730669b2426c5f8b44e6056797c30bf6ac3d2e0f82d21d81aa3557

Observation fe59b0df-1003-4619-bb2e-5994d0365be5 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.541012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.541012Z digest=sha256:6a2f19771041d6ea5fb5d6e49e241fe0dd9dbd529fe380e8bd85ffdb5ff935a2

Observation df50c6f1-8be3-4d9d-915d-7c18ab264605 · outbound

This paper cites Llm-sr: Scientific equation discovery via programming with large language models,.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Llm-sr: Scientific equation discovery via programming with large language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.507274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:22.670967Z digest=sha256:4266eb36e57e5f21dc9d2087072cf217ef0d2648a5f74693d28fffba5973d57a

Observation a364e9d4-33f0-4d5c-93a0-a456a9d46440 · outbound

This paper cites Brenner, and Eun-Ah Kim.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Brenner, and Eun-Ah Kim

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.854654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.854654Z digest=sha256:51db5b0bc4424e295bae097ab5ab4707eb3b9e65689325d11f1c3deb01ba4f03

Observation 51298dbc-c492-4a51-8376-bb4d517b523d · outbound

This paper cites LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-Feynman: Leveraging Large Language Models for Universal Scientific Formula and Theory Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.956783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.956783Z digest=sha256:6dbb9158c8dc0ca6f1ef86351174dc933c20e1ee4649a2cb9b720b3003f65490

Observation 32b624ba-dac3-4396-a211-711b3a7bd2a1 · outbound

This paper cites Large physics models: Towards a collaborative approach with large language models and foundation models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Large physics models: Towards a collaborative approach with large language models and foundation models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.042992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.042992Z digest=sha256:2f81d215b401a172db88308afe7f3b724fcf5c0f20898b38c3eb9c82c06fd062

Observation c4563436-9104-4f8f-87b4-a07a5a2783fb · outbound

This paper cites Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Advancing AI-Scientist Understanding: Multi-Agent LLMs with Interpretable Physics Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.142195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.142195Z digest=sha256:bf7c813ae6545ca88f05623702a4475a55280096a8aa9fb277f92c4587872a32

Observation 0b2493fa-4321-4818-a7ef-0128d76ea2ca · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini Robotics: Bringing AI into the Physical World

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.253023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.253023Z digest=sha256:fc187dd6727f21334332f1d7673eb8e8f7af474921adf1320227acd57ec09557

Observation 396fc029-6d63-4220-a6ad-614e4ce1b101 · outbound

This paper cites Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Theoretical Physics Benchmark (TPBench) -- a Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.366842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.366842Z digest=sha256:66e2b53bc43dc412b755a494d48224846ae453b68bfecbcebf02978aa7f4fb98

Observation 3f783fd1-951f-41cb-9688-782fa1291086 · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.478351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.478351Z digest=sha256:61b6cbf561108a840cd16dd2bf78e22b727403f390dbee884cdb3397355ada9e

Observation 803b530e-0457-400d-a200-728c2ce4442f · outbound

This paper cites Expert and novice performance in solving physics problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Expert and novice performance in solving physics problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.365459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:23.556910Z digest=sha256:21fa0cbdff34767d8ae42b4c05bfcedcd4ed3b01db9b209f839525133b65c0b8

Observation 9371a1df-69bf-44c3-85d2-988571c33163 · outbound

This paper cites an unresolved cited work.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.661672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.661672Z digest=sha256:ccb891a2dea2dd52ecd9c5e67948af72bdfe5e81fb239cfb63077c059c671afb

Observation f7c2d1ae-54f7-48bf-9368-62c51f1b0030 · outbound

This paper cites Cognitive load during problem solving: Effects on learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Cognitive load during problem solving: Effects on learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.740680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.740680Z digest=sha256:cdde22137fb7f41d3485511caae130f83b7dfe419cc9ace8acb3d83c849ad546

Observation 4b35877a-f8ac-416f-aebc-eba31f4ffe22 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.833827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.833827Z digest=sha256:c3c28c8990e278cbfa25d3b562df879a6aeef80af1b324b9713891c92f16dd03

Observation 521599b2-9065-4a7f-931a-2bbcfdfd0ee4 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.924503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.924503Z digest=sha256:3ea82af873509c1eaae0e986857bed95c5c0f97b02b4767cb21f195ab8ddeb53

Observation dbb57740-cea8-471f-bcbf-fc7a73709532 · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:23.996077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:23.996077Z digest=sha256:94cb1df833263bdfc28581ef25b35ca570c0e6a0dae4f0d5c70b322e2cde615f

Observation f4e69425-a990-4618-bc0a-e1b6462b638a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.060193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.060193Z digest=sha256:0ab8a647b69a56133078bfd53d86e37d8f65707c494a5c7bc026eb5fd162444e

Observation cdf12970-07b3-43f0-bc47-85ff7762fa00 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.129251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.129251Z digest=sha256:9ad19c35b0e32fc61e93fd19f0e9e2bd774c4da77bd435e89b7c1ec7182de9c0

Observation 79cc130c-bbb1-43d2-b840-5dcaa3d9d231 · outbound

This paper cites Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Have LLMs advanced enough? a chal- lenging problem solving benchmark for large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.207587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.207587Z digest=sha256:727651f27efc4bee7adfd2980ab42ea7046511c424a155dd7eb89e453913738c

Observation 02dda92c-2780-45b9-8f2c-1f865cb1b5ae · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.297967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.297967Z digest=sha256:14fb3fffaf7f888f568b7ca44a10e0cc72801d31d2683d120b495eb1e9034edf

Observation 0c387862-5d82-4e16-9e86-65313924be74 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.372229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.372229Z digest=sha256:a70505d4e9c95d351b5f1e8277fa937ec48139cbc1cce7e2f68d5fd303b16080

Observation a869ab84-d082-48b6-ba44-90cd9788c3f8 · outbound

This paper cites Scieval: A multi-level large language model evaluation benchmark for scientific research.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Scieval: A multi-level large language model evaluation benchmark for scientific research

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:28.130295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:24.440236Z digest=sha256:4a03fc3afff0ca2b7cfadf510cff55c947bd44c4d0523387443f015268817096

Observation 5a431d61-daf8-4a7e-897a-e5c964430a73 · outbound

This paper cites TheoremQA: A Theorem-driven Question Answering dataset.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models TheoremQA: A Theorem-driven Question Answering dataset

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.562397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.562397Z digest=sha256:45b28e25023c47bf3054cf9283b53c800261f5b2b698ad3ff4eba7131efa555f

Observation af06dde0-a12c-42a3-ade6-77517c781a61 · outbound

This paper cites Humanity's Last Exam.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Humanity's Last Exam

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.630990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.630990Z digest=sha256:51d1f792fc455b2af4744c8efbd4a624c6398a3fd9565f2db62e4730d1058061

Observation 4a96a0a4-9d2a-495f-b9bc-2f049d547835 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.704748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.704748Z digest=sha256:79d1e2471acf7cbac0d2d71eb4ae58e2dc9b17f9a88df75bda58a855421b7e7f

Observation 180c217d-8b14-420f-b85d-f8106afa1290 · outbound

This paper cites Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Olympicarena: Benchmarking multi-discipline cognitive reasoning for superintelligent ai

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.776896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.776896Z digest=sha256:61682c75cdc950c90d7c9fccf68533684665b8e7bcb454410caa5a3dd20bdfab

Observation 9b4783e0-6f33-4c67-a1ba-a05baec08497 · outbound

This paper cites Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.850334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.850334Z digest=sha256:23971bb0497a6c6245b633dbda70280c3ea76ae831dc47348717601e8a751612

Observation 03c59c30-0f57-4c2b-9221-927d494fe54b · outbound

This paper cites UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:24.926698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:24.926698Z digest=sha256:d8209f5d821bdf0b4604891b9e7808573aab7a90e4de4f8895dd768958012d12

Observation b1401a96-4fdf-4fce-8ad7-6f18ae4f77ba · outbound

This paper cites PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.013730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.013730Z digest=sha256:627a7cb81c10ef4a20100b4f14f9f79bdea523cb3e4f179f2ffe4fd3983a99b3

Observation aa447c21-1001-482c-9c3f-e2f2bc2942c1 · outbound

This paper cites PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.112064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.112064Z digest=sha256:fb37842e9b0bdc2794f213495c38ee2cee554393876a0ccb3df39d950e2ccd1e

Observation 9993ea64-1444-49b2-b6e0-7d4512665de9 · outbound

This paper cites CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.206982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.206982Z digest=sha256:12f2dee21a072db30743d109ae619e674f3af77b7e6f6a35657d1582c423aa73

Observation 18e58079-ba3c-4c38-ab28-ba13d822cc40 · outbound

This paper cites Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Mm-phyqa: Multimodal physics question-answering with multi-image cot prompting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.273400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.273400Z digest=sha256:7def6161b440519dd6a2c6c09005e2fee858f1d23b798151ffb9c24765eb46b4

Observation 6da479d7-7839-4893-9e43-1d28264fe77c · outbound

This paper cites FEABench: Evaluating Language Models on Multiphysics Reasoning Ability.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models FEABench: Evaluating Language Models on Multiphysics Reasoning Ability

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.362149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.362149Z digest=sha256:04350e1488010bddc92c3fd94d45284fed8edd156e17435ef2b2239572064d75

Observation 7a961bc9-1382-43b1-9101-ff2fc5c1d320 · outbound

This paper cites Learning to reason with llms, September 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Learning to reason with llms, September 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.886863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:25.433618Z digest=sha256:b58648546f9c9d287c16a0c741479e65d64e93abec03cec16a843dd33ec38976

Observation f3766474-28fc-4d52-aaee-35044a14fcff · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.520633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.520633Z digest=sha256:4e153895de05f327f625364d56a2bfd3b4dcf02918a1a04af79adaf5ff96359c

Observation 82d86910-6508-4c19-a1ed-f374995122c0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.610427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.610427Z digest=sha256:8fb4989c16b7c4cde026e9b9a1de4d828762349a2f4a4a2ae70f7175b94129b5

Observation df088cbf-23f6-4b89-ad03-62d780ba7c78 · outbound

This paper cites Claude 3.7 sonnet and extended thinking mode.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and extended thinking mode

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.712882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:25.698133Z digest=sha256:58aa483c4df4483d8b31d0cd6e462828ac569e861527e6ac2036cd8f4c9cdd0f

Observation 9c98da01-6cb7-4932-8275-2d57c4e85b04 · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.536957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:25.762655Z digest=sha256:0fbbd92229a293d5b4ee73a8d93334fe6a082d3cdfe5c431c0d3cdb6afadf412

Observation 610037b8-ea61-4d33-9e03-36693ba95cde · outbound

This paper cites Openai o3 and o4-mini system card.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Openai o3 and o4-mini system card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.353683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:25.833242Z digest=sha256:51bec14d2b8e616f52c2f0379199c20320f97bd5c0efc81d3e47376713427821

Observation bc3d9c65-2030-4145-a99d-14949198f77d · outbound

This paper cites URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models URL https://blog.google/technology/ google-deepmind/gemini-model-thinking-updates-march-2025/

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:25.907140Z digest=sha256:b5ef9c2de5b16b69cf07789ce2439398b5682957371a1f7f603793d6bcd2cdc0

Observation 0c3101cb-d4ed-4803-9e0f-7821cf0c4c4f · outbound

This paper cites Introducing gpt-4.1 in the api, April 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Introducing gpt-4.1 in the api, April 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:25.980401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:25.980401Z digest=sha256:2f45b49e16215d8ceda83620d342597036a747fe1c1f476734626b70caafbade

Observation 41e85d35-4d36-4afa-9448-d6f018c2cc8d · outbound

This paper cites Claude 3.7 sonnet and claude code, February 2025.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Claude 3.7 sonnet and claude code, February 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:22:27.046469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:22:26.052171Z digest=sha256:cd3595b81a245809eccc1a2f8a434f0f2121549a007e12bb040bb404e27e8c20

Observation 2d41297a-7b22-48f6-824a-25225da34672 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:26.127078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.127078Z digest=sha256:f803e7508f09c37b994c02e874b7e0c512d6d33f5ad61ffa770ba29ed54b55ac

Observation 56ffc97e-5c6e-4606-9adb-ba5a36adf4f6 · outbound

This paper cites DeepSeek-V3 Technical Report.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models DeepSeek-V3 Technical Report

Reference 46

Resolution
malformed identifier
no resolver link, observed 2026-08-07T12:22:26.217604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:26.217604Z digest=sha256:9d26586651f3dd003b5bb2071f02f6fc5cff60f1e78bbd5435cfa534f80b67f8

Observation 0e7d5fea-6140-40d0-9f66-44aae2f38c47 · outbound

This paper cites LLM-SR: Scientific Equation Discovery via Programming with Large Language Models.

PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models LLM-SR: Scientific Equation Discovery via Programming with Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:22.780339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:22.780339Z digest=sha256:62d214dfc2933cbbfc36a3502bf0ac368c1818a4075495178d84c9f8cfaf7442

Pith citing papers

No inbound Pith citation observations are available.