Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.07087.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07087 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:53:58.448516Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T08:40:40.910461Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.541945Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d385b-6226-4e31-a870-161a0705d6b8 · outbound

This paper cites write newline.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.343053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.343053Z digest=sha256:a6f5bc18e4b9caafc1f29fa23579811bd96e86475fd18e329baf990b844388dd

Observation e1a12ddb-7f55-45ba-ad4f-73451fe84156 · outbound

This paper cites Claude 3.5 Sonnet , 2024.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Claude 3.5 Sonnet , 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.774725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.349218Z digest=sha256:5ee87dc3a44f4a40377d4c76f10876d08221cb46dc34297f9c03a00703c2cf78

Observation d50af5b8-f711-45f8-98f6-893b9babe3ee · outbound

This paper cites GPT-4 Can't Reason.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring GPT-4 Can't Reason

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.354107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.354107Z digest=sha256:b748d884b9550ee925f267eae53e8625b7f41016c16f4f8127c72aa8038b3396

Observation 65fa7dd1-f9fb-4fb8-831e-9e143f8790c4 · outbound

This paper cites DeepSeek-R1 : Incentivizing reasoning capability in LLMs via reinforcement learning, 2025 a.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring DeepSeek-R1 : Incentivizing reasoning capability in LLMs via reinforcement learning, 2025 a

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.761963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.359270Z digest=sha256:b733c4b2c8af3433e8998941e25bd865135217f6a397eeb44ad86f8d91401d5b

Observation 9abc6e86-9f0a-46a1-a9ad-464abcfbd369 · outbound

This paper cites DeepSeek-R1 model card, 2025 b.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring DeepSeek-R1 model card, 2025 b

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.748806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.364348Z digest=sha256:9e1531ba0b3871e84c75617bc9f91885463b69b329ae902c902201308b35377e

Observation b2b7225e-4314-44fa-8c28-6ac86d1a5b4d · outbound

This paper cites L., Jiang, L., Lin, B.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring L., Jiang, L., Lin, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.735199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.368996Z digest=sha256:e3b28f503af67394a5f453a920b87a649dd43a096faef7f3f1e0d5a14c45a07f

Observation 4326786a-5b26-4ffd-a8d6-14dc41f2c514 · outbound

This paper cites Kimi k1.5 : Scaling reinforcement learning with LLMs , 2025.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Kimi k1.5 : Scaling reinforcement learning with LLMs , 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.721861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.373693Z digest=sha256:d21f93a195a0e291f5693ad1f13d313d46c3884abf500c54ddd4caaf55ea32a1

Observation 6fb7d72f-9a0e-45c9-86f4-ac70a80e85ce · outbound

This paper cites K., Dasgupta, I., Chan, S.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring K., Dasgupta, I., Chan, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.708350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.378515Z digest=sha256:edee5d727dd08ca81b59c8a3f5f0f2c6096318b292c0bdae52b5f3ce1ca4b0d8

Observation e6082ac6-0cec-4c51-8f52-da3b1e5774ec · outbound

This paper cites Introducing Llama 3.1 : Our most capable models to date, 2024.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Introducing Llama 3.1 : Our most capable models to date, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.694698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.382757Z digest=sha256:0e75ba2adb9b109d2aa8f5de1fc09a131c626c992f96ec58608308d3abe0b04b

Observation 48f5b795-eb6c-4f2c-90fe-82f8459f3a0c · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.387372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.387372Z digest=sha256:70e86d5cc4baf891a0ab853e3f5a9a96d0108e83ded5727f2f682a2322dcfb6b

Observation 1a570a10-0985-418f-b7c4-58685e3469b8 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring A Comprehensive Overview of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.392210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.392210Z digest=sha256:6ed79b7bd04d4b9f233efd18ae35b7afaa5fc1ed81299af5446bb037fc1d99b5

Observation e0b468e4-69b4-432f-b894-221578b91067 · outbound

This paper cites Hello GPT-4o , 2024 a.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Hello GPT-4o , 2024 a

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.679892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.397984Z digest=sha256:da45de716bd143d8c9edde89d5b53da90082435165c73ba5795ef8d3ca9b0080

Observation dd7cb98b-9f09-42b7-852a-0ce6e76eac11 · outbound

This paper cites Learning to reason with LLMs , 2024 b.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Learning to reason with LLMs , 2024 b

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.664477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.402501Z digest=sha256:71aded9fd2176667be7afbabedca383e7dd1e5a37278a4d5df23c4ef26f08ab8

Observation 9eace8b9-8428-49f0-a6a7-cacd9842b9de · outbound

This paper cites OpenAI o1-mini , 2024 c.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring OpenAI o1-mini , 2024 c

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.650040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.406833Z digest=sha256:3ece83b5abafb4e3ed58cd28d6a08b8c68dee9e8bae5d61f29d6fbf10e600116

Observation d099ac27-7d05-421c-83e9-4e7be89a6911 · outbound

This paper cites and Hassabis, D.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring and Hassabis, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.636196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.411341Z digest=sha256:36e73b59b2ca41bd872ff9e1876424528956ec895e24b70db488684595c6212c

Observation 434f714e-610f-4443-81eb-c96787744229 · outbound

This paper cites Measuring and narrowing the compositionality gap in language models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Measuring and narrowing the compositionality gap in language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.622448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.415729Z digest=sha256:de84c6a96c4a8219d7b182642700e36533963dad04b8c8e3189b2451a061c94f

Observation 308db8f5-f1ff-4f17-9dcf-044cde671c0e · outbound

This paper cites The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.420252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.420252Z digest=sha256:5bef60cd31ba79b7e54854f2064c2ee7670f4a800aa329d671f15158c60ed454

Observation 87e53046-c0a7-4f52-81dc-71d90f61234c · outbound

This paper cites H., Sch\" a rli, N., and Zhou, D.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring H., Sch\" a rli, N., and Zhou, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.609040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.425090Z digest=sha256:742afa84571cf0f401943c3f92f56d59a7505ceea3783b8811f1cce0184f389f

Observation 7981cc38-c4eb-4601-9ffc-71540aa28923 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.429611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.429611Z digest=sha256:6434f0b665157ab3715e7c4f9f03ba579d2fe92db4e4bf3ccea74acef6017e9a

Observation a252b1d1-2e84-4f6c-997d-bb3ee9f24002 · outbound

This paper cites LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring LLMs Still Can't Plan; Can LRMs? A Preliminary Evaluation of OpenAI's o1 on PlanBench

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.434511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.434511Z digest=sha256:a407323f5508ae792d165859166e8134891e9cea29c1c8dce8874a6505b871e6

Observation 50496835-9d32-4382-ad34-5be903cfd868 · outbound

This paper cites Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.439406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.439406Z digest=sha256:e57bd1398a5eaf15b5d78fd72b573d65865b48a8ec51682834476baaa3677a24

Observation c7cadd4a-96ec-44d4-bab7-8ac4d3c9a951 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Tree of thoughts: Deliberate problem solving with large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:53:58.444234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:53:58.444234Z digest=sha256:3624265e51580e7faaedb2ecc251533ddf3604a7816fe5c0ce0a97212657408e

Observation 0f15cfa4-2872-4e42-8026-5218c68e6246 · outbound

This paper cites Larger and more instructable language models become less reliable.

Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring Larger and more instructable language models become less reliable

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:53:58.586191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T13:53:58.448516Z digest=sha256:ac8f12caba3449647697ce51ce6f3d97aae597752e35152bd38d6540c14e70da

Pith citing papers

Observation 89aa5d30-c9ab-4169-ac72-b89355b454db · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring

Reference 259

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.545298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:279386a8aaa454c1614fdba0f8694fccca405d727cd5527788c3e44358c3906f