Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning in hyperbolic space for multi-step reasoning

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2507.16864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16864 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:24:27.883078Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T23:51:47.724033Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9f291b92-9ff4-42fc-892d-ecaf5592c824 · outbound

This paper cites Advancing Reasoning in Large Language Models: Promising Methods and Approaches.

Reinforcement Learning in hyperbolic space for multi-step reasoning Advancing Reasoning in Large Language Models: Promising Methods and Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.191374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.191374Z digest=sha256:dca5e6aa59708704ea573a92e056657d948ff295fe4559cd88219e1e4d94f334

Observation e791a1a2-9549-4505-a103-4353d0e036f4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.736509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:25.289751Z digest=sha256:6d9bc5cf6fbab93fec434a344d732d664259ff0ee4cf2b8b057d07f43e9e3dbc

Observation 413e08c6-4fc0-45ed-8c58-1983936e23ed · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.656103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:25.385382Z digest=sha256:56392f4ddc653ce98af3f43c32e186d151b9ce6684b0c0f1e339149240f1f4a6

Observation 03a384d6-6603-4d12-a5ca-d098dad036e3 · outbound

This paper cites Multi-step reinforcement learning: A unifying algorithm.

Reinforcement Learning in hyperbolic space for multi-step reasoning Multi-step reinforcement learning: A unifying algorithm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:24:30.474757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:25.600915Z digest=sha256:95dd50da8c9ae4909770ce12ec782788113e6c8120f05bbcdca7e284c0cd743c

Observation 6da6aa9e-b692-4407-a50d-263863717101 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Reinforcement Learning in hyperbolic space for multi-step reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.721846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.721846Z digest=sha256:c8238e2908f416e3fd2ff50cb92ad7d86bb0e41bf6d994545f96d16183cd2150

Observation aaffbf05-6329-479d-bf10-f10098639611 · outbound

This paper cites Process-Supervised Reinforcement Learning for Code Generation.

Reinforcement Learning in hyperbolic space for multi-step reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.819837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.819837Z digest=sha256:12d30339e6eaac91e43eccc00e972900a74b7d6f836ecde3c59b2a5e78b95f02

Observation 6c9a4a8c-d07e-4988-a050-86cb06ce2e57 · outbound

This paper cites System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI.

Reinforcement Learning in hyperbolic space for multi-step reasoning System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.874748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.874748Z digest=sha256:550811bba6a375a699c0f93f11b2f46db96345b6d9aa90b538543bf089de500b

Observation 47f297f6-ad03-4245-9655-2cd4e3090e17 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Reinforcement Learning in hyperbolic space for multi-step reasoning Mechanistic Interpretability for AI Safety -- A Review

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.984891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.984891Z digest=sha256:38068119c6141b8040a9c08260e2f74ff3b943592b67536a1eb7dd77392fbb43

Observation 0356ed62-c5dd-4f6b-91f6-952dee763cc6 · outbound

This paper cites Transformers in Reinforcement Learning: A Survey.

Reinforcement Learning in hyperbolic space for multi-step reasoning Transformers in Reinforcement Learning: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.050901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.050901Z digest=sha256:f4a04a942051683ca8772facc7c1b435eb879311887fdd4f8ebd4a3c2ffd6197

Observation 5700f723-f5e9-4840-b5d7-8340f3c30fef · outbound

This paper cites A Survey on Transformers in Reinforcement Learning.

Reinforcement Learning in hyperbolic space for multi-step reasoning A Survey on Transformers in Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.139099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.139099Z digest=sha256:3835962efee3ae824ad6837566e4eaf3377134bc59412aee2ac57090dd5a1740

Observation 3fbb0660-e539-4f30-93da-fcf734f3e99a · outbound

This paper cites Deep Transformer Q-Networks for Partially Observable Reinforcement Learning.

Reinforcement Learning in hyperbolic space for multi-step reasoning Deep Transformer Q-Networks for Partially Observable Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.214913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.214913Z digest=sha256:9a5c7b1446b7feaa6beb33bf5b057a2a3251a1b570ddae1cb6b8e3cf9671a524

Observation ae8e3e84-d7d0-42fe-886a-3e73e6e824f5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.322606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.312548Z digest=sha256:90521b00fedab68f99f7803a2f6d715c4c7418e0a9ec2b737bf0762b10cb0761

Observation 3febb1cf-f523-4cb9-8de8-7d08acdfd19c · outbound

This paper cites TransDreamer: Reinforcement Learning with Transformer World Models.

Reinforcement Learning in hyperbolic space for multi-step reasoning TransDreamer: Reinforcement Learning with Transformer World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.376531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.376531Z digest=sha256:7db385807eca33796aa35b009640ed1ce5c874211e2907be891100b3b66eecda

Observation 9f2d7c5d-0f5e-4bcf-94f6-8b4c9700e7ec · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.170128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.447389Z digest=sha256:d61cc37e478d15e49222adeeb09709d7b4763c90871f1a9236545a95a099d205

Observation 3a5f4cac-de24-471d-8db3-3581bef3df7d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.007428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.491830Z digest=sha256:14e3715b2555cba8e10f497dafb1e1cb177655bf61f863a3f1d6f6f244c73274

Observation a906c0a1-69e0-4378-930d-5c71a3242d7a · outbound

This paper cites Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism.

Reinforcement Learning in hyperbolic space for multi-step reasoning Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.585807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.585807Z digest=sha256:398ab8b1b0b4c6ea59926a7f1846c165fba2d34c71bdd8dc76ac97cba1e0a4c0

Observation b41e75c3-aeea-4c77-91fd-e1d67ff4f3eb · outbound

This paper cites Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization.

Reinforcement Learning in hyperbolic space for multi-step reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.666597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.666597Z digest=sha256:4fc5a2668691d9279c320877589e44039b1a9fecc073297e0c2c6de79b3baf0b

Observation de208ac0-a99d-4c4e-9d92-bd2b09d0416c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.852764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.793498Z digest=sha256:62170b13d936455ae336ecc7c390dccf5d734ee2b07f1cd3a04fc42f5cf35e0c

Observation f0e5f976-146d-40fe-9e6a-ced587c64e51 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.689007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.889102Z digest=sha256:37c743c51a0c84aaf1b7e85aa46e2353cc60e2dd0d88cde2ce325efcea11c77e

Observation efb3d764-7090-4e84-8a54-65a4ec6c0b1c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.547578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:26.986934Z digest=sha256:80cd64f86d2d6026b363bbd6eccb7930a128f12ee3d32080c797832eb49dde00

Observation 8cc50101-0e0b-45a8-81f1-4a2740de302c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.386796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.052694Z digest=sha256:3ce20c61d3240ad49048398c535e8d517c279552f842cd823489f102e93f6f35

Observation ee9628fd-abe5-4c13-92e9-51b449639072 · outbound

This paper cites Poincar\'e GloVe: Hyperbolic Word Embeddings.

Reinforcement Learning in hyperbolic space for multi-step reasoning Poincar\'e GloVe: Hyperbolic Word Embeddings

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.151997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.151997Z digest=sha256:b875a993dc7f1ba6dfb555661a6a53eb456b50ec701210ab3c41313ddea252d2

Observation 86128f63-7e99-4d89-b410-f9c31ad43f95 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.249154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.220722Z digest=sha256:9f34ec3adfd83730dac73ba1efe374135b7fd03999b6d3ba7356c28897228303

Observation 19b80823-8346-4964-95f5-fd171b1c823b · outbound

This paper cites TransMLA: Multi-Head Latent Attention Is All You Need.

Reinforcement Learning in hyperbolic space for multi-step reasoning TransMLA: Multi-Head Latent Attention Is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.310795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.310795Z digest=sha256:bb673e2129357bd9847a70f27748726afc27486b68c55c3f8474f73b8c49d7c6

Observation c910248f-3e95-48f7-b327-122eeb0e90bb · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.068005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.405059Z digest=sha256:6c8a7886fe47137adc06562dd3b9ee93ec2334dc3dafe82a51d8766ef4624368

Observation 0fd0795c-86f6-41e1-b087-1f98c1a8ac92 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.881615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.511247Z digest=sha256:289cb536133f77f5e4de62b677b5430d46918b9afc738c45cfcaaf406c12503c

Observation cf8a0260-50bd-43bf-a72c-03e63185ed2c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.715723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.637964Z digest=sha256:4255c4c700f8af0b03ca6aecbcaa838d179dbab38cf06b0797f9dd7dd77b3d7e

Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.704354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.704354Z digest=sha256:89713ad0728c1d9d5fee9001b2ad84581cd84846c6b6daf8e7aadaa1a182d268

Observation 3495c8fb-34ed-46ca-86e3-adfeb0792238 · outbound

This paper cites Transformers in Reinforcement Learning: A Survey.

Reinforcement Learning in hyperbolic space for multi-step reasoning Transformers in Reinforcement Learning: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.790125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.790125Z digest=sha256:9bd1471e9524ce86c6c849f5243165600b74db3c4773076e0054492dffbab803

Observation 5e63b8d5-7ed3-4a55-99b9-9c3cde6264ca · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.562546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:24:27.883078Z digest=sha256:35d3b3d6e9dd88dd359c462b24723b972a7b2d5a1c18300ea8abdd3f86ae9c10

Pith citing papers

Observation d5b84770-5e8f-45f9-bf6d-9cdc8a10a36d · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Reinforcement Learning in hyperbolic space for multi-step reasoning

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:54:45.604319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:90f2fb86d4a00c028b706d719795f12072d1d2bb1af3fd6a5d32d53acb36e838