Pith. sign in

Paper Citation Record · LEDGER

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.03545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03545 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:52:37.405393Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed278b8-71eb-4459-9c40-e8e4f961a4ae · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.363913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.125672Z digest=sha256:3c9ba30833d7c49889b5db8c89d1f504e03cd821f1b77675dba5a3472c89cc7d

Observation 35873716-33c1-4ee7-a5d7-687ed403528d · outbound

This paper cites , year = 1983, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1983, title =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.349715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.131388Z digest=sha256:b1b8af575b6809522b8126690e5d004a3f298e307d75e855b09047a978bccc51

Observation 185ea821-dcd0-4fa2-a70b-1c8587f17197 · outbound

This paper cites , year = 1984, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1984, title =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.335520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.138291Z digest=sha256:483d077952a9da8c73fbd7769f3964c0ed09511b2fa2f0f260ebd55e557df43a

Observation b3d29bff-8b19-4eab-85af-8753c59bde22 · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.143686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.143686Z digest=sha256:2ae1f0333c5cf9b4637da7d1f2ed86765b41f9e5a8792eb78f91ed4167a06c4e

Observation 3608bf1c-1911-4653-84ac-93be67822f81 · outbound

This paper cites , year = 1980, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1980, title =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.311285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.149323Z digest=sha256:6d0cda8fca82440a5f4b783b16f80f8b621f731052e36b5cc5d3033d0cdd9135

Observation 7fe45bb2-cba3-4838-b941-ae90ab071f1d · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Clancey and Glenn Rennels , abstract =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.154451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.154451Z digest=sha256:67e2a8fa8d18bb721cef6446ecafd8b410ff0dc42d196c6cd2841f7a9393635c

Observation 0ac35a06-d4bf-4e1d-9431-1f99db176d23 · outbound

This paper cites and Rennels, Glenn R.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning and Rennels, Glenn R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.296976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.160028Z digest=sha256:392468909a68b813472f02176238d311455b46923845dad1c03d932869f3acb1

Observation cfc2ccef-ffbd-4879-bdb2-229b356bac7a · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.282460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.164832Z digest=sha256:276c65d68e10ecb880140465bd9394a8e7bb438a34f5d9abc7d214c8b62850c5

Observation 6f41fcd9-91d7-4a90-85ab-5a1651d6b9e8 · outbound

This paper cites , year = 1979, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1979, title =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.268248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.169591Z digest=sha256:d0ddedad29c3919b0b0e9d90396535b62eefda5f6004b2009f8e70e1b8a1c25f

Observation 8423e12b-d49c-4ca6-8743-803e28c5f2ff · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.253850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.174492Z digest=sha256:86518f7af11c1b22697d442df180cc2bb6544d5292417ca73bf5dbfc168d5040

Observation fef0cc84-d6bb-489c-bc68-e029c50a9b71 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.179241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.179241Z digest=sha256:d1fa9a262172cd234f4ade680d605cbeb06c53520a389e6081fd4355e281e839

Observation f57f9d02-2e29-4399-a2a4-e07407b6d9df · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.183822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.183822Z digest=sha256:f14f70e12094ab55deb6f7f0733c544c2a669c970187dc0f6db667237a5f5152

Observation 7230285f-38ad-475e-9636-945b7b7b6387 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Eleventh International Conference on Learning Representations , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.188675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.188675Z digest=sha256:e8cf2fe528bd26bcea7eb4a873d914fe37e5bb22a73a2385f589795483fe8d96

Observation a192e1c2-29e3-421f-ba77-5439c7648c1b · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.193820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.193820Z digest=sha256:20f3774e9493ecba8e0cd93b089fc2adef357cd15a772724c856cad6fd23bb6b

Observation 28f925f4-298b-49af-bc80-cafc8694a046 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning , url =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.202867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.199002Z digest=sha256:fc1185e1b5e629ed89d1574b0d4374c409fdf1164cfff7c8d67202c3e1d422f3

Observation 7a9c20f4-41d6-4a78-8f6c-cc472feeaa1c · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.188738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.204174Z digest=sha256:116fba518020916d95c08f7d1151d739f72ca6d879cefd35761f12af8b7b4269

Observation 7f668440-b608-4435-8fa4-1ee9e9964a04 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.174641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.209028Z digest=sha256:e4dea360abfe17ef06d75e11ed6782d24f7f8925abe1c3f7669a772418fedbea

Observation a27ffac8-fbe4-4c6c-930f-d5d2e403c0b3 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.159714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.214354Z digest=sha256:9c67ec96cc939f730edb4f47434a625007d687792633438eced64d05c952a7fc

Observation adbaa938-1bcf-47fc-8526-8307834ead2d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.144454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.218963Z digest=sha256:87d1b93d7b818ac902d5c8cf880ebce16993ede99a4b903164216e469d46c61c

Observation 20e4baa6-33a1-486a-bb8e-f7f86972d041 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.130677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.224796Z digest=sha256:4dda415993ffb175a6aa87087fbd240d7c3ddfde4387c4aa4d9d1a35e00b43f3

Observation 79e68d40-b81d-499d-968b-c94afc4010e3 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.116799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.229542Z digest=sha256:dff1842e5e05eea3f4e9d026052d50c91882c938a58b3d81d7a716a0d6e306c7

Observation aad71683-7505-4632-bdae-ade1e461c9f4 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.103194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.234366Z digest=sha256:5613adecd4c958d09aab4d96fd4774330f5eb703de89854cdc36c07c48a0c20d

Observation 1718470a-a437-4913-a8fe-0b4324b23d6a · outbound

This paper cites Forty-third International Conference on Machine Learning , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Forty-third International Conference on Machine Learning , year =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.089662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.239627Z digest=sha256:c17fa2a275040dc5a0600e7439049a9edc757d85935ae5370ce960c34857164a

Observation 3c107409-0448-475b-91d3-245a5a6e1c69 · outbound

This paper cites DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.244689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.244689Z digest=sha256:ad739cdec973ae1c8ae231abcc7cff4af1b2d443491935225e0a1b14a2ea0895

Observation 72850efe-d844-4d56-b9bf-396ea6340ca9 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.249652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.249652Z digest=sha256:d9e79d62f4e8e6aeaa146dd3ceadb9c7abcd5e4ce5eef86f07dc4d61c69e77f9

Observation 7d739666-e02b-4c69-96ea-050634c320bc · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.254564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.254564Z digest=sha256:fec93805f66b49ea7d1fd54ba5341c59443a601c79b213c02f8908edbc6640b1

Observation fb731896-e67e-4e71-bd9b-a1385c1ed121 · outbound

This paper cites ACM Computing Surveys , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning ACM Computing Surveys , volume =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.058301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.259146Z digest=sha256:6745f2994d5e7d0524990c791af44293bf51da41970092412b7b0e936195f3fd

Observation a2328cc5-d3fb-41e8-bc74-bead1257befe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.043693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.264159Z digest=sha256:f6580c070689a2e28a8633513461fd6472245338b071830dfe89156aa6ee3c50

Observation b5eb781d-193a-4253-a3b9-c151273900a7 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.029749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.268766Z digest=sha256:75a257fe3e5997ebec7bce9898e27e48aba4d5cc28fff357e5332ee2ee6ec1c1

Observation 11539c81-0316-421a-b678-9166245911af · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.015451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.273663Z digest=sha256:48de78e8049fc59ef242df7b1cf95dbce76b90d756e3b55b61b000d2706746cb

Observation f788e82d-6b0d-4c88-9236-6177cfd148ea · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.000775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.278216Z digest=sha256:5960501e592770c1af87b3e733053edf8fa05365f776e3b886b44416f3cc823f

Observation 813a7dde-7224-4df4-b079-24cc74824a68 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.985176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.282965Z digest=sha256:9c04bb3910f54e1c6c9470a1159a990f3a9f6541c1b30a17f8703d2769743af4

Observation 19522351-5d6b-4f5a-8ea7-4483c3b21224 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.970491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.287476Z digest=sha256:82fc8affa7ae43df2d1484a6d278dc58b647bbfdecf26d416c12798b8d47532d

Observation d29273ac-bbb1-4f0f-a13b-caf93ea42e2f · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , number =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , number =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.955910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.292516Z digest=sha256:823f3933d9d8bc5996fd5cf3ce3ca4eb892883bf6c30428e6c86f1490f44b959

Observation 613632e7-a01d-4dda-a218-95bcc6f1caae · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.940948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.297300Z digest=sha256:3dd43e44765eb774711ea4d391899a9f1562c2078b02f3285135800075fa9605

Observation f8802d1c-dc37-4058-8c6c-ac8382597cb9 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.926536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.301938Z digest=sha256:c5d73cd0ca3847beb8a38a04d5eca0674993d60e2ed7c1bb290c55d23919cd83

Observation d9c7f3dc-d4a5-49c1-91cc-cad23708d9e2 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.306383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.306383Z digest=sha256:697a19312774a063a7d1a924218321edc109f78c4c4fcc0f903a1eef5c33fa53

Observation ae6db8fd-e887-4b58-822f-50e922298525 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.311151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.311151Z digest=sha256:de63c5acf64fa0cd9718182e27aa4a2160e08bc37f4d0c5290ffd42078369c01

Observation 6f09fccf-7350-4f5d-b300-9bb198b0639c · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.894932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.315756Z digest=sha256:47491f89add26dfe34358d2b4adae510475ed44425dc0f5b4e1f22642e9785dc

Observation 7b77432a-8fd5-407a-99e6-c7719486bfa8 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.881012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.320316Z digest=sha256:4a1e2b8fee0049b3d6692ee51903b8a13f248334b10ee5954efa59c8c7c919e1

Observation 685415de-be4c-495f-b9df-191984cf0c50 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.866652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.324949Z digest=sha256:fde9aebf3544e1aa5b818b7dfb14297fa8c2d980988cf6a7e174ac12405881a8

Observation c9123113-4c72-4338-bad5-3c1160a8ae38 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.852473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.330636Z digest=sha256:d760539bd05101b03e840f24d118924a4c28613188739603816bcf97919f2571

Observation c83c0caf-7d57-436f-a0d7-282ad478b756 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.838805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.335195Z digest=sha256:1f7b1cdec9e4aa2ad4e8f7fc74959ceb37c51ed42bb016e58e96fbb276e59333

Observation 206a6009-f473-414a-b89e-7b37689ea0f4 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.824996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.340347Z digest=sha256:0de899c9f2db7dc734508daef71b388affa288a8ee68175afffa77f31372f86f

Observation b5cfdda8-f7b2-4197-a10b-5932bf58afdc · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.811006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.345104Z digest=sha256:d071fc00e354c50c7f0b3f98f173c7f622a4addf4c26c78f9ebe825d280d4256

Observation 9f72fd16-6654-4ccb-8c74-5f5c6a2c1814 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.796660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.350015Z digest=sha256:8cf1e5842cf08da1d82fd971629b88344f522ee98adc5def39f68a59c209edba

Observation c4f1566f-16b2-421e-82dd-9063565dd468 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.781622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.354797Z digest=sha256:7793889d3e6643787eb587252aaebc120db8ee0092f3b8f6f305934705496ea7

Observation 4c46b567-af8e-4ca3-9634-9754c596e643 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.767694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.359492Z digest=sha256:22c3310e940a99915ef59eccfab0228dea308eff7c5706de71ee761b57139751

Observation 1579bc71-7677-4a15-aa62-b16ae9304b4d · outbound

This paper cites 2021 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2021 , eprint =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.363908Z digest=sha256:7513b3740aaa6db8b538052eb4bb87c65f3a532e3cc1fdfca9ad75c7ea9132ec

Observation 19720dca-b1e2-48e0-9c2a-1b824bad68a0 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.739927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.368739Z digest=sha256:24b52e89baa0f1f0f4ae4fe714fe21d12c43f09481cc1ccba81b26cc06c8b86d

Observation 7e02f688-fc05-453e-81b9-93c2e71db1c8 · outbound

This paper cites Hugging Face repository , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Hugging Face repository , volume =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.725796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.373259Z digest=sha256:0338e56e8f71956fb62e6db04f8cad711e1118e3fa8bc08849d329accaee6dab

Observation 4988a4e9-2759-4038-844c-33085495760b · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.377789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.377789Z digest=sha256:9cfd15eb9212e8ad8e8addd6b2cf6996a8817fc63b46124bdfee010afe2c024e

Observation 06de54c9-ba28-4e50-8f3b-0211714439d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Findings of the Association for Computational Linguistics: ACL 2024 , pages =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.701090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.382398Z digest=sha256:20ab2b0782ad43383f82776fcc5b531a8bc86fe81c926496e790b78a3fe8d921

Observation cbb5181a-3311-4f3d-8d9e-152a201fe851 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.686100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.387143Z digest=sha256:3afa569bf2097b4867d151ab7ddd0bccdfd09e7e71e9de1948eda45e9fe44f1d

Observation 88758300-b891-42c6-ac45-1711f3e6a1f0 · outbound

This paper cites The Hidden Link Between.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Hidden Link Between

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.391538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.391538Z digest=sha256:523c55161a8acfccf6ebbfaa22d1656c947af757e7d0ec7a566042c6d24aa9cb

Observation 306e9743-db19-4fe3-9559-470fa2c9c306 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Fourteenth International Conference on Learning Representations , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.671530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.395884Z digest=sha256:9eb08c2a89936411726e6a94116c6838b2c65b287463250de04b95f0fa0dd5dd

Observation 35b06d4f-4564-49d9-83a5-ab510ec2c64d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.656364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.400464Z digest=sha256:cfc4a9fcb7c8eb880dee05a334bdaf5c2101556dd9313ebba13957f1567f2fb8

Observation 7fe7ba90-5138-4f16-8f00-5575d2f34a96 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning A Survey on Human Preference Learning for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.405393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.405393Z digest=sha256:cd7387986e6bf58315194079e98898b8dd7fe34193815fb4f42ccf1d57ec7edc

Pith citing papers

No inbound Pith citation observations are available.