Pith. sign in

Paper Citation Record · LEDGER

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2608.03545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03545 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:52:37.405393Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ed278b8-71eb-4459-9c40-e8e4f961a4ae · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.363913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.125672Z digest=sha256:857acc086d65278e40ca49659565feb07d6d5e46c3ed449c56d23f89250c7fb3

Observation 35873716-33c1-4ee7-a5d7-687ed403528d · outbound

This paper cites , year = 1983, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1983, title =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.349715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.131388Z digest=sha256:142d2264a38d681f07a7610ae72124b724cf16abc167a2360990c8a1b5fe3562

Observation 185ea821-dcd0-4fa2-a70b-1c8587f17197 · outbound

This paper cites , year = 1984, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1984, title =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.335520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.138291Z digest=sha256:bc349a19355a94bedaab4ccafde201e83d3c166d9edf041aa74a30596a2a6432

Observation b3d29bff-8b19-4eab-85af-8753c59bde22 · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.143686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.143686Z digest=sha256:c04e75ed9b6411881980de8b711b154dde4c81ca557fd6e99ed320f9743c4b17

Observation 3608bf1c-1911-4653-84ac-93be67822f81 · outbound

This paper cites , year = 1980, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1980, title =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.311285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.149323Z digest=sha256:099e96dfb9c6f23d84908a568679c17cecd9dec02e92b0dbab6ade177bb5ebb3

Observation 7fe45bb2-cba3-4838-b941-ae90ab071f1d · outbound

This paper cites Clancey and Glenn Rennels , abstract =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Clancey and Glenn Rennels , abstract =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.154451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.154451Z digest=sha256:0a6b191f301fb64d1837fa3745b03da8be6cd38783ef2128178e40d1e5f11b0c

Observation 0ac35a06-d4bf-4e1d-9431-1f99db176d23 · outbound

This paper cites and Rennels, Glenn R.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning and Rennels, Glenn R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.296976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.160028Z digest=sha256:284f1f5539bfe7717fcb1d0a75e4af496a0e5c3be341bb1a2b33592341efa3e7

Observation cfc2ccef-ffbd-4879-bdb2-229b356bac7a · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-08T00:52:38.282460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.164832Z digest=sha256:e303bd358608f94348ce6f6d9eae0cb04c1d6d90c3db50bcbf99ab6167941603

Observation 6f41fcd9-91d7-4a90-85ab-5a1651d6b9e8 · outbound

This paper cites , year = 1979, title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , year = 1979, title =

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.268248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.169591Z digest=sha256:b05f1ca8920283c88b0fe47a65f3f3cf7e0b73b2d4e9c5281c7e29dcdc9ca63e

Observation 8423e12b-d49c-4ca6-8743-803e28c5f2ff · outbound

This paper cites , title =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning , title =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.253850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.174492Z digest=sha256:bb73b630ae1f28c1246cd975afc20d194272e4c9db50e0d8b1f2669cd72ad8bd

Observation fef0cc84-d6bb-489c-bc68-e029c50a9b71 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.179241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.179241Z digest=sha256:be6ee480dde974c882d258e941f3cb449916676b21bde975b528ecab598f753e

Observation f57f9d02-2e29-4399-a2a4-e07407b6d9df · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.183822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.183822Z digest=sha256:e7f74b22d93d3fb3db3077e00e04ea1d44b76db71d60d31f579d549952beb647

Observation 7230285f-38ad-475e-9636-945b7b7b6387 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Eleventh International Conference on Learning Representations , year =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.188675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.188675Z digest=sha256:9d10c2538524ff81f7253cb1861fdc8460286a89361fe8dbe724ff9169b39d5e

Observation a192e1c2-29e3-421f-ba77-5439c7648c1b · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.193820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.193820Z digest=sha256:db504516ccb41ac61242b63d19f0fbdd3be6b51ee67bb084b7ddec16d7060cea

Observation 28f925f4-298b-49af-bc80-cafc8694a046 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning TTRL: Test-Time Reinforcement Learning , url =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.202867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.199002Z digest=sha256:f4f07cd0e9a8d058a6db5f5b6080490865246bbffdf5bfb2dd392b55b0ccbc6b

Observation 7a9c20f4-41d6-4a78-8f6c-cc472feeaa1c · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.188738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.204174Z digest=sha256:df769a5ec6e01a6edd1eddc1d9d75e263606dc08bcf6e53d7da34eb0e153ac03

Observation 7f668440-b608-4435-8fa4-1ee9e9964a04 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.174641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.209028Z digest=sha256:4634082b39ded0a1e853a2bffc279e88e2213439ad40c08d3835f067725dcdff

Observation a27ffac8-fbe4-4c6c-930f-d5d2e403c0b3 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.159714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.214354Z digest=sha256:f59a85b065d10e1174732960ca0aafdf3d9185d053b972255fc0e47ee8a6a9a5

Observation adbaa938-1bcf-47fc-8526-8307834ead2d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.144454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.218963Z digest=sha256:05866e45e406492ac400d5465629bef180a26048d7361d41f3b7fba5acbfb7bf

Observation 20e4baa6-33a1-486a-bb8e-f7f86972d041 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.130677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.224796Z digest=sha256:7d3a4e9b69a920f21b0b7bd93f3fff5d502ad0ad88d61fcd919ddedf463a48d7

Observation 79e68d40-b81d-499d-968b-c94afc4010e3 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.116799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.229542Z digest=sha256:eb5bc77a0d2ed4f5f8019e5bf68414b822fb1ed2e8281fb8ae109c3c03158a07

Observation aad71683-7505-4632-bdae-ade1e461c9f4 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.103194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.234366Z digest=sha256:491f255b57efc5ba1afcae1fd713feeac8916889aa2caffea571e12212f73e54

Observation 1718470a-a437-4913-a8fe-0b4324b23d6a · outbound

This paper cites Forty-third International Conference on Machine Learning , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Forty-third International Conference on Machine Learning , year =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.089662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.239627Z digest=sha256:17640bb8414eb5883b979b4a158fc1731e3fb6f18afb757eda7c7c7ce4519a7f

Observation 3c107409-0448-475b-91d3-245a5a6e1c69 · outbound

This paper cites DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning , volume =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.244689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.244689Z digest=sha256:c8ddbbf77cc7e2eff81941f8620d348ec7298fffbadaa3cc4d54092aa7c319db

Observation 72850efe-d844-4d56-b9bf-396ea6340ca9 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.249652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.249652Z digest=sha256:3f3ab028b90ff335a1f8caf7a7d719a4ff5416e7fcb2f9fd497bb8ce061ad852

Observation 7d739666-e02b-4c69-96ea-050634c320bc · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.254564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.254564Z digest=sha256:e550b286edf8ed8cb221667ac7139dc777fbf79f94ec2e2be356a9589b6cc6a9

Observation fb731896-e67e-4e71-bd9b-a1385c1ed121 · outbound

This paper cites ACM Computing Surveys , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning ACM Computing Surveys , volume =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.058301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.259146Z digest=sha256:da6270972705da3c39c6355af417218edb946a3337d49ed5349a09ab0bf7422a

Observation a2328cc5-d3fb-41e8-bc74-bead1257befe · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.043693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.264159Z digest=sha256:1b5a50704e5379c4daefbe4a63be4b61b7c1ad10e1700e058708f526276cb6aa

Observation b5eb781d-193a-4253-a3b9-c151273900a7 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.029749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.268766Z digest=sha256:54e5d1f3436b8799c77e366a63da0eeaad1bc0305295ea6158a47c4940428f59

Observation 11539c81-0316-421a-b678-9166245911af · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.015451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.273663Z digest=sha256:34d4fc11b2755895a0450b410180a0f61c6a5d03e61b48fbfef1af36376164d7

Observation f788e82d-6b0d-4c88-9236-6177cfd148ea · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization , url =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:38.000775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.278216Z digest=sha256:3aecf158376e90e56ac9301b4f96161d10cb19ea9d82d5a2e1640f2c6a89b84d

Observation 813a7dde-7224-4df4-b079-24cc74824a68 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.985176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.282965Z digest=sha256:a05eaf23991b6e0214013354c98aa2cf06af044f5ee5cf738637a56f4610eb5a

Observation 19522351-5d6b-4f5a-8ea7-4483c3b21224 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.970491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.287476Z digest=sha256:33200029511776fd488a8372cba7b69f4eead4df4fc7143b1f11d9f71812bade

Observation d29273ac-bbb1-4f0f-a13b-caf93ea42e2f · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , number =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the AAAI Conference on Artificial Intelligence , number =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.955910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.292516Z digest=sha256:ec81b6f8e53cf73a7f4ea4bfdd59692fdcdb0144966932b7720501b866b29372

Observation 613632e7-a01d-4dda-a218-95bcc6f1caae · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.940948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.297300Z digest=sha256:f635c4ad108209f2fcc75deccccf5ab80cb48adb4ced0b1711e905b45786fa2d

Observation f8802d1c-dc37-4058-8c6c-ac8382597cb9 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.926536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.301938Z digest=sha256:36b5a9148bbf9b5ca08b8d8daedff6efb844325e36f4ff8f6115f8da497b27a1

Observation d9c7f3dc-d4a5-49c1-91cc-cad23708d9e2 · outbound

This paper cites 2017 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2017 , eprint =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.306383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.306383Z digest=sha256:3d0bd43dea7ba9d9c5b9d5dfd1941a714ccab96bf19e833c988e94e922b89f80

Observation ae6db8fd-e887-4b58-822f-50e922298525 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.311151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.311151Z digest=sha256:11b2810aecaae0fc6e1a0bcba24bb3ae8f2ac6b3986132428a70c90ba91368b8

Observation 6f09fccf-7350-4f5d-b300-9bb198b0639c · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.894932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.315756Z digest=sha256:1158fff32b9bc1874a231b6a5f25b52c0cac250579f845782620c2794e12deae

Observation 7b77432a-8fd5-407a-99e6-c7719486bfa8 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.881012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.320316Z digest=sha256:9f1b6a02975daaf4da86b0f4d246d59dc1b19e87b49c7068aa6c81701b127f2f

Observation 685415de-be4c-495f-b9df-191984cf0c50 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.866652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.324949Z digest=sha256:a67152dd3fd10829d67d329b8854f7d477acfa0ca78998c4472967dc0c9c1402

Observation c9123113-4c72-4338-bad5-3c1160a8ae38 · outbound

This paper cites Advances in Neural Information Processing Systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in Neural Information Processing Systems , volume =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.852473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.330636Z digest=sha256:b78df26bef56cb28a7438aa9a20ecb5d76074857cfe8aa36c75046ccd858c8d2

Observation c83c0caf-7d57-436f-a0d7-282ad478b756 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.838805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.335195Z digest=sha256:b6bdbc891ddf527dbb2142ce63f96f286b33f1f08be0523d9883f0923064e722

Observation 206a6009-f473-414a-b89e-7b37689ea0f4 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Reinforcement Learning for Reasoning in Large Language Models with One Training Example , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.824996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.340347Z digest=sha256:4abf0cf321c628477b78b7be483fa3e9cd7636d50f22e4e9031a142796edccd3

Observation b5cfdda8-f7b2-4197-a10b-5932bf58afdc · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.811006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.345104Z digest=sha256:c0be7cf7c6aba0a13d0025a0bf835f83ad730ccb939c7f2bf4175d4e10eb4337

Observation 9f72fd16-6654-4ccb-8c74-5f5c6a2c1814 · outbound

This paper cites 2026 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2026 , eprint =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.796660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.350015Z digest=sha256:69e43f48abd75f2a540575831ea98f6447f005824967bb458c0f3502ae79095b

Observation c4f1566f-16b2-421e-82dd-9063565dd468 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.781622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.354797Z digest=sha256:5ac3853688d007958676e5ecdbb8204052f02fba1d533e3aa2bb95701a42fe3b

Observation 4c46b567-af8e-4ca3-9634-9754c596e643 · outbound

This paper cites 2024 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2024 , eprint =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.767694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.359492Z digest=sha256:e4447cb6acf94154fed2008805f7b0f47bb0e4f596103b9620976ab3ace49b38

Observation 1579bc71-7677-4a15-aa62-b16ae9304b4d · outbound

This paper cites 2021 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2021 , eprint =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.753867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.363908Z digest=sha256:f522e0826628e3e73bd76e11d06540d2b43a508ff6c40d0060f1a93d58ded658

Observation 19720dca-b1e2-48e0-9c2a-1b824bad68a0 · outbound

This paper cites Advances in neural information processing systems , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Advances in neural information processing systems , volume =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.739927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.368739Z digest=sha256:4fde9dc6f37167914acd0d8e13f2979df5bf3a7ef657ee6f5bf0746c2a5b530b

Observation 7e02f688-fc05-453e-81b9-93c2e71db1c8 · outbound

This paper cites Hugging Face repository , volume =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Hugging Face repository , volume =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.725796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.373259Z digest=sha256:14c821aff965dce41316b2e2a265438e4b2ffa1876b9d015430a68b6eba82fe3

Observation 4988a4e9-2759-4038-844c-33085495760b · outbound

This paper cites an unresolved cited work.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.377789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.377789Z digest=sha256:e3d06a200e5db164be5ae07db182e8c32c3f30380d052cad05aa7f6c6b5fae88

Observation 06de54c9-ba28-4e50-8f3b-0211714439d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2024 , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Findings of the Association for Computational Linguistics: ACL 2024 , pages =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.701090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.382398Z digest=sha256:33355dfc3f204b259b980e8049434ff226a01fcc778f03dac888c19ce15b3245

Observation cbb5181a-3311-4f3d-8d9e-152a201fe851 · outbound

This paper cites 2025 , eprint =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning 2025 , eprint =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.686100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.387143Z digest=sha256:04b1ac76811d2aea623024630cd08d7706acd0f9791e7e05e57e04df6056c16c

Observation 88758300-b891-42c6-ac45-1711f3e6a1f0 · outbound

This paper cites The Hidden Link Between.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Hidden Link Between

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.391538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.391538Z digest=sha256:38236f68ba15f9ddbccb69535118aec07c50cb90b4485b688f24f97f86992caf

Observation 306e9743-db19-4fe3-9559-470fa2c9c306 · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning The Fourteenth International Conference on Learning Representations , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.671530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.395884Z digest=sha256:3999f3a4f9a48937d9eab9c27e27f9deaa1389cd19e2512702eae16d29eb2c66

Observation 35b06d4f-4564-49d9-83a5-ab510ec2c64d · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T00:52:37.656364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-08T00:52:37.400464Z digest=sha256:b735b6232b55ef88514c0a57a0d283eb7e5cc3085cf745c5736a027912c8ff32

Observation 7fe7ba90-5138-4f16-8f00-5575d2f34a96 · outbound

This paper cites A Survey on Human Preference Learning for Large Language Models.

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning A Survey on Human Preference Learning for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:37.405393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:52:37.405393Z digest=sha256:aec2ca570ac01b53bd9a43efdffe938ccae910cfc8619fc11afb5d9271b23b48

Pith citing papers

No inbound Pith citation observations are available.