Pith. sign in

Paper Citation Record · LEDGER

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2608.07371.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07371 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:32:47.290687Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43af2e35-8b6b-4a77-b853-75fb78390afb · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Advances in Neural Information Processing Systems , year =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.153181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.153181Z digest=sha256:b8a3393a2006e01f17766592d0f0d19266c12c6939279c605770aa39e71bfc15

Observation d882b2c6-7e5b-40d3-a6bf-9fbb63ddb08f · outbound

This paper cites Robotics: Science and Systems XIV , year =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Robotics: Science and Systems XIV , year =

Reference 2

Resolution
verified exact
doi, observed 2026-08-15T14:32:47.347370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.158990Z digest=sha256:07426794e4a8866c8de4187438e917aa5da2c8f0f40adeca11be0c1d2452ebd4

Observation 45955c89-a13a-4101-a2e5-fb155d762ac8 · outbound

This paper cites 2024 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.162667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.162667Z digest=sha256:e91667d24ead5551e06b7fb11a2dd9e0b814bab73771d0077366b507ce0b5808

Observation b1df07b0-d27a-45fe-b124-7cee04b408f8 · outbound

This paper cites 2024 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.165985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.165985Z digest=sha256:fe9fd19ca41c06a1330d01e63a6c66eaf2421fc543ffdf1c663d63760560905d

Observation 40cef15e-f13c-4663-947d-2283b7be1d0a · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.169375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.169375Z digest=sha256:9337ddba6c2520093ba48371a19d6db925a3f0654a82852523d6bcf0e38a9455

Observation 590e8453-8ced-4a3d-9d75-1bbe6cf84e16 · outbound

This paper cites Group-in-Group Policy Optimization for.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Group-in-Group Policy Optimization for

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.810510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.173342Z digest=sha256:62bfd2ceca114ab1def264549eab60682a4016257d36767bdc15ea734e4de3d6

Observation 892b1006-3112-42c9-a7be-5675cb4b0bee · outbound

This paper cites 2026 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.176697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.176697Z digest=sha256:f5420c5627482fb73610db12c7998efcb72e8f66bd90c3a7fe0024aef9bd31bd

Observation 598177fa-89c7-4cda-9663-6546f2e75202 · outbound

This paper cites 2026 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.179849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.179849Z digest=sha256:4cfc854d5b025bbcb6959bdb1ac93311672a2811e30569726ba096fedb025834

Observation 2004b62c-5367-4e00-8db3-131d370d0c8d · outbound

This paper cites Self-Distilled RLVR.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Self-Distilled RLVR

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.183433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.183433Z digest=sha256:dc49b6e7a5b013110bbe4d6c31593402a13c1777607b803f189e262bcd183f0c

Observation 4c22479f-2c14-4c62-bf45-b6947037748e · outbound

This paper cites Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.188247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.188247Z digest=sha256:85d56e6617f9867c10ff015415fe9218f50b68203548990ff68bad34c2d43182

Observation 6a497627-3db7-4365-9863-3381caafb392 · outbound

This paper cites TIP: Token Importance in On-Policy Distillation.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TIP: Token Importance in On-Policy Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.191944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.191944Z digest=sha256:0490e5eb85fc33ce9598824cb01897a735938a26416da256b37ef87b8edffd1e

Observation 40697789-9769-4ddc-8676-b87b8aadbe91 · outbound

This paper cites TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.195468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.195468Z digest=sha256:88641440261cf4e83e1bca64b9e7e36054bc737683480430e5fc1f5617c40bd6

Observation 1457809d-0d90-4605-a119-62cdfcfb6904 · outbound

This paper cites 2026 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.199425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.199425Z digest=sha256:7fcd08cfbfa864f617fb48aaecd8c8201ee8d7727dd459b40d18a3e9c1590a5e

Observation 49d7e6e7-d96b-499d-aff1-cbe99ab49061 · outbound

This paper cites MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:32:47.645907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.202666Z digest=sha256:f6c78a96da176bf53922d7fe8bf86335463679490668925014593bef503c0472

Observation 7e4e1bc5-677a-44e2-8e03-d79747a8f281 · outbound

This paper cites 2026 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.206077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.206077Z digest=sha256:a29c4ab4d0b49922539d4fbcd24c18f1e93666946ce0216919498d31ef038a06

Observation cdfa1f23-c784-4d8c-b20b-647e48fbc0f3 · outbound

This paper cites StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.209479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.209479Z digest=sha256:be09434498257cc39531f6e633f45b2ad58c0d65b8b5aca99c63d6f5785443ad

Observation e7841065-f830-445d-b8ff-e7406788012d · outbound

This paper cites HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.213039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.213039Z digest=sha256:ee702fe91093750aa866762fe3fe4ee713d89830d9e3b7d8b9f661f7345b587f

Observation e6576559-7c75-4aab-835c-1b2429f6e0a4 · outbound

This paper cites GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.216744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.216744Z digest=sha256:a54a57a0bc1346f2b932421cb6d6d9a5a03d070f8ba00e0d498fe8367fa7bbb2

Observation 32442963-7db3-4cc0-a4f2-71475cb1dfb0 · outbound

This paper cites OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.220446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.220446Z digest=sha256:b59d774fbf158f606e0487beb343ec6f559a71237ffe6151961a36a811cdcf74

Observation e93cda6d-58b7-491a-956b-429b01ad9891 · outbound

This paper cites PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.223649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.223649Z digest=sha256:1ef542b40b6ff33616cb005b8c95421145393a5c49ab41752172162251c59210

Observation bdbfe7a7-1e4e-4aed-af23-6a706b3a777e · outbound

This paper cites 2603.18683 , archivePrefix =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2603.18683 , archivePrefix =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.227342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.227342Z digest=sha256:fee219c4be54a878572a9e4d37799bb3694b6f6ed450827e1875d81829368fc5

Observation 9bdf08b6-28bb-43c5-af36-bf921743eded · outbound

This paper cites 2026 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2026 , eprint =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.774080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.230480Z digest=sha256:a8dfd83f72f6ed5d0a075adec5ad2ef1f150f3105fb61d7933088a3db43f66ba

Observation d65bb8c2-9c52-485c-97f6-ef1902e6c21d · outbound

This paper cites HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.233941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.233941Z digest=sha256:c73a2456b7bdb8751aefedaa55907b9d70a81818144f5f2d8d69d5e1e8321f82

Observation 77e4ed2d-77f4-48d5-801d-8fdba5c16bac · outbound

This paper cites CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.237322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.237322Z digest=sha256:96a2274e85f865d430d1f14095be694516b7e00586fa55545a362491e89059a8

Observation 5823c01d-e450-4f7d-936a-f8d519616ee6 · outbound

This paper cites ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:32:47.336865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.240840Z digest=sha256:e4ee8d4bdf39bb4beed882c9e5ef8b1a78e7e01e2bb03a300fc3eaf06626bdf7

Observation e3b83a2f-7ea5-4647-958d-16182753004f · outbound

This paper cites SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.244217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.244217Z digest=sha256:e2727e99a1d3579f520aabc55c03dce10b9626c5924afee5333f01845c18f815

Observation d74bd473-f4b2-426f-8612-6caf4e99b8ba · outbound

This paper cites TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.247975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.247975Z digest=sha256:80344de79849628c380244e698de043145b2fdfbd610f7b05d14f086fba6d446

Observation fdd59071-fe34-4b74-b6ba-4dc6288b7777 · outbound

This paper cites 2021 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2021 , eprint =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.251571Z digest=sha256:1e45e7c92fe307489a8a6ab2925267c6f95ec3b24994f9de2cd31b3336fafa58

Observation 4debe24c-edf4-44a9-84e5-b13ee961a9eb · outbound

This paper cites 2022 , url =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2022 , url =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.754518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.255143Z digest=sha256:f6d7a9613cae200a1e6db968b3885005ec6efcd2bb160ffe8abe7425da063496

Observation 4f4d4130-eaeb-44d1-a9bc-c22894d1521b · outbound

This paper cites 2024 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2024 , eprint =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.744503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.258648Z digest=sha256:8f462d42e460ee29e2e58929ce8090e340e9df1dbb62dbb633d2458bfde0618c

Observation 0f9e17d9-cb83-41f2-884e-ca019fbb0777 · outbound

This paper cites 2025 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2025 , eprint =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.262158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.262158Z digest=sha256:e6d4a9368a18f8ebf4b6924f706c140b53336a8ffac2593ac76bae1a4f28f2c9

Observation fb21a203-6dc1-4dfc-8b53-05e5b062ea59 · outbound

This paper cites 2017 , eprint =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2017 , eprint =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.265583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.265583Z digest=sha256:4dc72862b7bed7363e94e97499314cd6242c43c19b6c1ecec5f14f94ed62794b

Observation a70b3bff-5819-4816-b7e2-c691bcaf9183 · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.268646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.268646Z digest=sha256:49324d67a0959ecbe62f1832ef4fc920214835136e7337dc87f430a75558fb55

Observation ae0882d1-496a-4562-8be6-b1a4dcaf22aa · outbound

This paper cites 2510.23603 , archivePrefix =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2510.23603 , archivePrefix =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.272797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.272797Z digest=sha256:ffa59a6b422d583dab56d89c06fa31cc161d4678e1441fef5b7a860a7373b487

Observation 3d8b50da-c695-4d3d-accb-dcc189b748ef · outbound

This paper cites InstructSAM: Segment Any Instance with Any Instructions.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning InstructSAM: Segment Any Instance with Any Instructions

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:32:47.387496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.275858Z digest=sha256:91e673a176c54dfb1d85e9ffcf9e3d354aed43a9a334b3ee5fb68bad8cca9130

Observation e6cf9bf9-8a6c-4b66-8880-d69a209ba3a9 · outbound

This paper cites HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:32:47.279314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:32:47.279314Z digest=sha256:004ce2b689f307cb2338b7994f90e638778e7a01b4fad905d52f64eb3ac8b8af

Observation fb9f8a0c-ff79-4004-9b1d-67eab96bef69 · outbound

This paper cites 2025 , url =.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning 2025 , url =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:32:47.721027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.282982Z digest=sha256:ccb1c31f92750111ea6abee39929b86ea0bfc0ad7366c33b9a44cee50795810f

Observation 185b5178-ac17-4822-a8ed-223a8c273ce1 · outbound

This paper cites VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:32:47.362059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.286487Z digest=sha256:df98cdcbe33c65880aaf16a9f34c17975c9c0286513f3942d379bdbba643b83d

Observation 85e9fcbd-71ac-4e3f-89fb-feff40384c9a · outbound

This paper cites an unresolved cited work.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:32:47.710479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T14:32:47.290687Z digest=sha256:26e0d037db7234aa4b9950999bdec97ba0104cac6694aff1c22d6117d21932e1

Pith citing papers

No inbound Pith citation observations are available.