Pith. sign in

Paper Citation Record · LEDGER

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2606.09348.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09348 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T14:37:28.660097Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:06.584525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:39:08.145196Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 271df1c1-b80c-4a90-8d7a-cb2c275ae08e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:0dd8f7472c7b4fde81b66e84e33f4afcde70cdaf9e504963ff3670ce1b001809

Observation b7838f1c-c8e0-472d-a45d-f95616640fa4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:270c939b9cf2aa9bfa68e5d0ccf8d9723107891c127151a43857a357a64a2e2d

Observation d41e2f70-1e2a-4bb4-982c-1fc5d55cb135 · outbound

This paper cites Tongyi DeepResearch Technical Report.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Tongyi DeepResearch Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:8c647e1d8396bc94dea337a8505dd10e86911c17f6341601f33572a08630edaa

Observation 8eaafdb3-a240-47d2-b17b-b3e3fc1c1559 · outbound

This paper cites DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:a97f85ba004dbbaaf47ca41e856cf80519ef9da60b54cb6b5ec5fbea0cde5bc3

Observation 99a86f73-10da-42c2-8c65-79136130593a · outbound

This paper cites OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:f1f6c3533c08ab788d2b05b8e29dc01a2a4c50238ad00ff608974f17336feb49

Observation a3dbaaa8-63a2-4dda-a5d8-d83ef8d110e2 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Kimi K2.5: Visual Agentic Intelligence

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:81392a1e0193f208a21df2b2ad0d3471ae95468805b082852bd778876c54e266

Observation 58275e48-cd67-4043-9df3-9981930d8b7a · outbound

This paper cites Mind DeepResearch Technical Report.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Mind DeepResearch Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:f8a64977a0072892a73a35d2819595936fb034cd8c4fb4f158f477c2b7598ad5

Observation 3bc5bcf0-576b-4e6e-9b8d-6c589443d14a · outbound

This paper cites Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240, 2025.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:7d66176ac86332b2af5cbb1c2e6abc20b84d6b286d13be0417573ddb340bb34b

Observation bcba510f-3d50-43de-90b2-ecabedac86b2 · outbound

This paper cites Treerpo: Tree relative policy optimization.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Treerpo: Tree relative policy optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:42ea175a4dd6720988879412b1e058224029ec356c78085ce7f41f7356f51c82

Observation 9783179d-9fc5-4e62-a5b6-1fb73bd651b9 · outbound

This paper cites On-policy distillation.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-policy distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:980dfd2f1f013daf458dd48df341ae2bebed394941e9fabff35467d1b8c038bd

Observation 56b852a0-4995-46d8-96a7-ab4d6300a29d · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:ec1a7b5761cfff77cfcc8e6402bd5d075fc819819f9ff6624bb9b75ef2254950

Observation 3f5ce916-b75e-4304-b226-4d294cf7e800 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:76e4f2649b0307c6f0b2c7c605c4a3020095360e5469b7175bcfcd03e521115b

Observation f2dfa813-82f0-481b-a68e-f9b11a85bd46 · outbound

This paper cites Self-Distilled RLVR.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-Distilled RLVR

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:0d5533b79d14828f2f1ea05a7f16d770d458ce88c3008cceb3f55d9f73901ab7

Observation 18c5bcec-c127-4fb6-85a2-f1f253c40672 · outbound

This paper cites Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic.arXiv preprint arXiv:2511.12159, 2025.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Criticsearch: Fine-grained credit assignment for search agents via a retrospective critic.arXiv preprint arXiv:2511.12159, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:6c9c0b1f0358c81d1e83e9ed3fedc620b8e5bedeea63de608ab3e0b3a6528ac9

Observation bb4595ac-9b8f-46c8-83b3-5f2d32aaf597 · outbound

This paper cites RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:48d2df274bd50082414c16b7e31653659382627605d59234f745d992c6c4128e

Observation 76d81e20-033a-4f34-8979-905514cf5d4a · outbound

This paper cites Reward Hacking in Rubric-Based Reinforcement Learning.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reward Hacking in Rubric-Based Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:55e58843a295efea8048fc62d91f5cf2cbb4759ab74c4973570ff2866fa3c81f

Observation d52464a9-a930-4b0c-8ca9-24415466a351 · outbound

This paper cites StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment StepSearch: Igniting LLMs Search Ability via Step-Wise Proximal Policy Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:d026c95464def65eace6da22185a6d9be2698f29d21d4e1bba9f5da380637b83

Observation 41f54242-efd5-4ed6-86d4-1e1e6898a2bf · outbound

This paper cites Reinforcing multi-turn reasoning in llm agents via turn- level reward design.arXiv preprint arXiv:2505.11821, 2025.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reinforcing multi-turn reasoning in llm agents via turn- level reward design.arXiv preprint arXiv:2505.11821, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:9c10ade99648e15a0ddd31cc7760df375b7c7eaac874a198aa9f64f402ac7300

Observation 9d89e011-167c-4caa-8076-607574450878 · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-policy distillation of language models: Learning from self-generated mistakes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:83a4a92999dea85d096959673bf9d8cab8a6aad37ceb1f232807f1fd974d76ad

Observation 50974593-1df5-4bcb-9e25-714769de15db · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:07e50fc57fc5927330389ca60fb196b6282622d54b0f004559ae10cabb4d2c34

Observation 80f0789b-3f89-47f5-beef-54f3a5c6a7bf · outbound

This paper cites Reinforcement Learning via Self-Distillation.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Reinforcement Learning via Self-Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:27db0208baa0a106aa12d1fd17d8bbee5ef01dbe956b80a754536e1c5b48fe77

Observation 4d30b6a0-8e11-456a-bc58-0883f2291ac4 · outbound

This paper cites On-Policy Context Distillation for Language Models.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment On-Policy Context Distillation for Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:c83cc6d2e289efa49bebc276136d6a34fcd91c070e6b69162fd3a806a6419357

Observation 450dffd7-2cd8-452e-9bbb-3b7aa4e0e614 · outbound

This paper cites Self-distillation enables continual learning.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Self-distillation enables continual learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:7842c00520006a3f23f039b6db70e1571e0ce0a8531b3f5ddb279b1697c1943d

Observation 45ea532e-4e9c-45a2-a9f5-3c661072cda0 · outbound

This paper cites Privileged information distillation for language models.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Privileged information distillation for language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:722a354d3ed515dfbc71f993aba310a519d0c56a25d636a68a713e957b30d3ec

Observation 2d906c2a-4658-45fd-99ba-3880d5d2aa2d · outbound

This paper cites CRISP: Compressed Reasoning via Iterative Self-Policy Distillation.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment CRISP: Compressed Reasoning via Iterative Self-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:12b527fcccf8dc1d64341b7a25b6140c8a36bfdfc9b4f68f0d197a04e90bcdb6

Observation fcd0fced-f996-43da-a381-533a4679ee28 · outbound

This paper cites Mirothinker-1.7 & h1: Towards heavy-duty research agents via verification.arXiv preprint arXiv:2603.15726, 2026.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Mirothinker-1.7 & h1: Towards heavy-duty research agents via verification.arXiv preprint arXiv:2603.15726, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:c772a800838d6d08c966640ee95ee9d3e243c37b4cf79ca27017f68b07cf9f1f

Observation f6fd1f05-858b-4506-967e-3d4fcd0c2ded · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:a938836fbb7264439df45abeb3ea52e36bc701adf5ce11804af8e36f7e5e900e

Observation ea4bd47a-ef63-47f5-abe3-b9d1912d3f56 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:e5a9ed7e9e46b9a821cb1e1985ffe821c7d9c4e55590265d5ff7e86b57bb80a2

Observation 2ad92bf4-9342-408e-bb14-2e64635209c3 · outbound

This paper cites Deep- researcher: Scaling deep research via reinforcement learning in real-world environments.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Deep- researcher: Scaling deep research via reinforcement learning in real-world environments

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:50a1e8a426dd894177d615e617a55814c76ab7695c03a13b2c64635927b8bc39

Observation e37dbe36-d3a8-4805-b5b3-099961d9c22c · outbound

This paper cites Stabilizing moe reinforcement learning by aligning training and inference routers.arXiv e-prints, pages arXiv–2510, 2025.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Stabilizing moe reinforcement learning by aligning training and inference routers.arXiv e-prints, pages arXiv–2510, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:7662db8af1fcd7b9a1c5677efe849b671ef46fa3250a119eab51cb1f9684dc0d

Observation f70d8dba-4c17-4acb-a0ff-3d1b11f825a1 · outbound

This paper cites Openseeker: Democratizing frontier search agents by fully open-sourcing training data.arXiv preprint arXiv:2603.15594, 2026.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Openseeker: Democratizing frontier search agents by fully open-sourcing training data.arXiv preprint arXiv:2603.15594, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:77d18c467b3034535a9d2b621cc30407a4507116a8fe2eceb107d74cb6ee4886

Observation 797b6d42-dc69-4b74-8c2c-8de9f105a1d6 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:00e14911aff489cadcd22c842925824501d1299a3aec60a1022de19149b90aaa

Observation bfae2e42-f81b-471e-862e-5d10667694f3 · outbound

This paper cites Browsecomp-zh: Benchmarking web browsing ability of large language models in chinese.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Browsecomp-zh: Benchmarking web browsing ability of large language models in chinese

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:0e60ced6ca893c37deb1697a6e4efb1cc3cd349e59e6f8957443a22556dd56d8

Observation 919d60fb-cd25-4d14-90cf-075ace980846 · outbound

This paper cites Gaia: a benchmark for general ai assistants.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Gaia: a benchmark for general ai assistants

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:113d6d844a62e7b5bc4b022af56343fa68496bf41e0fd487f14ce96c1abf15cc

Observation 154f2951-df48-4eeb-9675-884235253d8d · outbound

This paper cites xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:a13325a19e0ad3fbe7fde99e02f253fe9fadcbad30af2281ec08af0d27a18949

Observation c575b97d-a4a9-459a-b111-0819f8277cba · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment gpt-oss-120b & gpt-oss-20b Model Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:e9ce68bbc74da5270da189e0996e262e844b6d5d28afa3b8745d4a3842b10388

Observation 4f27c5fc-cbe5-48ff-b64e-ab4e7b4beb19 · outbound

This paper cites Qwen3 Technical Report.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:1fc9af845c339928746481d59dc11d3da33dc3039fc5d9dd1b10e27ad61863cb

Observation 8e5b875d-1eb9-4d53-b3e7-1a80f05fc2bb · outbound

This paper cites Llamafactory: Unified efficient fine- tuning of 100+ language models.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Llamafactory: Unified efficient fine- tuning of 100+ language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:c1efd1344bb9a878139e42a08e07ad1a1220244badfb881dd079a885ae4d5034

Observation d9bfc50b-5b94-4024-9a0f-5603a5456c53 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:2b7722863fbc20dfe5d176a9cbd2f8b2e4f5b33aebe8972561601a2928d9f526

Observation 937f6d50-f92b-4c9a-8b42-ef06584cd3f2 · outbound

This paper cites param1":.

PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment param1":

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T14:37:28.660097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:37:28.660097Z digest=sha256:ccd7d5c1b0dfe6e845a4caba01e8f6c0d264769a3a8fc88d78cd37922f543b68

Pith citing papers

Observation 78c16054-15cb-4cd2-b031-3895e2104ec9 · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.703451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.703451Z digest=sha256:4b1614c2897d6aeb5cabfc57cb5b939d19b8e3d290c16ca450932b0bd1f8fab4

Observation 38de8315-b88a-46a4-82dc-19a2ec9350d9 · inbound

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation cites this paper.

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T16:39:08.247916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T16:39:06.584525Z digest=sha256:eceef3b54b3779863f083bd2198a788d40e82651ae7c30d5c9e4e1209e8fa5c7