Pith. sign in

Paper Citation Record · LEDGER

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

As of 4 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2606.25556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.25556 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:04:09.237686Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact11
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a852350c-56eb-4c8f-a697-f7d81fee59d2 · outbound

This paper cites 3SPO: State-Score-Supervised Policy Optimization for LLM Agents.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents 3SPO: State-Score-Supervised Policy Optimization for LLM Agents

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.692706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:77e7a2060fa871e84f6678e60fb182974b38dfa8649f6155cbc6ab00738294f7

Observation 78dd07e1-454d-43d5-a82c-1294cc757c56 · outbound

This paper cites Hierarchy-of-groups policy optimization for long-horizon agentic tasks.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Hierarchy-of-groups policy optimization for long-horizon agentic tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.710290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:34a3d76a779d2bf76c17aaa37fc4906d58d0f1657c45f360a8fa63a4b68b04d0

Observation d11bf8e8-69d3-4f59-9f74-c090727beb5f · outbound

This paper cites Counterfactual Credit Policy Optimization for Multi-Agent Collaboration.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:40:07.682048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:9c9cf747dbc1ee89e7b71337beadef96a341da0cb7e86518bf718b801ff9d3bf

Observation 3836ffba-fb46-4ec7-9a06-ecfcd9361482 · outbound

This paper cites EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.676769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:6f4993465eea2df81e73dea626e297309de6d03cc79227907514a603e8cf31ca

Observation 8512bf8a-c8f5-4cb9-b9a4-09104d992594 · outbound

This paper cites ADaPT: As-needed decomposition and planning with language models.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents ADaPT: As-needed decomposition and planning with language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:e9806adbf186ea75c64f86b53821fb74779dc573e8a0098da7d9e209b3d463a8

Observation 38c871c4-6484-4b15-b2c0-f7ec0c06baef · outbound

This paper cites Proximal Policy Optimization Algorithms.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Proximal Policy Optimization Algorithms

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.679210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:9d0b418ecb7a7c93141a79ab7a60d25ab7ed5ada1c54cf97deee835ffd0198ea

Observation ed63e954-a7aa-44b3-bdad-2749b8b4bac1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.684520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:fd56e549351d2bcec8a1da350b7fceba72731f551d384e51db6cddd0394e9f06

Observation 2d4e19af-aff9-440c-b889-168c4c22c501 · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.712600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:7428676f71f2e3472467f0eed392e93992fd2855ab84d7c20b25b869fb8c2e39

Observation cc7153b7-09e3-4ffc-a0cb-6bf605b8d134 · outbound

This paper cites arXiv preprint arXiv:2603.08754 , year=.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents arXiv preprint arXiv:2603.08754 , year=

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.707019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:079926915251abf906e83b9fc74bc75d5ca8fe8d4a49224574e328c3580f35be

Observation 84eb6661-9c47-46ac-8675-9f5b7b1459d8 · outbound

This paper cites arXiv preprint arXiv:2505.22338 , year=.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents arXiv preprint arXiv:2505.22338 , year=

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.703613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:9000508c1ef3b3402c96a97bafd666cb3c36616d9be357390299a0f5a177765a

Observation 42f73ad9-1957-40da-a546-5dc98e4bd52c · outbound

This paper cites Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.700720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:80a45346fffe102c80e9e44f0e6aa12e56ee29d3cd19d38a3b9110240dafcd7f

Observation 4adb449d-0218-46a0-a110-4838e5e871f2 · outbound

This paper cites Qwen2.5 Technical Report.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Qwen2.5 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:40:07.695197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:6a34f2297d0bda75cf0e90924df0b18a89860353f3f52d7f9c29cf3c806092ae

Observation 352d998d-fff4-42ed-a36d-3956d1af82a9 · outbound

This paper cites Learning Invariant Representations for Reinforcement Learning without Reconstruction.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Learning Invariant Representations for Reinforcement Learning without Reconstruction

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:40:07.693168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:0582140fbfc4535d328d1f2368bf63728179fc05400130a4f0052cb85c658ea2

Observation 89dc4a3d-182d-438f-a3b8-d4fe67dfe03c · outbound

This paper cites an unresolved cited work.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:2f8bc6baa43f4502def7781cc2ddc1760d161a12a7f5372668090000670f6943

Observation 1f81d9b8-fd43-4ade-9871-aa0d40c06af6 · outbound

This paper cites For each promptp we sample G trajectories τ (1),.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents For each promptp we sample G trajectories τ (1),

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:41922762e1527bb84fc2accfec3c1bcf17abac62d94a199d21d85121b5808bf4

Observation 16ec4b14-d23e-42c8-b276-fe726f82e4b5 · outbound

This paper cites see Table I.1.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents see Table I.1

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:9da2534243109a3201563b030e2b1aa7798f96b926eff0d513a967b26e91f8dd

Observation 9dc4a11c-36da-43a7-be1b-f26b391035e2 · outbound

This paper cites OnALFWorld,BiPACElowers the single- ton cluster fraction by 9 .3pp and increases mean group size by 1.6×.

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents OnALFWorld,BiPACElowers the single- ton cluster fraction by 9 .3pp and increases mean group size by 1.6×

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:4b3dbaf8b45c5193b586de59e7785b3754c7f7a2c72b9cc8108182aa2f577445

Observation cb74eda4-80b1-4b18-8e70-9a7bc3ceb053 · outbound

This paper cites Val/success-rate (binary aggregate, |V|=128) across three seeds: 93 .8%, 92 .2%, 94 .5%; mean ±std = 93.5±1.2% (reported as 93.5 in theAllcolumn of Table 2).

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Val/success-rate (binary aggregate, |V|=128) across three seeds: 93 .8%, 92 .2%, 94 .5%; mean ±std = 93.5±1.2% (reported as 93.5 in theAllcolumn of Table 2)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-25T21:04:09.237686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:04:09.237686Z digest=sha256:9445262e858fc5187fcd3b72c53478432edb2b6a94eb1f6e0f8a8c7dd437b39a

Pith citing papers

No inbound Pith citation observations are available.