Pith. sign in

Paper Citation Record · LEDGER

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2604.03562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.03562 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T13:05:01.450957Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8778e897-b65e-4c41-b1cc-e6563e6e9da0 · outbound

This paper cites SpaceX Starlink,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling SpaceX Starlink,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:bb62434fd1f35714bfab12a3f4be68456359d2edf27681ae744ac7c1770896e0

Observation 8e0ccf51-d8e2-4b09-b645-143595936ab4 · outbound

This paper cites Beam hopping for multi-beam GEO satellite communication systems,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Beam hopping for multi-beam GEO satellite communication systems,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:27049b56f48f39b1bc48f13513700bdea2fffad732cc8338a20c37ee828488ef

Observation 3d614cea-32e2-415c-b96c-0708011ce9b4 · outbound

This paper cites Deep reinforcement learning for dynamic spectrum access in satellite communications,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Deep reinforcement learning for dynamic spectrum access in satellite communications,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:846ee088401292f3615130f547785f276751e2089c22d34a13ddb1c514b8b6a8

Observation edea8cad-01d6-431c-92cb-3ed4f736261f · outbound

This paper cites Eureka: Human-level reward design via coding large language models,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Eureka: Human-level reward design via coding large language models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:ff2baae58e9b697d1870985965274b68a487faf8135dd3e1dd3bb2a9bc75c361

Observation e1cd6ae7-12d7-4c8e-95d1-32ab17174fc6 · outbound

This paper cites Deep reinforcement learning for resource management in network slicing,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Deep reinforcement learning for resource management in network slicing,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:8bd019a42921a87549fb02e20c5c7e2b5d9596898c9ffda1b91c6e9feb6be558

Observation da16fb74-2dff-4773-97f6-73fe454f95d3 · outbound

This paper cites Deep Reinforcement Learning Architecture for Continuous Power Allocation in High Throughput Satellites.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Deep Reinforcement Learning Architecture for Continuous Power Allocation in High Throughput Satellites

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:bb48a6be3f1d03b9925a386659b8eee3e6780753ad22c91a9c7a5cc3621d33eb

Observation 22c11e2f-ecfd-439a-ad3b-b358a0bfd956 · outbound

This paper cites Multi-objective optimization for cognitive satellite communications using deep reinforcement learning,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Multi-objective optimization for cognitive satellite communications using deep reinforcement learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:3c72efa21f86624a1ec7e017a41c2cce5cd4d03ed27486468467ff5a42a9ae58

Observation 36dd3ab4-dd4a-4312-b72c-94a7dd0c7651 · outbound

This paper cites Deep reinforcement learning for satellite communication: A survey,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Deep reinforcement learning for satellite communication: A survey,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:6bd8c0f28e658972dc86dfddb59daff52d2b756c213b7e61931d0f214d5d40ec

Observation fdef6d9b-0790-4b2d-8b4c-29a80c808a92 · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:a6069648a4d133a6a48438dfef9ed5a2b5ced146b0b43b2d5f76ebb5945d4b96

Observation a02ac5e8-931f-4039-b021-ca2e421b24d6 · outbound

This paper cites A practical guide to multi-objective rein- forcement learning and planning,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling A practical guide to multi-objective rein- forcement learning and planning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:3003ca2ad753c17f937de912e3d2b65800e827007f5279ffd2bf050bb97c2199

Observation 31a6e1dc-8a81-4995-86bd-218aeb620d77 · outbound

This paper cites What Can Learned Intrinsic Rewards Capture?.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling What Can Learned Intrinsic Rewards Capture?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:46ab48afd47bb9f461ee69cdf046b15a605412d3bb9b37fb8d9a2dd46b4cbef8

Observation 3735d2af-d4b2-48a3-aac5-eacecfc2de4e · outbound

This paper cites Large language models for telecom: Opportunities and challenges,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Large language models for telecom: Opportunities and challenges,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:6aa30e2883968432398ca755696475bf06942da5917fdc019ce7a8c4d94e3df0

Observation da95fcba-51c7-493d-955b-588eb84223bc · outbound

This paper cites Networking with large language models,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Networking with large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:b5d877fcdac2cd43fe63518f6edf394eea389d2de24a3d766f85317a91d93d6d

Observation 392c7006-c0b6-4315-92cd-d93ffa57cc57 · outbound

This paper cites Large Language Models for Telecom: Forthcoming Impact on the Industry.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Large Language Models for Telecom: Forthcoming Impact on the Industry

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:30ea554083317071968ad94f291631786d4961890540d682325e38a27d01fe0a

Observation 8e2ecb67-045b-4c29-b263-50b17bc91190 · outbound

This paper cites WirelessLLM: Empowering Large Language Models Towards Wireless Intelligence.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling WirelessLLM: Empowering Large Language Models Towards Wireless Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:7cfecd4c09a2f9d70c784d442cbc566a365630f10c32c90c208786b865881e2a

Observation 9b97aedd-78bc-4c1f-928a-4c7361c03e17 · outbound

This paper cites The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:7899f4e6d12918924d20632bf78a611bf823d27b1d35794850ecd723fa1273c3

Observation 19818408-4229-48fc-88a5-d38a02cafde1 · outbound

This paper cites Continuous inspection schemes,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Continuous inspection schemes,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:ae4432e5ad0fdcd8c6bb33e6ce28b295cae8872a263e453e44238b06404deb76

Observation cc6084af-2ea0-4f23-8a1a-243e09fa68f0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:d0dde46bae9ce35286a36c84f89fa86c2433fef662f7266083782e20938b11a5

Observation 2f22e198-fa1e-4ad6-9e89-ead6537b9627 · outbound

This paper cites A definition of continual reinforcement learning,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling A definition of continual reinforcement learning,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:00b0f70fd50baa1eb796c8ade215d49e83b0c846b6a760e637a3070ecfe5589e

Observation b7ff86de-5294-4358-a5a1-363b6f949572 · outbound

This paper cites Gradient surgery for multi-task learning,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Gradient surgery for multi-task learning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:767ec10a6f264acfb87d76fec019b93711817e5659ee10767e96b09118fde732

Observation b640f302-29e5-411a-9048-de4f8e15c978 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive NLP tasks,.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Retrieval- augmented generation for knowledge-intensive NLP tasks,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:c718d08ee62ec51c10f98776d54a4cca056b2b9957421a447f66c2d3830acbe0

Observation aa0ae670-76ed-47e1-9c13-6d8bbf177cd0 · outbound

This paper cites Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval.

When Adaptive Rewards Hurt: Causal Probing and the Switching-Stability Dilemma in LLM-Guided LEO Satellite Scheduling Grounding by Trying: LLMs with Reinforcement Learning-Enhanced Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T13:05:01.450957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:05:01.450957Z digest=sha256:d5c6c070a53947633ff138de699ea4802a9aca5ec522af7e5a27fd5a5a20f23f

Pith citing papers

No inbound Pith citation observations are available.