Pith. sign in

Paper Citation Record · LEDGER

RvS: What is Essential for Offline RL via Supervised Learning?

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2112.10751.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10751 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:30:59.372776Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:57:25.909944Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c520d41-c0a7-43c9-accf-38eef9cf9e44 · inbound

Adaptformer: Sequence models as adaptive iterative planners cites this paper.

Adaptformer: Sequence models as adaptive iterative planners RvS: What is Essential for Offline RL via Supervised Learning?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:36:16.193382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:36:16.193382Z digest=sha256:2d930ba71636f4a73623dc20ac5ba07e5e93ef3978195a3c71ec079df0a84899

Observation 2154a41c-a724-4ca4-87ee-e2dc68cd691b · inbound

Are Expressive Models Truly Necessary for Offline RL? cites this paper.

Are Expressive Models Truly Necessary for Offline RL? RvS: What is Essential for Offline RL via Supervised Learning?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:40.401240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:40.401240Z digest=sha256:89e9006377d88b0c73366660c16329996d317e2bfb978e376f3f9bd3b9a23bf4

Observation d16239f2-7e0e-4b7b-85f1-50654e568540 · inbound

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning cites this paper.

MGDA: Model-based Goal Data Augmentation for Offline Goal-conditioned Weighted Supervised Learning RvS: What is Essential for Offline RL via Supervised Learning?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:02:44.088611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:02:44.088611Z digest=sha256:52c7c15d93ba3d075a730a462f45279cdb615049ee956c73defb17418b19a474

Observation 18e5abe6-5693-46f8-87b6-b08c8561b1bb · inbound

OMGPT: A Sequence Modeling Framework for Data-driven Operational Decision Making cites this paper.

OMGPT: A Sequence Modeling Framework for Data-driven Operational Decision Making RvS: What is Essential for Offline RL via Supervised Learning?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:23:51.177242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:23:51.177242Z digest=sha256:86514db97de9e1ea67022cf123fa7c49916e369222365345390324384a1e6193

Observation 7dfa7fc1-3245-47dd-9d6d-545b04214591 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL RvS: What is Essential for Offline RL via Supervised Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:45.370476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:45.370476Z digest=sha256:c7cb671ab90b75566752f1d9a446a917b9a3ae1148506aeedc80386daabb3f0e

Observation b26c610a-459b-44af-9acb-69a25dc55670 · inbound

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization cites this paper.

Closing the Gap between TD Learning and Supervised Learning with $Q$-Conditioned Maximization RvS: What is Essential for Offline RL via Supervised Learning?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:25.775369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:05:25.775369Z digest=sha256:a9109a68fbf50b3e66acefb9f724e73e6b2d09614817cfdb172dbe474ad8aa25

Observation 77a78522-f49a-4eb8-96e5-bd86295f3c90 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation RvS: What is Essential for Offline RL via Supervised Learning?

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:43.933445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:43.933445Z digest=sha256:f6fe83369172fe20c7b92b1d504ea40569dbea8b9dd7c9e04e40a5775772460a

Observation 54a25989-6fbc-430f-8a9f-c16e31be8935 · inbound

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service cites this paper.

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service RvS: What is Essential for Offline RL via Supervised Learning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:13:27.000122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:13:27.000122Z digest=sha256:d3870b70aab4051a848351320d80decd470d0c8a4b1cb6dd36de268f364665fd

Observation 97e7a599-1719-4386-aaea-c5b1a5d02a0e · inbound

Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making cites this paper.

Hybrid Sequence Modeling and Reinforced Verification for Controllable Target-Conditioned Decision Making RvS: What is Essential for Offline RL via Supervised Learning?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:23:27.070458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:23:27.070458Z digest=sha256:ef5030b6a6efadc1e79370c5837e595bcdd900a4ff0a24e7ed94058203a794d0

Observation abfabb09-a6fb-48e8-8cc1-5b0bbb4ce9c9 · inbound

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers cites this paper.

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers RvS: What is Essential for Offline RL via Supervised Learning?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:57.836508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:57.836508Z digest=sha256:d25faa86862f6f525043e84a5e1605a80279c89b373a185dc0aee451800d5b3d

Observation 7d9d66ef-e6a5-418e-b356-5d4829068f98 · inbound

Nonreciprocal current induced by dissipation in time-reversal symmetric systems cites this paper.

Nonreciprocal current induced by dissipation in time-reversal symmetric systems RvS: What is Essential for Offline RL via Supervised Learning?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T09:47:32.460955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:47:32.460955Z digest=sha256:ed8a6f680384d67cb5e2e08bbe9151ef0d057a77dbd02412a286472dae15ae12

Observation 79207e7f-55cb-4d85-bec5-370e17248870 · inbound

Receding-Horizon Control via Drifting Models cites this paper.

Receding-Horizon Control via Drifting Models RvS: What is Essential for Offline RL via Supervised Learning?

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:35:51.780977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:41:20.941243Z digest=sha256:f2bf5881393d6c8fb5032f2b076303aea88b3836ed1a4cd67c5a9bc9395d47bc

Observation e164e9d8-980f-4d36-b7ed-737e89913485 · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL RvS: What is Essential for Offline RL via Supervised Learning?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:58.020372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:67bf6f88566c7bdd7f71b12c47c306f3c18482ceb21c10db7f47544140af23d4

Observation e0f9405d-fcd0-430b-9810-a5586012827f · inbound

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos cites this paper.

Dash2Sim: Closed-Loop Driving Simulation from in-the-wild Dashcam Videos RvS: What is Essential for Offline RL via Supervised Learning?

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:57:10.317901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:15:08.063308Z digest=sha256:ebd3b031ebf31e1d365b13936abfb2de1f08d8038fc63c6c83bd1fc8050f926d

Observation bd9b792f-5df6-4528-b42d-a1c20966e1ec · inbound

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies cites this paper.

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies RvS: What is Essential for Offline RL via Supervised Learning?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.911641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T19:23:32.966503Z digest=sha256:84d9c18c235c2b74f38d12b9da811ff725ec6f334233bf05b25dbe026218ca52

Observation fa82bc0e-5d86-4685-ac2b-9c256cdc4859 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation RvS: What is Essential for Offline RL via Supervised Learning?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.940244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:62b4c16c4ce7d75eaa8a1051df23ad2d18a4df27c38deb6aaedafeba93f2d0d6

Observation 3303363c-3c5c-4159-bde0-22b1bcc557b5 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation RvS: What is Essential for Offline RL via Supervised Learning?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:046bb81f7c5d3a409f93e5be5581a6ee4a701f056d7335c4e01f73dc1d3997d6

Observation bc371d39-e35f-4d23-a2d4-32ee4f41e98d · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation RvS: What is Essential for Offline RL via Supervised Learning?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:20.993278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:20.993278Z digest=sha256:078c0f31bd3b175ee1c3df39cc95e0de388ba85375cadcd307bebb3765276c6b

Observation 9adc1b62-d588-4645-b977-26d4ae74af3a · inbound

Reinforcement Learning: From Algorithms To Foundation Models cites this paper.

Reinforcement Learning: From Algorithms To Foundation Models RvS: What is Essential for Offline RL via Supervised Learning?

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:14.035207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:14.035207Z digest=sha256:56c00f7c3b8a6660a9cfddbac9331c575edc3d915dfd3a3f0a0f7b83e36173db

Observation e8cb85e9-1b8f-4505-bafe-81e03d552cd5 · inbound

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play cites this paper.

When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play RvS: What is Essential for Offline RL via Supervised Learning?

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-01T12:36:10.784085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:36:10.784085Z digest=sha256:bb67e091df5caea74a92bba433d261007198a3eb0bf2c2810f8d3c87331ac2f4

Observation fd174571-0244-4399-a8de-a133003556ab · inbound

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression cites this paper.

A Factor Graph Approach to Scalable Multi-Output Gaussian Process Regression RvS: What is Essential for Offline RL via Supervised Learning?

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T00:30:59.372776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:30:59.372776Z digest=sha256:209184e9d7dcff68c79652747d6083429204e1d65e28462fe10474d48d50586f