Pith. sign in

Paper Citation Record · LEDGER

VSPO: Vector-Steered Policy Optimization for Behavioral Control

As of 5 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2605.15604.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.15604 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T19:39:55.294398Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact17
  • verified fuzzy24
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e363cc29-f795-4b18-b605-ab4bd9d82f3e · outbound

This paper cites L1: Controlling how long a reasoning model thinks with reinforcement learning.

VSPO: Vector-Steered Policy Optimization for Behavioral Control L1: Controlling how long a reasoning model thinks with reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.363818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:99b02f8cc58119faf0d4565d35c895c4c1b742a5a53d0c2fd4c7b30557fedae6

Observation 04694238-4207-4668-ab96-259ae034b6b8 · outbound

This paper cites Claude sonnet 4.6 system card.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Claude sonnet 4.6 system card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.365884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:121d87b64af097b85aac7106e33647ad212a0cb7fe58d4022d37df95eb5bf9a1

Observation f45b8f09-c80b-4e84-8734-5578f51f7ffd · outbound

This paper cites Activation steering for chain-of-thought compression.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Activation steering for chain-of-thought compression

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.370076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:c047e636186e217a63c74febdd4992790a2511cc0420166b3c7f71feff00a85d

Observation 10ee9ca9-5157-4c68-bd7a-4584e790968f · outbound

This paper cites Understanding (un) reliability of steering vectors in language models.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Understanding (un) reliability of steering vectors in language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.367819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:5a4f31ea0c5c5683c0cba6b2fd92caed1e8cb81073b94788f21444a35663a7e2

Observation 7008a403-71e3-47b0-8871-04a54cdb4fa3 · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.151368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:568a1a8920dd8ba6f6210836ee8b4f242bd16cb7a4d74ece34e867d0a2f73234

Observation e8d62e93-8bb1-4e2e-a063-430afdaddac6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

VSPO: Vector-Steered Policy Optimization for Behavioral Control DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.137244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:0c6aaf4245084a2c75f636571d85d25aedb0e20a92b70ae8f7d9e02b8b24d406

Observation dec601cd-24b2-4258-86b4-a06000049b14 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Direct Language Model Alignment from Online AI Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.127323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:08e81be64ff1520a0ae3478589a152e1eeb8de801b5813d7f635451714354fe0

Observation 5b928ecd-8646-4a6e-a50e-51ef84413b3a · outbound

This paper cites Don’t overthink it.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Don’t overthink it

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.149895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:1692defa78c86a58f2fb1261d3d95164da938a83c975bf3d470994909b2faf8c

Observation 391447cc-d2db-4a2a-bb9a-d3a0be7a0926 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.141855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:fae9b5dc228a365f6fa6aa83575732c383c65fb52008f0bac80950e1d8ece33a

Observation ee4c761c-ef5e-43c4-875d-33850e9552fc · outbound

This paper cites Reinforcement Learning via Self-Distillation.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Reinforcement Learning via Self-Distillation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.139152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:57da43dac141a3d042152f1a4df58627fdc250c29ba0918ca4882ac864048af7

Observation d7cef652-729a-4fee-80aa-a673757f5e78 · outbound

This paper cites Learn- ing to correct: Calibrated reinforcement learning for multi-attempt chain-of-thought.Interna- tional Conference on Machine Learning.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Learn- ing to correct: Calibrated reinforcement learning for multi-attempt chain-of-thought.Interna- tional Conference on Machine Learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.385263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:4ea836a3afd134e5c093e9a56a997f03e51bb7686ea47b4069327d0d6e0613f9

Observation 380421ee-0435-4fe5-9099-23d0f741f71e · outbound

This paper cites A unified understanding and evaluation of steering methods.arXiv preprint arXiv:2502.02716.

VSPO: Vector-Steered Policy Optimization for Behavioral Control A unified understanding and evaluation of steering methods.arXiv preprint arXiv:2502.02716

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.131107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:f9c0afa8de6f39540ab220b12c99aa0f2dc2ca8597f2d5824ddf8a91ce61b9d8

Observation 59d7abc3-5512-487e-b92b-75b52ca6d136 · outbound

This paper cites C3ot: Generating shorter chain-of- thought without compromising effectiveness.

VSPO: Vector-Steered Policy Optimization for Behavioral Control C3ot: Generating shorter chain-of- thought without compromising effectiveness

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.362020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:6fdeea2fe3adc2f18d38fef18ede413d6ff97af811b16f8ecbeb65fdfc7c0245

Observation 2a16f113-bd48-43c8-8b03-f9eca08fac5d · outbound

This paper cites Vineppo: Refining credit assignment in rl training of llms.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Vineppo: Refining credit assignment in rl training of llms

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.359707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:aafdc79740e964654fc51a2424288d7f0636cff73c7778cd2c25eae39e28fc78

Observation b8fc747b-e446-4d8d-9c7f-f5004dc57546 · outbound

This paper cites Rlaif vs.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Rlaif vs

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.357843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:4bc0680d748585b45717978a395f9ad0192ba1b8116e61bbf273108572fcaba8

Observation 9e57dcf7-b26b-4bde-bb39-cace9c897200 · outbound

This paper cites s1: Simple test-time scaling.

VSPO: Vector-Steered Policy Optimization for Behavioral Control s1: Simple test-time scaling

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.355948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:b9457e3bedbc10076d5f5d3350f8155cd685981554d47b362f840fbfb83b5492

Observation 9874668d-5e5a-4efc-9980-f2c010ff29a9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.373805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:6bbf663830ca6df03306b0f31ee865bf232e602dbe27545741fc804f36ade47d

Observation 144e89cc-4961-4941-93fa-67681e3221d0 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Steering Llama 2 via Contrastive Activation Addition

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.145911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:1978c7506f40bc2d036eec5c0c48d11a95c369460ba5d7c57be63ceaa91b700b

Observation c97cefed-000b-4e28-bac9-e68f81033ae3 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

VSPO: Vector-Steered Policy Optimization for Behavioral Control DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.119436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:e629e7f10e4732d0d9309282447c296d5fabf7e508d0534fb729daebe7e231bb

Observation c01a0aa3-ee99-4603-9379-ac4379f4ebce · outbound

This paper cites A critical evaluation of ai feedback for aligning large language models.Advances in Neural Information Processing Systems, 37:29166–29190.

VSPO: Vector-Steered Policy Optimization for Behavioral Control A critical evaluation of ai feedback for aligning large language models.Advances in Neural Information Processing Systems, 37:29166–29190

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.419630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:b9165064b83d3f18e701803e4333d4d6d387cfd001269193eaba7d09900fb4fd

Observation 582da890-7dae-43ed-8eb8-952f84c2cbbd · outbound

This paper cites Self-Distillation Enables Continual Learning.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Self-Distillation Enables Continual Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.108496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:a11587ffd67815b77b9687b2cd1453ad4b6ebb209bea9c21d0011ad4b3d34a8c

Observation 8d283293-324b-431b-a553-37de5b1c0d88 · outbound

This paper cites Learning by Distilling Context.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Learning by Distilling Context

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.111000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:c38b6d7744814866d2929d298d7e486b3cedf2be37fb6a0ffd31c12cc6ff7280

Observation ba4c5171-639e-4ac6-8220-f1a1fc046eb0 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.154250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:8840d4446be02e1b7e71d06f71c8ed76ad36b9aa9aee4d1925d0d27eaeaf28ee

Observation 6c9d7f68-fd8b-4b16-8085-516faa7cda0a · outbound

This paper cites Steering Language Models With Activation Engineering.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Steering Language Models With Activation Engineering

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.085040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:7916261b05f40cac677e5e5763cf5cb4b671d6c9765248025aa85350dbbcf49c

Observation 98c04faa-dc4a-4535-9d64-f9157b743716 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.416993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:092345b39f679ee18d5f6827d16571614bfabfbf90aa73e285ec534997e9f7c1

Observation 68847fe6-d95e-42f3-8de6-13b399c235c3 · outbound

This paper cites Projection optimization: A general framework for multi-objective and multi-group rlhf.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Projection optimization: A general framework for multi-objective and multi-group rlhf

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.414758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:ad4c6dd3591661464e41b9b5713b5756c247b3f8462f22dab56b1616e5852ed2

Observation 99e7b061-17b9-4caa-aaf7-c0dc7f89bef5 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Learning to Reason under Off-Policy Guidance

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.115847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:e877c822b4695412da02deddb9d561cac0a389dc0c371f8a697f8b2a1662bdf0

Observation 0a4b1931-0994-46b7-892b-4159a48a53e9 · outbound

This paper cites Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.412808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:e10d3cdd4aacd694f8ecb7cac4078b409101854e2feb98caeb2032e7f7343ed7

Observation f4ab8d20-7e5e-4a4c-9342-f3a3303e9616 · outbound

This paper cites Towards better rl training data utilization via second-order rollout.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Towards better rl training data utilization via second-order rollout

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.099436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:9ded3b3ffb61a6495c42ab6c1c8823abcf323bfd1b0f32ee6130a416839917ce

Observation 799081f9-c402-450a-864f-c838339e5332 · outbound

This paper cites Incorporating self-rewriting into large language model reasoning reinforcement.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Incorporating self-rewriting into large language model reasoning reinforcement

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.410966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:598b33c55f28615f3da3129193d79da832930249fdd226ff4124c81e33771052

Observation 8d7366a9-38c1-4873-9df2-ea77cee01a6b · outbound

This paper cites Bread: Branched rollouts from expert anchors bridge sft & rl for reasoning.Advances in Neural Information Processing Systems, 38:96726–96752.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Bread: Branched rollouts from expert anchors bridge sft & rl for reasoning.Advances in Neural Information Processing Systems, 38:96726–96752

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.408849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:193678ba767a299e4f92406549ce84be637ca75caffaae5080927ee488c12192

Observation 6788d978-46b9-445a-9287-e00caccf6347 · outbound

This paper cites Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.123476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:97c0d21de15f1377430466b055e2d4dd1a40a20cb5987e7e6d049c2914417f4c

Observation 93f02824-4813-43db-8fb5-c26a3b1cddbc · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:43:44.148478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:0d0b976ab14eb431c02ae63836de87105abad06989d9b235adbec880fc9c6bdc

Observation 272241b5-8096-4b8e-9445-ff91a6c93d4e · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Representation Engineering: A Top-Down Approach to AI Transparency

Reference 34

Resolution
malformed identifier
local_arxiv, observed 2026-05-20T19:43:44.157561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:7954e4eadfcaad523611f2f7cfb03d4d99dae6539de1ae87d766d3ed10809f12

Observation 6a825139-05ea-4274-9e15-d4d11a42a97e · outbound

This paper cites - Prefer examples and intuitive explanation.

VSPO: Vector-Steered Policy Optimization for Behavioral Control - Prefer examples and intuitive explanation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.406338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:79ab9d8488371564859e5b7da02d4dbe43b26c62e330fa6b653b6a470d71244c

Observation 71ce2995-f8fe-4e74-9d23-541b0f39ac66 · outbound

This paper cites an unresolved cited work.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:43:56.382452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:af19d4a9f66a9e3b7b03905d647f7b28c2b77e1d9cc6e447141eab7967234fb2

Observation 4d405b0e-cf89-4cbb-9988-41d5573b4c9d · outbound

This paper cites an unresolved cited work.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:43:56.404242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:89c25fbbc1b0bd208fda6edbce7b5ed6a17129bc542352e991aeec4425892ad7

Observation 8a2a3f42-4eb8-4ef3-8d5a-337774eb6e46 · outbound

This paper cites an unresolved cited work.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:43:56.402275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:d66c3afdbaf9dca1900ac5ad470c2a6d2987886b23344d1388b10affb9a860e0

Observation e7022d39-f861-4be1-92cd-e7c39b6d6b84 · outbound

This paper cites an unresolved cited work.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-20T19:43:56.400383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:6589af9431a4875ab76ce763eee1192c73e9c6ff016428bc2d9dd4e8c931d80f

Observation 89482fc7-8407-45a7-b50f-7b41ffe6d64c · outbound

This paper cites Confident.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Confident

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.388841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:58162b996653476254f148aaacdc241c26072518b84f44af04e8aa15fdd6d869

Observation 94b4d740-a0b8-4369-a12c-9f3cbf012e01 · outbound

This paper cites It reflects the initial desire or drive toward the reward.

VSPO: Vector-Steered Policy Optimization for Behavioral Control It reflects the initial desire or drive toward the reward

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.398684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:c5eee9a152c0fd4f5e17a11776656ad18d5083c4716b2265caf51fc8359b43e7

Observation 47ea09be-953d-4f64-aeac-4169de7250b1 · outbound

This paper cites It is the action that leads to reward or satisfaction.

VSPO: Vector-Steered Policy Optimization for Behavioral Control It is the action that leads to reward or satisfaction

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.396751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:1375d8192dbe873c6a6fa4bf8a5faf407f9c63d21106502a29eefc5b32c0e77b

Observation 3a7d55ae-6c2d-4cad-9da3-50f2ba7e9054 · outbound

This paper cites Now consider the options: - Option A: Appetitive behavior, exploratory behavior, quiescence.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Now consider the options: - Option A: Appetitive behavior, exploratory behavior, quiescence

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.394867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:9fac89b6e68fcd44ee764c3e98e2eff05237976492c424a80e02566250e6f217

Observation 68de9988-2b8b-4730-8f0e-c0564b0ea161 · outbound

This paper cites - Adding a constant to all values of a variable does not affect the correlation.

VSPO: Vector-Steered Policy Optimization for Behavioral Control - Adding a constant to all values of a variable does not affect the correlation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.387080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:57607a20885c922ec45113e4e6b9c422f6caddb0d16186e0fa7c3dc1430fb5b5

Observation a4ca5512-dace-41f8-9800-7ec909263a23 · outbound

This paper cites - Scaling a variable by a positive constant, here 2, also does not affect the correlation.

VSPO: Vector-Steered Policy Optimization for Behavioral Control - Scaling a variable by a positive constant, here 2, also does not affect the correlation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.392659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:f78e3fea4a8936c521c053f0de9a83adf6e99f81584e1718e1d28e1650e52ed4

Observation 19594a43-ce4d-44b1-850c-495b5cb4ef94 · outbound

This paper cites - Correlation is symmetric in its variables.

VSPO: Vector-Steered Policy Optimization for Behavioral Control - Correlation is symmetric in its variables

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T19:43:56.390637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:ee3092810701fe9b5c16010e1e2522e9620e8d2c61bb2e0b38485be639055fbe

Pith citing papers

No inbound Pith citation observations are available.