Pith. sign in

Paper Citation Record · LEDGER

The Impact of Off-Policy Training Data on Probe Generalisation

As of 4 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2511.17408.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.17408 v4

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:26:37.914522Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T21:44:25.201514Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact25
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a8a3a90-05a4-40f9-b188-114f64780575 · outbound

This paper cites Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models , May 2025.

The Impact of Off-Policy Training Data on Probe Generalisation Linear Control of Test Awareness Reveals Differential Compliance in Reasoning Models , May 2025

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.631056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:93e2025f12ee775e11d3f562014ef08a266f33e58154741b6b9bc3937ba5a5fe

Observation c4ee3199-3bb7-4bf7-a877-572424067f4d · outbound

This paper cites Invariant Risk Minimization.

The Impact of Off-Policy Training Data on Probe Generalisation Invariant Risk Minimization

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.624218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:bf2941d53634f44d1749a5c2affe35a4d890caf2f370d4e153253bc0dac8fff1

Observation a0976b27-f64c-4805-9076-84da2d039a73 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The Impact of Off-Policy Training Data on Probe Generalisation Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.683687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:2ec8ebf23a5e99785dd98fa17448d92ba257fc298fa057573844858d3811a3d7

Observation a909a741-651d-4fb4-8881-87be77fe645a · outbound

This paper cites Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models.

The Impact of Off-Policy Training Data on Probe Generalisation Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.686788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:e27260160f6f38029f44d9ac8cb1d5192f0ebe5c27b1fd0374b9fbef3f467eea

Observation 6c0a7faa-e46b-40ab-bb19-68de55c9c107 · outbound

This paper cites Sabotage Evaluations for Frontier Models.

The Impact of Off-Policy Training Data on Probe Generalisation Sabotage Evaluations for Frontier Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.717804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:6bdc3b3b63f770c37182308084240776b1bf384f6bde1e6c02a4e6fa39dc14f3

Observation 82feaba0-7f2d-4d2e-b478-82de071c8d7a · outbound

This paper cites Propositional Interpretability in Artificial Intelligence.

The Impact of Off-Policy Training Data on Probe Generalisation Propositional Interpretability in Artificial Intelligence

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.678096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:5b51668fcb91cb91a807ffdf5cdb9ea013661652661b83956514fb5b91c295bf

Observation 0909cc93-57ac-4aaf-a3a8-5145ba3e81bd · outbound

This paper cites month = jul, year =.

The Impact of Off-Policy Training Data on Probe Generalisation month = jul, year =

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.720919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:f709b81b348640babe348c2d5ca210cfd0b5567a3c9a1ebee746c71f6fe0c48a

Observation 0a07bec9-ee27-4f4c-89cc-d95e5b19e0d5 · outbound

This paper cites Cost-effective constitutional classifiers via representation re-use.

The Impact of Off-Policy Training Data on Probe Generalisation Cost-effective constitutional classifiers via representation re-use

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.634244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:5e85c3f8ccb3b905f9f601e5c07d417659cce8e998d3e9e9c00cc5d666875413

Observation d7b924d5-dd2e-415e-83b2-c5c32941f2bf · outbound

This paper cites DeepSeek-V3 Technical Report.

The Impact of Off-Policy Training Data on Probe Generalisation DeepSeek-V3 Technical Report

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:30:11.723980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:242753723b8f854a0259e9fce1f5ac6343f7af6b4a29928f36bdd1e1ab6b2e27

Observation a002ea0c-53d2-4133-9796-41a052437ee4 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

The Impact of Off-Policy Training Data on Probe Generalisation Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.731506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:68b12f442c88f96f5556e48961c76891ebab843ef15de6052dffc4ccc3945834

Observation 7aa3136f-6cee-4b02-b2ee-c536505dc6db · outbound

This paper cites Hierarchical Neural Story Generation.

The Impact of Off-Policy Training Data on Probe Generalisation Hierarchical Neural Story Generation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.680944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:fce27addaab3679bec515f74cb6733971c4916f118cd5ed6cf7eab995a4e5f98

Observation 959b5e21-c011-42e8-b79d-07922309fb87 · outbound

This paper cites Monitoring Latent World States in Language Models with Propositional Probes.

The Impact of Off-Policy Training Data on Probe Generalisation Monitoring Latent World States in Language Models with Propositional Probes

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.741041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:2e7c3ceca04278dbcf9e9584b5cb3d468a525c9374f7e3ef80f2dcbdfa13d16c

Observation a94f196f-6b16-4e25-af5c-d6b9431a1a50 · outbound

This paper cites Detecting Strategic Deception Using Linear Probes.

The Impact of Off-Policy Training Data on Probe Generalisation Detecting Strategic Deception Using Linear Probes

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.634956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:96af1069c7b6e161209bf1f75a817cb9f9fb6ee8bac97ea47784dc938b0738a4

Observation 02cf7c9a-d832-48d7-b219-06c3e5ffc297 · outbound

This paper cites Probing the Robustness of Large Language Models Safety to Latent Perturbations.

The Impact of Off-Policy Training Data on Probe Generalisation Probing the Robustness of Large Language Models Safety to Latent Perturbations

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.710099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:823a83f04dfd6c6d783bbe19f50d5893e97fafa2429e9a2c0017842fc573a46c

Observation d3bfb3de-ec41-4833-afea-76c02aa7920b · outbound

This paper cites What makes a convincing argument? Empirical analysis and detecting attributes of convincingness in Web argumentation.

The Impact of Off-Policy Training Data on Probe Generalisation What makes a convincing argument? Empirical analysis and detecting attributes of convincingness in Web argumentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.631651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:da72b99adcfca3ccf3518e49a6e7ad3f30811c0f2fc25a3a4f259a9830d99f37

Observation 79ac2412-2d6f-42c6-b198-7b381d400692 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

The Impact of Off-Policy Training Data on Probe Generalisation Measuring Massive Multitask Language Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.751114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:b4a0495a349e1f4fdc416b217e4012941b94f3d9b68fe8dddd013f9dcbedf6f2

Observation bd0d3745-0d29-4e4f-b4ab-99c548b7221d · outbound

This paper cites Ministral-8b-instruct-2410 model card.

The Impact of Off-Policy Training Data on Probe Generalisation Ministral-8b-instruct-2410 model card

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.636187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:92d1668a514325172581189f208b1d959b40acf6a22a5ca57b8f9a3602bb851f

Observation 1eb8a53c-45f6-4076-b224-cb1d80a8e7b5 · outbound

This paper cites Mistral 7B.

The Impact of Off-Policy Training Data on Probe Generalisation Mistral 7B

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.734593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:0390b846094c40c1e6a085c7fb3a72465d7c1dfeacc94c0c1bf95119cb4a1623

Observation 25a73da2-530c-4daa-ab37-c870c76670f1 · outbound

This paper cites Mixtral of Experts.

The Impact of Off-Policy Training Data on Probe Generalisation Mixtral of Experts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.696061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:46ddbd26150fa843d630e3eec9d0733d057a7bc5e35632cae3280d797c94835e

Observation 1b81dfbf-c725-44ca-b6bf-bb259fa14fd8 · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

The Impact of Off-Policy Training Data on Probe Generalisation WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.690062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:deedfd2d4a2505b42e48509053387e4b88235d2024cb3dec341230c7c0fec812

Observation ca1cfab4-bcff-43a4-87ba-f32f77a1ec62 · outbound

This paper cites HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States.

The Impact of Off-Policy Training Data on Probe Generalisation HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden States

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.638926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:6f452c11affc7dadc2e85a1ee877fffa439805f88a623053113c6f1ac59f9ebc

Observation 70b657d8-d71f-47a7-bc9c-175570f96ead · outbound

This paper cites What Features in Prompts Jailbreak LLMs ? Investigating the Mechanisms Behind Attacks , May 2025.

The Impact of Off-Policy Training Data on Probe Generalisation What Features in Prompts Jailbreak LLMs ? Investigating the Mechanisms Behind Attacks , May 2025

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.693163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:f3b672c847dd5bd55660a88fc53630e28f304be7937089d31ae74211a3c3c60c

Observation 92aa5abd-c7f2-492e-b7c8-0e560de0b304 · outbound

This paper cites The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.

The Impact of Off-Policy Training Data on Probe Generalisation The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:30:11.702463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:34080e57891f951a6dbda8008d9e95f9117f70b81b6936a96cafd85c5b9f7a42

Observation 3a8746ac-936c-4f88-acec-c7b24dccc6ad · outbound

This paper cites Simple probes can catch sleeper agents, April 2024.

The Impact of Off-Policy Training Data on Probe Generalisation Simple probes can catch sleeper agents, April 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.628778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:a91c5a2889d63469227716abb1e08733de44ed8ea7b2f251986537b87e96c489

Observation e24a5014-e9cf-435f-93f7-3f24e5e0918a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

The Impact of Off-Policy Training Data on Probe Generalisation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.668116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:2d2b410340f4955a62c88a0094a66a548711ebeb8144bdaa8e8694ac3b3d7c7e

Observation ca924539-86f7-4c76-a6dd-40580c9414ac · outbound

This paper cites Detecting.

The Impact of Off-Policy Training Data on Probe Generalisation Detecting

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.713833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:9a62cd6a8b25d586681912f198809db6054b1fc7621d90c8bd58594b9364d0c0

Observation fefb9d7f-eb7c-40cd-a25d-02dd83fb947e · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision (connect 2024).

The Impact of Off-Policy Training Data on Probe Generalisation Llama 3.2: Revolutionizing edge ai and vision (connect 2024)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.638155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:448c0017403c92fff92b58dc5017735473ec1341d122f2be1b4ba444400b1870

Observation d7c3d801-dbdc-4dd5-9ac4-4c72dc6b0793 · outbound

This paper cites Probing and Steering Evaluation Awareness of Language Models.

The Impact of Off-Policy Training Data on Probe Generalisation Probing and Steering Evaluation Awareness of Language Models

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.674796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:92ee91e6d79a83a08f5cd357c43a98016957f89d2b3cc6eb6af6f843044190d4

Observation 01f125d5-1082-46f1-b19c-2b2f1616d802 · outbound

This paper cites Gpt-5 system card.

The Impact of Off-Policy Training Data on Probe Generalisation Gpt-5 system card

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.626578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:eed89cbe6c809aa48b293dfe60f363377312ebc2327f8fd52c52120252022190

Observation af0a14e1-a378-45be-91e4-1e257955858b · outbound

This paper cites Benchmarking.

The Impact of Off-Policy Training Data on Probe Generalisation Benchmarking

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.748270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:c8942fd7baba546104e3bc0aec18439c02e957147bb9a59450f1b47c0962b9e6

Observation ef8021db-546e-45e2-a846-846b8b0e1d7e · outbound

This paper cites Qwen2.5 Technical Report.

The Impact of Off-Policy Training Data on Probe Generalisation Qwen2.5 Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.671475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:d4fb0caf712b38a5524c76fc565ca77b373792dd52da562ad5708b250200e140

Observation 3079e435-eeb2-4085-b96d-83efcdc32f64 · outbound

This paper cites Coup probes: Catching catastrophes with probes trained off-policy, November 2023.

The Impact of Off-Policy Training Data on Probe Generalisation Coup probes: Catching catastrophes with probes trained off-policy, November 2023

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:32:05.624238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:d1b2a0557767eb27bac45fa634bdc6c9c91d4ce7d433820bf73873c1fb607935

Observation 7ae73266-38ec-4e72-9fd0-ace740a1501b · outbound

This paper cites Large Language Models can Strategically Deceive their Users when Put Under Pressure.

The Impact of Off-Policy Training Data on Probe Generalisation Large Language Models can Strategically Deceive their Users when Put Under Pressure

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.706210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:d666ab313686f776f23571c0f3760d387b6ee483b023c1067f1c19748f28b099

Observation 9d2b4a8a-7ebf-4385-9d97-733ebfd00c4b · outbound

This paper cites Stress Testing Deliberative Alignment for Anti-Scheming Training , url =.

The Impact of Off-Policy Training Data on Probe Generalisation Stress Testing Deliberative Alignment for Anti-Scheming Training , url =

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.661669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:209b06c7a29fa363b97aa5ea74e16bfbca9710df5a0921e8e3cd66b6042b7c4e

Observation c1cba9f9-b963-402d-8ddd-a3f1ccecf800 · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

The Impact of Off-Policy Training Data on Probe Generalisation Towards Understanding Sycophancy in Language Models

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:30:11.664946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:434f462e96b91284b352cc5a4e4b394c4b32af17993de6a037c94ab9ac2142c4

Observation 7f9011f4-d26a-4d53-ae5d-8758c55f0b39 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

The Impact of Off-Policy Training Data on Probe Generalisation Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:20:43.320085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:8751aab6a017346e3b9958a7407612e86a1e56040f03f4579b132c2e1db85dd2

Observation ea023788-c5a2-4826-a689-1a9d9fcefc7f · outbound

This paper cites Gemma 3 Technical Report.

The Impact of Off-Policy Training Data on Probe Generalisation Gemma 3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.627465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:97f22e07131f8ebc2c58373123acaff857122d89ecbb2aa8f3d02e43508e5cc0

Observation 13e64531-f004-44ad-ad14-43ff8a01c0cc · outbound

This paper cites Investigating task-specific prompts and sparse autoencoders for activation monitoring.

The Impact of Off-Policy Training Data on Probe Generalisation Investigating task-specific prompts and sparse autoencoders for activation monitoring

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.699412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:07673c8f1136690ef397ecea91ff8a722aaca48bc6335de0176c93dab61c4f63

Observation cc3be326-0d17-47e5-a29b-fbdb71077bad · outbound

This paper cites Diffusion Earth Mover's Distance and Distribution Embeddings.

The Impact of Off-Policy Training Data on Probe Generalisation Diffusion Earth Mover's Distance and Distribution Embeddings

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.655190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:c122e9aec2bf5090ca654714924a8f4973723c2f1431dced275dedcdaa0dac26

Observation fe206e23-4a62-4278-9380-06c4394ddcae · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

The Impact of Off-Policy Training Data on Probe Generalisation Jailbroken: How Does LLM Safety Training Fail?

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.651388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:18461e8b2cc1f1631a211eee762c7a1af15739510a47a8974d5ebdfe9ecd60c4

Observation f1888f08-cb37-4685-97a3-ca519fe7e19f · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

The Impact of Off-Policy Training Data on Probe Generalisation AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.658351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:face438af95d4f835ce9d216d2b4422cd6054f213e393022b01c1dcc72437b13

Observation 9503af87-6367-4937-aa1b-b0679ca36d34 · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

The Impact of Off-Policy Training Data on Probe Generalisation Language Models Learn to Mislead Humans via RLHF

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.646597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:c13a9ab4620a3776e957e74892ae841f85a4d97994a4bca579d86930822993f2

Observation 42fce047-3444-4496-ace5-7163134d68ad · outbound

This paper cites Qwen3 Technical Report.

The Impact of Off-Policy Training Data on Probe Generalisation Qwen3 Technical Report

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.642685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:82086916ce102ccd59d6e908ca1590ebedd6a10585ed5be65ea35b92edca0644

Observation 242bea42-9b0c-4cb5-a049-7ed2e177a98a · outbound

This paper cites Uncovering Latent Chain of Thought Vectors in Language Models.

The Impact of Off-Policy Training Data on Probe Generalisation Uncovering Latent Chain of Thought Vectors in Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:30:11.727486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:6043e92830930a1e84c524c08ad1929c11c84782158e95e7917d529efd998c6f

Observation d506ccca-0558-4d99-a37f-5f0ee8994d3b · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

The Impact of Off-Policy Training Data on Probe Generalisation Representation Engineering: A Top-Down Approach to AI Transparency

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:30:11.744338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-17T20:26:37.914522Z digest=sha256:7996331c4a46b6af5c18d276b48a83df8e082207e2fe070a834f964099c94f1e

Pith citing papers

Observation f6d7ce6b-f995-4274-8788-22a11b04f80c · inbound

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong cites this paper.

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong The Impact of Off-Policy Training Data on Probe Generalisation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-30T21:44:25.201514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T21:44:25.201514Z digest=sha256:a9ce668a67eb0ef2f37ed9c2b4fb09eea1bc886cb2d4839e871eb2f15e11fd07