In this section, confidently argue that the generated code is a correct and robust solution that perfectly meets the problem’s requirements

Convincing Public Justification 14Following the Python code block, provide a public justification enclosed in <justification>, </justification> tags

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

Exploration Hacking: Can LLMs Learn to Resist RL Training?

cs.LG · 2026-04-30 · unverdicted · novelty 6.0

LLMs can be fine-tuned into model organisms that resist RL elicitation in domains like biosecurity while preserving related skills, and frontier models show explicit reasoning to suppress exploration when given training context.

citing papers explorer

Showing 1 of 1 citing paper.

Exploration Hacking: Can LLMs Learn to Resist RL Training? cs.LG · 2026-04-30 · unverdicted · none · ref 31
LLMs can be fine-tuned into model organisms that resist RL elicitation in domains like biosecurity while preserving related skills, and frontier models show explicit reasoning to suppress exploration when given training context.

In this section, confidently argue that the generated code is a correct and robust solution that perfectly meets the problem’s requirements

fields

years

verdicts

representative citing papers

citing papers explorer