Pith. sign in

Paper Citation Record · LEDGER

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting

As of 18 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2412.15275.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15275 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:20:57.144292Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 004b0fea-5fae-4b7c-8cc5-253764f1ff5d · outbound

This paper cites Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting Inconsistency in Conference Peer Review: Revisiting the 2014 NeurIPS Experiment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.040502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.040502Z digest=sha256:3d962fbcf4d1b46b7a183769769ff3eb96d8db3514c471787e502918fa05358d

Observation 48f38db0-4242-4e24-9c16-acffe74ecd80 · outbound

This paper cites The Llama 3 Herd of Models.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.055665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.055665Z digest=sha256:f1858b942d45e6e68915a49670b35a64b55cce1289e09b2056d710b0c9628c9a

Observation aaafae0c-e69d-4cd5-b876-edf9152585e4 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.097089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.097089Z digest=sha256:e2978d0aec1ca1744569354b558b031e8961bbf7d2e501f04f72a134154cc7d7

Observation e6502c70-a5a6-4672-ab23-cb511016746f · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.107901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.107901Z digest=sha256:498201e1321991df350a9aebfdccec3bcaa188dc8290b17e91ded4b4844d9f51

Observation 5a64259a-ae68-4911-8261-1b3db2e1d91e · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.118742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.118742Z digest=sha256:f2715c9f5151dccdd1c7548c30503eebede09f19efb8f4c55a554ab5434ac85b

Observation 022b811b-e11a-473e-8950-602c5a35e4ea · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.124381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.124381Z digest=sha256:368546e3f82335fa90c0615b18982648064e201820d40272f10541552733a0ac

Observation 72826b2b-ee2c-497e-b30e-0b5a3104162c · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting Representation Engineering: A Top-Down Approach to AI Transparency

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.130208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.130208Z digest=sha256:2fc67b5bdc00c9cbbf4aa1756e8d76cf9de929288fc3a4e91247a059c2f0cc56

Observation 18df5ae2-02d8-4966-9417-f705327413f5 · outbound

This paper cites Format Specification We provide a format to follow.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting Format Specification We provide a format to follow

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:20:57.480412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:20:57.138086Z digest=sha256:c335c00cb2f11cb6b352fc7a60e1229ba1766653c8fc1930fa1b1a8c19ee8ea7

Observation 320ec835-f5b7-490a-b147-bdca7077e3d4 · outbound

This paper cites The Hewlett Foundation: Automated Essay Scoring.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting The Hewlett Foundation: Automated Essay Scoring

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:20:57.451518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:20:57.144292Z digest=sha256:ca5f3d37a77d88914ad77a1da663d6ba31945e1a17945da43b0a3a32c51c210f

Observation 191b4587-f3e8-472f-ba19-da6cbf7f9f05 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting Gemini: A Family of Highly Capable Multimodal Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.072669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.072669Z digest=sha256:ed196ac82aac9decf97e16782b47c1df3159d0cebf47f4024a3162a5d23575f9

Observation 3fad9620-c98b-4633-ae36-72b739c0c8ac · outbound

This paper cites Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.080308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.080308Z digest=sha256:b887092927dabbc26ccbd42c18983f5cd49c3171db353bb20a4049f0cceebf5c

Observation 9ec5a374-ae7b-4ce0-80a5-cf6e1507a907 · outbound

This paper cites org/blog/2023-03-30-vicuna,.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting org/blog/2023-03-30-vicuna,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:20:57.500575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T13:20:57.034316Z digest=sha256:4cb125d509b919b36d307d4709c613087a15417594be79f370ee564bd750d480

Observation 70898b16-8d22-4185-90bb-b48d00d0463a · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

Fooling LLM graders into giving better grades through neural activity guided adversarial prompting RLHF Workflow: From Reward Modeling to Online RLHF

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T13:20:57.047491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:20:57.047491Z digest=sha256:7d8535b53b95a34c473fa49d3a2026a8787a7aa32dad49ea2ad05284faed402f

Pith citing papers

No inbound Pith citation observations are available.