Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2402.14992.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T02:32:05.153277Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 23c0e894-68e5-428f-b2d2-5cecdb0b491d · inbound
Refusal in Language Models Is Mediated by a Single Direction tinyBenchmarks: evaluating LLMs with fewer examples
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a858ae38-1d42-40f8-a4b0-022b15c8c58e · inbound
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3cdd8a53-8bd8-4ec4-bb25-68b8e5613c90 · inbound
Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation tinyBenchmarks: evaluating LLMs with fewer examples
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96a60ddf-57d1-4ac5-8a55-b1319e00fd58 · inbound
LLM-Safety Evaluations Lack Robustness tinyBenchmarks: evaluating LLMs with fewer examples
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cde4c61-88c4-4323-91dd-789a8c441a12 · inbound
Small Language Models are the Future of Agentic AI tinyBenchmarks: evaluating LLMs with fewer examples
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 28acc0ee-d2c0-4e64-a7ce-0c9b8f526a57 · inbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users tinyBenchmarks: evaluating LLMs with fewer examples
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3687ec11-4f3e-44fc-830e-ae480d056225 · inbound
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48a3726e-5ed3-4e65-b1cb-d09fa4b169ce · inbound
Activation Steering with a Feedback Controller tinyBenchmarks: evaluating LLMs with fewer examples
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 715633ba-fe23-42a6-88b9-f2bbad4fb6ff · inbound
OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference tinyBenchmarks: evaluating LLMs with fewer examples
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c3bb8d-f922-48d4-b2f3-d4fc817460f9 · inbound
Efficient Evaluation of LLM Performance with Statistical Guarantees tinyBenchmarks: evaluating LLMs with fewer examples
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d118dd08-a9ed-402e-bd87-6793e5ecf684 · inbound
Learning More from Less: Unlocking Internal Representations for Benchmark Compression tinyBenchmarks: evaluating LLMs with fewer examples
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f43890e-7482-4390-8fa6-89024502c78e · inbound
Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees tinyBenchmarks: evaluating LLMs with fewer examples
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65884e4b-0b9f-4958-b4d9-11e5dc1f3602 · inbound
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents tinyBenchmarks: evaluating LLMs with fewer examples
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c4b1c7dd-001e-4c8a-9fb3-f6fe57fc1339 · inbound
ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation tinyBenchmarks: evaluating LLMs with fewer examples
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08b2782c-5c22-4e3e-b7b6-a4efd02b3113 · inbound
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment tinyBenchmarks: evaluating LLMs with fewer examples
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 542617dd-1eff-4d5a-8380-8b8024bf898e · inbound
Minimizing Collateral Damage in Activation Steering tinyBenchmarks: evaluating LLMs with fewer examples
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04920c39-ade0-40b9-91e2-44758f2239be · inbound
Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1aa01361-75d2-4f56-a12e-30cdb76ff703 · inbound
Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1abf399e-a176-4f39-b27d-8c6d06d05f65 · inbound
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation tinyBenchmarks: evaluating LLMs with fewer examples
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d25b9345-7f1f-403f-8cc2-6cdf67260e69 · inbound
ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression tinyBenchmarks: evaluating LLMs with fewer examples
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1544c1b7-2e17-4048-9fbe-8589d2ae9473 · inbound
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs tinyBenchmarks: evaluating LLMs with fewer examples
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e736fccb-b1dc-4c3d-b76b-6e31402ef5c6 · inbound
Validity Threats for Foundation Model Research tinyBenchmarks: evaluating LLMs with fewer examples
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0a3aa7fa-d48e-4be2-b8ba-7dae1bbc6520 · inbound
Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation tinyBenchmarks: evaluating LLMs with fewer examples
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c73fc89-5080-46de-8b55-9bc5acde3551 · inbound
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks tinyBenchmarks: evaluating LLMs with fewer examples
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67372468-6ce8-4916-909a-35b35b35fc27 · inbound
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results tinyBenchmarks: evaluating LLMs with fewer examples
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5324da0f-33cc-4946-a04a-de14cd5cb161 · inbound
MMGist: A Comprehensive Multimodal Benchmark for 2027 tinyBenchmarks: evaluating LLMs with fewer examples
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 384563cc-7ff5-4176-a42f-ded7d0652a5c · inbound
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2aa155b7-9468-4bbb-b8c3-46635b803cb7 · inbound
AGC-Bench: Measuring Artificial General Creativity tinyBenchmarks: evaluating LLMs with fewer examples
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54b9154a-cbb3-497a-b809-0d1e1887101a · inbound
AGC-Bench: Measuring Artificial General Creativity tinyBenchmarks: evaluating LLMs with fewer examples
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac233903-70d6-4f06-93c1-39640b514312 · inbound
OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters tinyBenchmarks: evaluating LLMs with fewer examples
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29a58174-185b-410b-bdb0-3459e2be244f · inbound
EduArt: An educational-level benchmark for evaluating art history knowledge in large language models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4503ec75-b14e-4b7c-b05f-561353a21df0 · inbound
Efficient Safety Alignment of Language Models via Latent Personality Traits tinyBenchmarks: evaluating LLMs with fewer examples
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5e8326e-8f62-4e2e-a43a-a5d02eaba7a3 · inbound
Prediction-Powered Active Testing tinyBenchmarks: evaluating LLMs with fewer examples
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1a20664-1446-4ebe-925c-e8ca3f0760c1 · inbound
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8d61ba-cc8f-4051-9fdf-a147ec13c4bb · inbound
Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator tinyBenchmarks: evaluating LLMs with fewer examples
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2146d11-490a-4c0c-acbc-e1f9e6672da9 · inbound
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer tinyBenchmarks: evaluating LLMs with fewer examples
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04007ee3-631a-47fc-b775-bda4630151e2 · inbound
Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset tinyBenchmarks: evaluating LLMs with fewer examples
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94f108e-e739-49d0-8054-c2a53b1c4712 · inbound
CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection tinyBenchmarks: evaluating LLMs with fewer examples
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.