Pith. sign in

Paper Citation Record · LEDGER

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models

As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2412.15524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15524 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:25:14.799196Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7706cd94-1b05-435c-b286-190ff7bc001b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.703507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.703507Z digest=sha256:bbbb5948300c48365e8baaa55b2f7648fd1fbbcfecb803dbd12509f06e39eef5

Observation c881e4d1-d6f0-4a63-85f5-1a841c3a9640 · outbound

This paper cites GPT-4 Technical Report.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.706680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.706680Z digest=sha256:6f0ce01454b3e777efff6b8bcd19350e63c5b2a5c86bd5823a60ea3fb7e3c135

Observation bf12983d-0afe-48b6-aa6b-e7c7a769ba9d · outbound

This paper cites Qwen Technical Report.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.709402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.709402Z digest=sha256:1d662ea4437cefe20c8b2cad18314d410044c5c3fc13344756352d114b2da314

Observation 031d8da6-90bb-47b1-bb8a-c01afa940d30 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.712008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.712008Z digest=sha256:27cc6e3f24380d66dff2ca65862e9bd54a4a38b68487fec12d0a2fa857c43c06

Observation cad45276-edd1-49ee-8e1e-b7ba18e741a8 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.714266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.714266Z digest=sha256:2e97fbebe3872828d0a7a1c8e7b8d2356306e12fc9144782f790f008dd2b7a2f

Observation a7184216-70a3-482c-84be-7c9953d09c0e · outbound

This paper cites Language Models are Few-Shot Learners.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.716475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.716475Z digest=sha256:05082913a1ac6c5e1bdb062b35e526d0f8033a3dbecad217898e5b98ee6eb5b7

Observation ad41664b-7b22-43df-8473-1e4923bf080f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.719144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.719144Z digest=sha256:d0506078964954357cdcec0bbedd41aeca4f7183295d7bf7e8a419f41d7106e1

Observation f0bc3b7b-27ee-4374-8938-66702cf5f376 · outbound

This paper cites Gonzalez, and Ion Stoica.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, and Ion Stoica

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.721166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.721166Z digest=sha256:bd399e9628e5ed59c75496850c7d7416159ab4a94def53883b5ea367bedb08c8

Observation bab4b4d1-647b-4872-b176-3c2fa75077ab · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.173850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.723180Z digest=sha256:319c15b9ebf5a235758f6cef4791d75bcbce06f3f5d84989aa5dbd4a6098541e

Observation 789a0590-c104-4f09-9d69-a9a27dc3041f · outbound

This paper cites The Llama 3 Herd of Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.725327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.725327Z digest=sha256:72d6354b7a66c7f962cd07a5ff30af76d1a1d3242959eab3c390c99e5b71bfc9

Observation dbb5bae7-83ed-47af-a34e-4ddb375f7903 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.729721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.729721Z digest=sha256:70d5d12cd00fd2bcef03fa083f2b7d4a841fe4269e7d678f1cd08d2f81c39222

Observation b40a680b-4166-4a2e-973b-ba933a162bf3 · outbound

This paper cites Prolific first, 2014.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Prolific first, 2014

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.166275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.732761Z digest=sha256:52ae33692655b68278849c24eead54f2b2666e4926157286e756be8903bf8717

Observation e8f6d6a5-2a32-4ffe-ab85-ad6c5c1dc123 · outbound

This paper cites Koala: A dialogue model for academic research.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Koala: A dialogue model for academic research

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.735127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.735127Z digest=sha256:41f51e4040e04ee2659baf16ffd774830c8a1ec2c9b5dced4726bbf88a4e22ea

Observation c61b52af-9bf4-40e9-8c4f-45cac6a4323e · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models OLMo: Accelerating the Science of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.737679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.737679Z digest=sha256:04e869b48ea91403d779163877b775713d6e4eca608d4bdbd10df6fb0df8de78

Observation 5acb2964-ccae-4b79-b572-bf6f503148cf · outbound

This paper cites Mistral 7B.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.740531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.740531Z digest=sha256:c6f8f46543f2b0a73a979e33f209db616d3fd01b6ba17777dcb5577a5606adfe

Observation e7031387-07ef-4e83-8c48-a1284c4773d8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.743211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.743211Z digest=sha256:93a8011d84da036e97135ab795ab66d2199d70bb5b39e544ba53d5b46afcc8d8

Observation 20da4237-8daa-47d7-b4a6-bda86fc023f9 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.745975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.745975Z digest=sha256:3fd0dbdacbbcc4c381e1192d8adc86647d278d62938725624720067d42c6b0a0

Observation 572ca1ff-ee8c-4166-bf5b-4459465408c6 · outbound

This paper cites Hashimoto.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.748830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.748830Z digest=sha256:bd185ec17c0611ece2c111091de709ec7f955163c6a0f43abb36c65efa33a5b7

Observation 41f67fbc-9b0f-4496-9aed-7946f7789fc2 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.751441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.751441Z digest=sha256:4da470157e1415194ab54b40f3603e08de100cfedfd90ddce1a3bbc867d05a39

Observation ed911785-929d-4109-8c39-72eb30362baf · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rouge: A package for automatic evaluation of summaries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.754255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.754255Z digest=sha256:e09fd40a2700074a119d1b0503f82036ecab52c15f90e5902450af030fd12709

Observation 36f8f95d-5a6e-4b0a-9224-0cc542b41a37 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.756813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.756813Z digest=sha256:911ec9250c9089adbb6c4c53f35e5f3f546df49b1c61488ca93dd4efed55ec3c

Observation e061115e-7682-4609-af9e-a0d808278e08 · outbound

This paper cites Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.759366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.759366Z digest=sha256:df60c1f51e8c74690b3e5162de506d9474260eb10616e3b01e0152ad6e2b0173

Observation 992f0be7-0a2a-4244-b48a-5a6aa77b3024 · outbound

This paper cites Training language models to follow instructions with human feedback.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.762244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.762244Z digest=sha256:31375ec183b122b9dedab26549aa4aee695d3265af1ade0bb5eaa8758933db67

Observation 0a8d359b-4a7c-4894-b8c0-a6f3b0a20d75 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Bleu: a method for automatic evaluation of machine translation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.764678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.764678Z digest=sha256:f3e139280d66e040e8dec352c4740505331eb4355073877322e813870b4200d3

Observation 2cfb50f5-8bd9-4199-974f-3f991d7c63ec · outbound

This paper cites Instruction Tuning with GPT-4.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.767175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.767175Z digest=sha256:4002fb25a0ef2782f86b6c44b9d52a95233e912a87ee8e30a0739b92af2d7970

Observation 338f78d9-b7f2-4e35-b073-b1b18bbb01da · outbound

This paper cites Rush, and Thomas Wolf.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rush, and Thomas Wolf

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.139419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.769774Z digest=sha256:3673ad470679ce93baa0ee4fb356223489d475286dbda67206ce4b0052069f75

Observation dd800184-77d9-4d0a-ae27-631d3eb0e644 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.773071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.773071Z digest=sha256:589b1c985123b03ed04b932442556f38abe640720219872abca4dcf8ff393f22

Observation 5cae3dbb-297c-4350-bb5e-e42e7a5e23d8 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.775460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.775460Z digest=sha256:9c54444064a0c86e7972a0dd188de253a284487fc5b7dc4ad6b056eec9f804b2

Observation 23fba9b4-7939-44c8-ac85-a20c49a25d39 · outbound

This paper cites Wizard LM : Empowering large pre-trained language models to follow complex instructions.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Wizard LM : Empowering large pre-trained language models to follow complex instructions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.777734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.777734Z digest=sha256:9b43f80a6ad2e4be4725ca947e2b69397d803a746056437caa6a4697e335d1b6

Observation 3f726c6b-5af5-4628-8a61-c223d586a93e · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Yi: Open Foundation Models by 01.AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.780280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.780280Z digest=sha256:c5391baf727497a6037fa1aa6811b025c18a226a9d6db859bf9663da7a2bc2e8

Observation 78d5f30d-0b37-4831-8052-de1308fed6ba · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models BERTScore: Evaluating Text Generation with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.782701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.782701Z digest=sha256:2cc374b7340dd24aa85bfeb0b9e3c12bb23c51aa06ddb416d989b1e15af141f9

Observation f7bbac03-8951-477a-8718-8faaf4e2e05d · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.785234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.785234Z digest=sha256:3c7c42a737e9a8578ef62ddff835b6c2ff1b1c7e87c50ab6c6a08379c36a57f7

Observation 191909aa-9571-4ddf-a930-03927830be7e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.787566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.787566Z digest=sha256:fb49b645cd4ab9c66fccd8cbbddf843605f81b7e3df294872b144cf42a256a8f

Observation 6f5bd239-60f3-4dc1-a9f7-fec408cad0e7 · outbound

This paper cites Lima: Less is more for alignment.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Lima: Less is more for alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.789520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.789520Z digest=sha256:82e674709e0f8d761f9b95fcb2ecc328e4ce96d29ca40f68915ae639b1c00858

Observation 8705ba5c-2c97-44c5-8f55-c37a6eab485f · outbound

This paper cites write newline.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.791614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.791614Z digest=sha256:ef3a910dcd71a94d61c396b79cede0f88a9a1f36e7137231a93ddf3e62537b56

Observation 6d59554a-79f9-4417-9fc0-d1133d89f1b1 · outbound

This paper cites @esa (Ref.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models @esa (Ref

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.794326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.794326Z digest=sha256:4d476208ea12d42b1241a6d944b28a8bb19237dadb8d7ffdba1b7e4a887502d4

Observation 776f9b1b-3835-48ef-b18d-8f9d18609a68 · outbound

This paper cites an unresolved cited work.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.796665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.796665Z digest=sha256:6e428b42077da1a8083e425b87b97eecf5ee0358fa07c299f505161174472d1a

Observation e33ad84e-b466-403e-b816-10e04b97dc4a · outbound

This paper cites Hide and Seek.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hide and Seek

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T11:25:14.799196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.799196Z digest=sha256:af7547a862ff405646f28589ec0a18456a8a9984e2d5cb0a344fc1f7b8d4aada

Pith citing papers

No inbound Pith citation observations are available.