Pith. sign in

Paper Citation Record · LEDGER

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2412.15524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15524 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:25:14.799196Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7706cd94-1b05-435c-b286-190ff7bc001b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.703507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.703507Z digest=sha256:175a3b2561321d406c6875aa715ab83a1b288d1cbb02b9d20aff093b76b47d7c

Observation c881e4d1-d6f0-4a63-85f5-1a841c3a9640 · outbound

This paper cites GPT-4 Technical Report.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.706680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.706680Z digest=sha256:82f5f5509187fd457367d907ce5344217ac8567b7d513a0ace021fe1df5445cd

Observation bf12983d-0afe-48b6-aa6b-e7c7a769ba9d · outbound

This paper cites Qwen Technical Report.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.709402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.709402Z digest=sha256:4eb1a37786b20d71899369cf61b9268c2c035c2f38a3e420224396f50a93f845

Observation 031d8da6-90bb-47b1-bb8a-c01afa940d30 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.712008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.712008Z digest=sha256:fc1ff116144e7003c584b66f07a7beeb0b136981970695880e0d195c913abf79

Observation cad45276-edd1-49ee-8e1e-b7ba18e741a8 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.714266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.714266Z digest=sha256:5b776a7aca23ab12b27517e944e4c5c94b58197a10c6bee2872f2224b47353b0

Observation a7184216-70a3-482c-84be-7c9953d09c0e · outbound

This paper cites Language Models are Few-Shot Learners.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Language Models are Few-Shot Learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.716475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.716475Z digest=sha256:f75d6bdb6086fd016df5fc5eae63c7199bdafeb5cc468540786a55a569983882

Observation ad41664b-7b22-43df-8473-1e4923bf080f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.719144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.719144Z digest=sha256:5e27704182bbc1a289d060bc6df742c6e669592a4eb7f47f17864610edc1c298

Observation f0bc3b7b-27ee-4374-8938-66702cf5f376 · outbound

This paper cites Gonzalez, and Ion Stoica.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, and Ion Stoica

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.721166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.721166Z digest=sha256:066234367e8aeaf435e0f26d924ac4a1db1f4cd05c99a4b6ff3eb507a6d690f1

Observation bab4b4d1-647b-4872-b176-3c2fa75077ab · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.173850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.723180Z digest=sha256:7cd1f65e375657f5c1e33afd2bf2d81b0d7a8f67e870c9713ca09894da1b1bc0

Observation 789a0590-c104-4f09-9d69-a9a27dc3041f · outbound

This paper cites The Llama 3 Herd of Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.725327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.725327Z digest=sha256:63b75e1198b1963b31b38e536c3c0c5550ab4724f03123751d800c3a58ea3761

Observation dbb5bae7-83ed-47af-a34e-4ddb375f7903 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.729721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.729721Z digest=sha256:defcce0b0fdf650e1447d45df138a9ab979a94803057408ade68e550f6a0d2e0

Observation b40a680b-4166-4a2e-973b-ba933a162bf3 · outbound

This paper cites Prolific first, 2014.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Prolific first, 2014

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.166275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.732761Z digest=sha256:f0471444f78e0e7b2812f20622303e655f92dd3099817075a5f68e00b611da00

Observation e8f6d6a5-2a32-4ffe-ab85-ad6c5c1dc123 · outbound

This paper cites Koala: A dialogue model for academic research.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Koala: A dialogue model for academic research

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.735127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.735127Z digest=sha256:0e681c12bc6f42a3044aa0936b2daf17f4fe16335e05c84f22d804343ee982c5

Observation c61b52af-9bf4-40e9-8c4f-45cac6a4323e · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models OLMo: Accelerating the Science of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.737679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.737679Z digest=sha256:ae2079dff52e493ff28e5e0a83bbef08478ec340551e3ecb1852d008be7d85cc

Observation 5acb2964-ccae-4b79-b572-bf6f503148cf · outbound

This paper cites Mistral 7B.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.740531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.740531Z digest=sha256:cb97c8966bf39f91ea33815b41d890372aa59a7e98f047e82bcd41f37d381fed

Observation e7031387-07ef-4e83-8c48-a1284c4773d8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.743211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.743211Z digest=sha256:1140f89d50c1ea85881e8d69d00b22ee17782b9b431f4d9afba2c49acc5aab65

Observation 20da4237-8daa-47d7-b4a6-bda86fc023f9 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.745975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.745975Z digest=sha256:da77c6a92da62e39f97c90c39d3e4d7a86b9173ce4c59f3920e9e00e656b09f6

Observation 572ca1ff-ee8c-4166-bf5b-4459465408c6 · outbound

This paper cites Hashimoto.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.748830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.748830Z digest=sha256:08cc29d47cfad83cf1a9b88bc98d12e4e357199b28851aed02bcdf9aac41102a

Observation 41f67fbc-9b0f-4496-9aed-7946f7789fc2 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.751441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.751441Z digest=sha256:52220f4f973c6a494caafe445e1b471c1f7105ce7ce0e2e9a9b7d8a3052c53f0

Observation ed911785-929d-4109-8c39-72eb30362baf · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rouge: A package for automatic evaluation of summaries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.754255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.754255Z digest=sha256:35a96139436b57b11d23a3374bcb0b0be94f3ba3d466fcadd24be38a48c19604

Observation 36f8f95d-5a6e-4b0a-9224-0cc542b41a37 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.756813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.756813Z digest=sha256:45f1aeb53ec7963651801b70431a7d8ac3ac9e151efadbaa45708acf04f7674d

Observation e061115e-7682-4609-af9e-a0d808278e08 · outbound

This paper cites Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.759366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.759366Z digest=sha256:c7a6bfd2d18e7e6f1a53d1b7095512f1b7bd53778fd129e471f61e4cfa13ee8d

Observation 992f0be7-0a2a-4244-b48a-5a6aa77b3024 · outbound

This paper cites Training language models to follow instructions with human feedback.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.762244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.762244Z digest=sha256:f8e5fb2d76a4ffb0f3156f9a7d92aef34d43a345c28eb26bc7cdea660c64de5e

Observation 0a8d359b-4a7c-4894-b8c0-a6f3b0a20d75 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Bleu: a method for automatic evaluation of machine translation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.764678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.764678Z digest=sha256:d6ac37e89ed27ac6dccd4cbb6007a0f9e5ece5e3521e6f0d91c37690a5a4817a

Observation 2cfb50f5-8bd9-4199-974f-3f991d7c63ec · outbound

This paper cites Instruction Tuning with GPT-4.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.767175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.767175Z digest=sha256:59e79d4b6d044b59dcfe8ade696d5cb104125ce0d387d7f872dd814191b25e70

Observation 338f78d9-b7f2-4e35-b073-b1b18bbb01da · outbound

This paper cites Rush, and Thomas Wolf.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rush, and Thomas Wolf

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:25:15.139419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T11:25:14.769774Z digest=sha256:fbf7a34c9935c77e0b2ad84d3c6247bb2617a68ad3e588968df8c6e293f1d323

Observation dd800184-77d9-4d0a-ae27-631d3eb0e644 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.773071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.773071Z digest=sha256:94f1df2306495f9c34b61211677f287c8ea863b7fa4d9c6139d71321e4af4d3d

Observation 5cae3dbb-297c-4350-bb5e-e42e7a5e23d8 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.775460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.775460Z digest=sha256:4c46d60284dece55dccab22db310f2133c0075b371fd058407caa5ed96b7be37

Observation 23fba9b4-7939-44c8-ac85-a20c49a25d39 · outbound

This paper cites Wizard LM : Empowering large pre-trained language models to follow complex instructions.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Wizard LM : Empowering large pre-trained language models to follow complex instructions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.777734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.777734Z digest=sha256:b2026a7699a40668b2eeb6cc420eefe76a746358fc862cd432fc945c8770606f

Observation 3f726c6b-5af5-4628-8a61-c223d586a93e · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Yi: Open Foundation Models by 01.AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.780280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.780280Z digest=sha256:7ced27e9718aa49e3b6f4fddf17d544cc6cab158476bfd6f0fad817367df0e0c

Observation 78d5f30d-0b37-4831-8052-de1308fed6ba · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models BERTScore: Evaluating Text Generation with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.782701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.782701Z digest=sha256:d98e9601016c990e05fe47e87974ab2ace04f13167a09a9bc5d94804962d4b12

Observation f7bbac03-8951-477a-8718-8faaf4e2e05d · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.785234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.785234Z digest=sha256:b7a38f9417258cba27d173b281011c444bf9ac33f4b5398c12d96b49ef492f8f

Observation 191909aa-9571-4ddf-a930-03927830be7e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.787566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.787566Z digest=sha256:decaa0f0ac8e4fd445a0df042bd16b7c95c9b222bd40f48bad0dd33f58bfed19

Observation 6f5bd239-60f3-4dc1-a9f7-fec408cad0e7 · outbound

This paper cites Lima: Less is more for alignment.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Lima: Less is more for alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.789520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.789520Z digest=sha256:569051850cfe1b5e3745b12b86a66e6814f6dde101c0e3743ad02fc812c5ceda

Observation 8705ba5c-2c97-44c5-8f55-c37a6eab485f · outbound

This paper cites write newline.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models write newline

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.791614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.791614Z digest=sha256:29635e573c7d4b0e24ba5019cb9706a2df2c9b9b0910db5d03fdc1af063048e5

Observation 6d59554a-79f9-4417-9fc0-d1133d89f1b1 · outbound

This paper cites @esa (Ref.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models @esa (Ref

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.794326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.794326Z digest=sha256:5c9d3c1412cab41fcd267add93978473ffbde040aa4644adcb1c39a156c20398

Observation 776f9b1b-3835-48ef-b18d-8f9d18609a68 · outbound

This paper cites an unresolved cited work.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:25:14.796665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.796665Z digest=sha256:16031bed6cd54bb290ad469ecdb7d069b3d40dfdfd3dc10c810df20124bd2442

Observation e33ad84e-b466-403e-b816-10e04b97dc4a · outbound

This paper cites Hide and Seek.

HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hide and Seek

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T11:25:14.799196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:25:14.799196Z digest=sha256:722209feb811b6d0d534299b5586037d3a281d6774c4f173ec2cff41b75bd99f

Pith citing papers

No inbound Pith citation observations are available.