Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2306.11270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.11270 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:30:13.021272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.677920Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ab441bd-c0a2-4bea-be64-65678e75e439 · inbound

Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding cites this paper.

Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T13:57:29.722915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:57:29.722915Z digest=sha256:293dd372e77f2e9adb746209d568cf07cc637b0065feaeff5bd4678c6786d739

Observation 7efef2b2-85b0-4390-957a-33ec2e23c547 · inbound

Beyond Prompt Content: Enhancing LLM Performance via Content-Format Integrated Prompt Optimization cites this paper.

Beyond Prompt Content: Enhancing LLM Performance via Content-Format Integrated Prompt Optimization Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T22:54:48.347005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:54:48.347005Z digest=sha256:1e4555ca91117e2108ed674fbfe39a1bbad16f420904e2e4235ac6876f168ef8

Observation 21ace7bd-d718-4556-80b4-7b331f75d834 · inbound

Leveraging LLM Inconsistency to Boost Pass@k Performance cites this paper.

Leveraging LLM Inconsistency to Boost Pass@k Performance Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:30:13.021272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:30:13.021272Z digest=sha256:53c3e55880217374c0107abcaffcfc18dcc58832093acbc50842811e36b7c5bf

Observation 99af6741-ddb7-48db-8fce-5ec9a5da0b7e · inbound

Investigating the Robustness of Retrieval-Augmented Generation at the Query Level cites this paper.

Investigating the Robustness of Retrieval-Augmented Generation at the Query Level Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:54.909052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:55:54.909052Z digest=sha256:dd52aa536ee2781b6c821a15b51c254833ab1d2847256f147735ca978b5ffc32

Observation 65904cfc-6982-4dd2-b3c3-78d6f22486b2 · inbound

PlotTwist: A Creative Plot Generation Framework with Small Language Models cites this paper.

PlotTwist: A Creative Plot Generation Framework with Small Language Models Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:40.507466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:40.507466Z digest=sha256:affb31744bd4e5e291076d9eef1f12de6faf581c5e2046af617872a26355fd30

Observation 7443e39b-e8e7-4138-9734-159fa2c6c416 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:56:27.679391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:1bca11ca93fdd2d601babbf2883fece4e629fb434a2e0bfb59c8c4619c9e900a

Observation bb2d9f26-1830-4dd4-af72-4cb87a97d3f6 · inbound

Towards Context-Invariant Safety Alignment for Large Language Models cites this paper.

Towards Context-Invariant Safety Alignment for Large Language Models Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:39:40.890191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-21T05:36:12.562807Z digest=sha256:d685f2fd79175cd60196b5977b86d12bc9c86fcc002be0e96ff3026c5bad9085

Observation e58e76aa-9eb0-4a23-ad73-5cbda748458f · inbound

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs cites this paper.

Same Patient, Different Words, Different Diagnosis? Evaluating Semantic Stability in Clinical LLMs Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:13.274114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:14:22.444020Z digest=sha256:812039f4e0d9e4daa10d22efee73b7099af1a3ac5cfc77cb1bf31fe901f0a0b3

Observation c103f8f5-6aaf-4a19-8e22-f6956c057a90 · inbound

Prompt Perturbation for Reliable LLM Evaluation over Comparison Graphs cites this paper.

Prompt Perturbation for Reliable LLM Evaluation over Comparison Graphs Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:58:57.679681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T01:03:20.261268Z digest=sha256:75ebb08893abfc2a41765db1615546297618ebad65edb847f5cda23943a174b2

Observation 6d28b749-88ac-4093-a4af-38eab487840a · inbound

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise cites this paper.

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T01:10:52.578408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:10:52.578408Z digest=sha256:ed7a2638a14a2d9de40be8ff235a7885b5cf2f9d5be46c2cf4eeaef1fac2f3a9

Observation 688187b5-de1e-443f-b59d-069303e3f4ca · inbound

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise cites this paper.

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise Evaluating the Zero-shot Robustness of Instruction-tuned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T07:04:18.645464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:04:18.645464Z digest=sha256:785bd2cecf7478c9beab2515fa1007458d45f4d5ed626b7c45121fa92a2f5119