Pith. sign in

Paper Citation Record · LEDGER

tinyBenchmarks: evaluating LLMs with fewer examples

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2402.14992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14992 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T02:32:05.153277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23c0e894-68e5-428f-b2d2-5cecdb0b491d · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction tinyBenchmarks: evaluating LLMs with fewer examples

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:47:56.080794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:eeeb7f21daef9843bf1612a740a1fc988859be56ce6756ec553e8db3c8075d87

Observation a858ae38-1d42-40f8-a4b0-022b15c8c58e · inbound

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models cites this paper.

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:19:22.473153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T05:19:22.423762Z digest=sha256:ff9d448d1e108c8dc800d5c9dabcb783532390ba0150629b5d26c80a13214838

Observation 3cdd8a53-8bd8-4ec4-bb25-68b8e5613c90 · inbound

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation cites this paper.

Beyond the Singular: Revealing the Value of Multiple Generations in Benchmark Evaluation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:45:21.881227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T03:42:30.220989Z digest=sha256:42d731475f9c01a2e70bd30d043a435334b4c026d1fe7a27db96ebc42b3f7313

Observation 96a60ddf-57d1-4ac5-8a55-b1319e00fd58 · inbound

LLM-Safety Evaluations Lack Robustness cites this paper.

LLM-Safety Evaluations Lack Robustness tinyBenchmarks: evaluating LLMs with fewer examples

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T01:27:21.237513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:26:45.402983Z digest=sha256:4ce7cd46f9b99f3886b71a3b0b8cc1810751fa6bb5389ee6c780ab2ec4722257

Observation 1cde4c61-88c4-4323-91dd-789a8c441a12 · inbound

Small Language Models are the Future of Agentic AI cites this paper.

Small Language Models are the Future of Agentic AI tinyBenchmarks: evaluating LLMs with fewer examples

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:55:50.972225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T11:55:50.897500Z digest=sha256:02e1539025a24e6de0a6c88b1368c6bec834a241d2e7520daad35f373c4788bd

Observation 28acc0ee-d2c0-4e64-a7ce-0c9b8f526a57 · inbound

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users cites this paper.

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users tinyBenchmarks: evaluating LLMs with fewer examples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T06:02:07.842309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:58:17.452837Z digest=sha256:bda8a19529a822118f8f0cd687576239cd10a88b6733cdc3d7f762fc18b97791

Observation 3687ec11-4f3e-44fc-830e-ae480d056225 · inbound

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks cites this paper.

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:16:24.095399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T13:16:19.744864Z digest=sha256:614669b413e631a53874bb12eac98c316238bd35a654137e5ad6b50830813546

Observation 48a3726e-5ed3-4e65-b1cb-d09fa4b169ce · inbound

Activation Steering with a Feedback Controller cites this paper.

Activation Steering with a Feedback Controller tinyBenchmarks: evaluating LLMs with fewer examples

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:50:41.340553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T21:50:05.186376Z digest=sha256:7bf3352777747649f55b7416c9745f7fe1d0d1d30b6f85460044a5c01ed55e75

Observation 715633ba-fe23-42a6-88b9-f2bbad4fb6ff · inbound

OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference cites this paper.

OI-Bench: An Option Injection Benchmark for Evaluating LLM Susceptibility to Directive Interference tinyBenchmarks: evaluating LLMs with fewer examples

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T09:39:13.877212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:39:13.877212Z digest=sha256:5c969b1ba1de1e2ca5b9638db28d45006c32e650a4f446d1586f6d1333cba0dd

Observation f8c3bb8d-f922-48d4-b2f3-d4fc817460f9 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees tinyBenchmarks: evaluating LLMs with fewer examples

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:00:51.726570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:708d32c6531f02d7829c7019d6527f35223c654a56db979b011f8783bedb730d

Observation d118dd08-a9ed-402e-bd87-6793e5ecf684 · inbound

Learning More from Less: Unlocking Internal Representations for Benchmark Compression cites this paper.

Learning More from Less: Unlocking Internal Representations for Benchmark Compression tinyBenchmarks: evaluating LLMs with fewer examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:58.615389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:58.615389Z digest=sha256:708e370c68ea3b4fe12baf9c024403392f3eca056001ab1610672c45c75fbe03

Observation 4f43890e-7482-4390-8fa6-89024502c78e · inbound

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees cites this paper.

Select Smarter, Not More: Prompt-Aware Evaluation Scheduling with Submodular Guarantees tinyBenchmarks: evaluating LLMs with fewer examples

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:50:59.586582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:47:45.677717Z digest=sha256:6cacf5435c5430f601fd45b98289337bebcef5dab88adfc4d41315dadd75b897

Observation 65884e4b-0b9f-4958-b4d9-11e5dc1f3602 · inbound

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents cites this paper.

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents tinyBenchmarks: evaluating LLMs with fewer examples

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:56:31.186666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T04:25:27.381142Z digest=sha256:d103220983393fe34e1f8148cfcf460be3baf18955888f10fc8087ea6e7f03fb

Observation c4b1c7dd-001e-4c8a-9fb3-f6fe57fc1339 · inbound

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation cites this paper.

ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:10.344342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T08:21:55.648930Z digest=sha256:6fb44ece2a001a6c3198013004fad58487e7fb767302355353b039c55bfe6ce7

Observation 08b2782c-5c22-4e3e-b7b6-a4efd02b3113 · inbound

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment cites this paper.

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment tinyBenchmarks: evaluating LLMs with fewer examples

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:51:09.782572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:50:05.920842Z digest=sha256:aa14ba513089a0469f51815332351815f697678bebc828694fe6f4d3b3bfff6e

Observation 542617dd-1eff-4d5a-8380-8b8024bf898e · inbound

Minimizing Collateral Damage in Activation Steering cites this paper.

Minimizing Collateral Damage in Activation Steering tinyBenchmarks: evaluating LLMs with fewer examples

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:56:48.964867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T18:58:30.056670Z digest=sha256:f1f9159bac41ed4ff199473ae2b2b50c99763891305d30f8ed394868b4a98cce

Observation 04920c39-ade0-40b9-91e2-44758f2239be · inbound

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models cites this paper.

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:13.759114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T10:13:35.777910Z digest=sha256:a5caba2889356923d428fb044c7fa6d3cec28fb3654fff47d8a1f6dfda6fb03c

Observation 1aa01361-75d2-4f56-a12e-30cdb76ff703 · inbound

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models cites this paper.

Jailbreak susceptibility prediction and mitigation via the behavioral geometry of models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:47.703641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T17:43:47.849960Z digest=sha256:573f37b9a16b1dc8889b98f6e54a34f9f463988138b9706bb1e029ff9df4d636

Observation 1abf399e-a176-4f39-b27d-8c6d06d05f65 · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:52:49.484995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:d8239793f01c2eb134abca5339566f8d09daf83a19a1199e13dcb3baad3d8279

Observation d25b9345-7f1f-403f-8cc2-6cdf67260e69 · inbound

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression cites this paper.

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression tinyBenchmarks: evaluating LLMs with fewer examples

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:12:34.475288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T19:10:58.845006Z digest=sha256:9f7683eae768799eb270429c451cee1d267afc1da6999f29d8fa0f0c2fc8e04e

Observation 1544c1b7-2e17-4048-9fbe-8589d2ae9473 · inbound

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs cites this paper.

Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs tinyBenchmarks: evaluating LLMs with fewer examples

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:02:24.486335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T16:56:29.536912Z digest=sha256:fd8d570f5354907149506e54da10cde228e08a38c5d122e7b0b1bd16f3f69424

Observation e736fccb-b1dc-4c3d-b76b-6e31402ef5c6 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research tinyBenchmarks: evaluating LLMs with fewer examples

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:36:44.859788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:4af89fbbb01bd7ad7ba27afedfa783accfb5c220d73616e0c1784352c6cd5abf

Observation 0a3aa7fa-d48e-4be2-b8ba-7dae1bbc6520 · inbound

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation cites this paper.

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:52:45.592125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:48:20.193699Z digest=sha256:5666b2daedd511e1a5b4d3adf0269293729276914b028466c2c0f960dc181037

Observation 5c73fc89-5080-46de-8b55-9bc5acde3551 · inbound

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks cites this paper.

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks tinyBenchmarks: evaluating LLMs with fewer examples

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.635188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:39:51.630229Z digest=sha256:41f4388f3f544d41bf0f031bd3b83bb48a220c6911a794dc8cab7014bef92a83

Observation 67372468-6ce8-4916-909a-35b35b35fc27 · inbound

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results cites this paper.

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results tinyBenchmarks: evaluating LLMs with fewer examples

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.293701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T04:45:20.445703Z digest=sha256:da06dc7f656bfb09f0dafa5093f71849d7fc3cd2117b4850543074ec5a82924c

Observation 5324da0f-33cc-4946-a04a-de14cd5cb161 · inbound

MMGist: A Comprehensive Multimodal Benchmark for 2027 cites this paper.

MMGist: A Comprehensive Multimodal Benchmark for 2027 tinyBenchmarks: evaluating LLMs with fewer examples

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:49:41.646194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:05:14.573386Z digest=sha256:5c6e8b247cee7b0423005b6dbeb522bcea947b3e42e420abce63ea8d57a95ca3

Observation 384563cc-7ff5-4176-a42f-ded7d0652a5c · inbound

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models cites this paper.

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:07.541242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T19:58:23.594907Z digest=sha256:4e8acd23a19e1b0747765411817ee695f9ae50a04859dd45de745760d7aa1d92

Observation 2aa155b7-9468-4bbb-b8c3-46635b803cb7 · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity tinyBenchmarks: evaluating LLMs with fewer examples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:56.060287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T12:33:53.029578Z digest=sha256:db4d995644c2142589ab4d24798cd980a883053519600f1b72228acace8e918b

Observation 54b9154a-cbb3-497a-b809-0d1e1887101a · inbound

AGC-Bench: Measuring Artificial General Creativity cites this paper.

AGC-Bench: Measuring Artificial General Creativity tinyBenchmarks: evaluating LLMs with fewer examples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.086500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T21:25:14.030920Z digest=sha256:1213cb7d2e9ffa698b08b57c74dabda36eb4e176aada9776073119a862692efe

Observation ac233903-70d6-4f06-93c1-39640b514312 · inbound

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters cites this paper.

OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters tinyBenchmarks: evaluating LLMs with fewer examples

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T06:37:42.188290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-03T06:30:30.308713Z digest=sha256:b99c3d2f2eed6e5993477e5aa81085d31fc77f2b011d5fad48ba2798c573de1e

Observation 29a58174-185b-410b-bdb0-3459e2be244f · inbound

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models cites this paper.

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models tinyBenchmarks: evaluating LLMs with fewer examples

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.254103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T14:48:06.514693Z digest=sha256:fa59d5e58c61d159dcb36e32064da5d88e0c5c927d8d371c1646c29f5b5a8e3a

Observation 4503ec75-b14e-4b7c-b05f-561353a21df0 · inbound

Efficient Safety Alignment of Language Models via Latent Personality Traits cites this paper.

Efficient Safety Alignment of Language Models via Latent Personality Traits tinyBenchmarks: evaluating LLMs with fewer examples

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T15:27:20.086993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-10T15:26:23.290009Z digest=sha256:e822e6403e4278b38e691814c9376a55df4e66ac4035962e9ff636e5bcb2419a

Observation f5e8326e-8f62-4e2e-a43a-a5d02eaba7a3 · inbound

Prediction-Powered Active Testing cites this paper.

Prediction-Powered Active Testing tinyBenchmarks: evaluating LLMs with fewer examples

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T09:06:59.145174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-10T08:59:10.780536Z digest=sha256:06c59d5ec48fefe493fb16125a438e3e7d67da30279737eab631a71404a1b6c2

Observation d1a20664-1446-4ebe-925c-e8ca3f0760c1 · inbound

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks cites this paper.

Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks tinyBenchmarks: evaluating LLMs with fewer examples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T16:38:50.302371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:38:50.302371Z digest=sha256:0be3fc932e592142dc5daa6af90503f946115b939922a9bbc5274cb84c009521

Observation 1b8d61ba-cc8f-4051-9fdf-a147ec13c4bb · inbound

Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator cites this paper.

Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator tinyBenchmarks: evaluating LLMs with fewer examples

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T14:50:59.373155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:50:59.373155Z digest=sha256:ef931d60dc6e7cf7362891f10a2682e3fe3808c7fe8b5080807abb60812095c9

Observation d2146d11-490a-4c0c-acbc-e1f9e6672da9 · inbound

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer cites this paper.

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer tinyBenchmarks: evaluating LLMs with fewer examples

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T16:17:41.484228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T16:17:41.484228Z digest=sha256:3cce9616880934a4be451ee37614a6680fc4f55edb936cc875db6382b83dd875

Observation 04007ee3-631a-47fc-b775-bda4630151e2 · inbound

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset cites this paper.

Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset tinyBenchmarks: evaluating LLMs with fewer examples

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T07:40:14.904090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:40:14.904090Z digest=sha256:e451c228b94bee8818475fc9b6b7181ee094f7a9688af8f8be2a4b2209812151

Observation f94f108e-e739-49d0-8054-c2a53b1c4712 · inbound

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection cites this paper.

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection tinyBenchmarks: evaluating LLMs with fewer examples

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:05.153277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T02:32:05.153277Z digest=sha256:5cf3cf721f27352728d5ecaab7e2ef687ef7491e1371eb4c1a2b677aca5a4e57