Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:29.631887Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.09351.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:29.631887Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a7faa10d-4099-41b5-a650-21c53b820358 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 267c094b-32aa-4078-97f7-149f5ed65537 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274cfb02-528f-42a0-9697-9dcd0ab32c86 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6fa001-a216-4e73-9165-179998869704 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Exploring LLM Reasoning Through Controlled Prompt Variations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e28a2820-d83a-4870-a3e4-ad5d68cb7093 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Teaching Large Language Models to Self-Debug
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dba996a4-ac77-43ba-b281-90b9f444d1e3 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute No Language Left Behind: Scaling Human-Centered Machine Translation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138c76ef-b0c1-414d-9ba9-ae9aaba98965 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Frustratingly Easy Test-Time Adaptation of Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6734f8-74a2-4378-b1f6-cc53c7332200 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a6e6da89-ec2d-4ea2-8465-ff9cf098849e · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Measuring Massive Multitask Language Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b7c528a-9b4b-4149-94d9-e81195fc1496 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c73951-a3d0-45ea-9c3a-2043d58f2330 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Calibrating language models via augmented prompt ensembles
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cff92cfd-bdf2-47bf-a7a1-fc5c5e406a62 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Test-time Augmentation for Factual Probing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4bd67d26-faff-47b6-bb9f-2fcad9165c9d · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Improved Text Classification via Test-Time Augmentation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6f71437-e619-4f1d-afcf-889046fc80fb · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute ReFT: Reasoning with Reinforced Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bbf7138-9f82-4521-ba43-69ffb2f84fcf · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Humanity's Last Exam
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71adee13-e605-4423-b1c8-3a5ba9bb01dc · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Boosted Prompt Ensembles for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925c9d0a-1c63-4e28-9703-13a0e88c3936 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e0924ec-8f7b-4c4e-99a4-539415e9c089 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ccfa0ef-9c32-4dcb-bb91-eb9c8f7731fa · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Better Aggregation in Test-Time Augmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4a00aed-f634-41ce-9bf7-0cc24d88e6c7 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8e096ca-cdad-4c11-a066-6d3d0b38de34 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa0e2dae-a3e0-4897-ad3b-cf91d50542d9 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5391658d-cc28-48e7-8402-61fa8f74db9d · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 999d48c4-4613-4f7c-8c55-3b4b207a152b · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74699a4a-1fd0-4095-98af-9a2eed0326d4 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99260707-88f6-4fe3-86fc-3d63e485bc8a · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21afcc6d-3ff9-4e5e-a30c-b57482d666d9 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40454ea8-570d-4c8b-97ec-ca34c0314d78 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f0135a-0377-40f5-bc64-063830293537 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute output as array
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0139a27c-89ff-4c33-9b54-781123c46005 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9acfdb9f-c477-417a-b29c-08e840dc63b5 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8d72bf-8bd1-451c-bc23-a2fae51aa21b · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 966d0eb5-9ada-45fe-b15f-ed69797d3ebc · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af04e517-c00e-417e-9224-a77163cee234 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64b5ec6-e754-426b-a1af-6a6b80e2f5e4 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 84b3b5fa-2ca9-4e13-b78f-4a9887b3e700 · outbound
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.