Pith. sign in

Paper Citation Record · LEDGER

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.09351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09351 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:29.631887Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact7
  • verified fuzzy1
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7faa10d-4099-41b5-a650-21c53b820358 · outbound

This paper cites Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.532259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.532259Z digest=sha256:2c11b2795a0f400813907f2842a2111e34849525e0fcf19f920c1c163ece5e3d

Observation 267c094b-32aa-4078-97f7-149f5ed65537 · outbound

This paper cites The Surprising Effectiveness of Test-Time Training for Few-Shot Learning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.539084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.539084Z digest=sha256:5f4d76ec91cfb5a3a57a9b4c5ce3aa64b157e0215ad47237b7a72df96cad1f37

Observation 274cfb02-528f-42a0-9697-9dcd0ab32c86 · outbound

This paper cites Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.545090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.545090Z digest=sha256:9138c398252a878852b9244bc33b681424d3877f775e2a8f8c3587b54e52b40f

Observation 9e6fa001-a216-4e73-9165-179998869704 · outbound

This paper cites Exploring LLM Reasoning Through Controlled Prompt Variations.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Exploring LLM Reasoning Through Controlled Prompt Variations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.548200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.548200Z digest=sha256:aeea483cd24e869cdd0affc2379f313b197bb677a314ba21efc3ffd64483ceb9

Observation e28a2820-d83a-4870-a3e4-ad5d68cb7093 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Teaching Large Language Models to Self-Debug

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.551448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.551448Z digest=sha256:8c645d3e412d074bd167a777540ce63fe5f546e617ece667d97a827cf6d4f227

Observation dba996a4-ac77-43ba-b281-90b9f444d1e3 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.556706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.556706Z digest=sha256:ab27e99f012d4e41623f06ddb6c5d080fed8d0dc42da718028f6abcb60f7dceb

Observation 138c76ef-b0c1-414d-9ba9-ae9aaba98965 · outbound

This paper cites Frustratingly Easy Test-Time Adaptation of Vision-Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Frustratingly Easy Test-Time Adaptation of Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.562315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.562315Z digest=sha256:40cd6122b624b8360c22332e781fdd0cf7db5ea5f23b2f5606bbc52df3cd0f15

Observation 1d6734f8-74a2-4378-b1f6-cc53c7332200 · outbound

This paper cites M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:25.243674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.565203Z digest=sha256:0fbdc4e108b5736765e4429f79b0cec83fc4322b6dbc4abe3441ef25c78c276b

Observation a6e6da89-ec2d-4ea2-8465-ff9cf098849e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.568253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.568253Z digest=sha256:29be4c52551e2ec0dd29fe187fc2bb9aed16328145e79e668f14feaf3759849e

Observation 3b7c528a-9b4b-4149-94d9-e81195fc1496 · outbound

This paper cites Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.573473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.573473Z digest=sha256:3fde5ae8aeb05f36daba85143cd33964c5e990b14a50cdd40671834bea375cbe

Observation b3c73951-a3d0-45ea-9c3a-2043d58f2330 · outbound

This paper cites Calibrating language models via augmented prompt ensembles.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Calibrating language models via augmented prompt ensembles

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:40.437350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.576659Z digest=sha256:9319834f29d6eda152c351b2649fb910894f964f2982a899f2d4f91746bbc0cb

Observation cff92cfd-bdf2-47bf-a7a1-fc5c5e406a62 · outbound

This paper cites Test-time Augmentation for Factual Probing.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Test-time Augmentation for Factual Probing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:10.183656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.579337Z digest=sha256:eab830c491b7fb4e7e35f5f8d2cfe75639dfffbaf0bf8cf526226a69c6942b79

Observation 4bd67d26-faff-47b6-bb9f-2fcad9165c9d · outbound

This paper cites Improved Text Classification via Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Improved Text Classification via Test-Time Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.582126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.582126Z digest=sha256:add993bc697149d64fb62ed75d1e13d3960305446ff7bccd37247cdd30c8614e

Observation c6f71437-e619-4f1d-afcf-889046fc80fb · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.584800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.584800Z digest=sha256:bdae851cbfac2b5043b35545b01e960dba5ec1dfac0fc95fc514b4b3ec07b9b3

Observation 1bbf7138-9f82-4521-ba43-69ffb2f84fcf · outbound

This paper cites Humanity's Last Exam.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.590449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.590449Z digest=sha256:fbd39f99750153400741e5f3e2c8bba5c0f1943e54ca6a97483f914e0d8e3e6f

Observation 71adee13-e605-4423-b1c8-3a5ba9bb01dc · outbound

This paper cites Boosted Prompt Ensembles for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Boosted Prompt Ensembles for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.593127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.593127Z digest=sha256:decddc9f4b4a8286d6f4e38677f1c7d7a5efb91168816ffdfba6066cdc9e4eb8

Observation 925c9d0a-1c63-4e28-9703-13a0e88c3936 · outbound

This paper cites When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.595863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.595863Z digest=sha256:9ffb14afa0c938c6be513de4efafce7ecf3729880425d0a6ca1fe322fba75432

Observation 7e0924ec-8f7b-4c4e-99a4-539415e9c089 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.598569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.598569Z digest=sha256:d68d3ca9a150aeac6c49eccca4c985408db2bfec3296dd91076ea689895744cd

Observation 9ccfa0ef-9c32-4dcb-bb91-eb9c8f7731fa · outbound

This paper cites Better Aggregation in Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Better Aggregation in Test-Time Augmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.601164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.601164Z digest=sha256:6216a04074fbbce6855bf788881538d2e509f8529ae1d03d9d11f92845d1b67a

Observation f4a00aed-f634-41ce-9bf7-0cc24d88e6c7 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.607177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.607177Z digest=sha256:2a469e0a787c0091a28f593af01371181e4195bb06f35d0b8259518d6ae1a941

Observation f8e096ca-cdad-4c11-a066-6d3d0b38de34 · outbound

This paper cites Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.609878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.609878Z digest=sha256:67dee9a4da0d55efd1f18189eefdbec1f11d01f6d3ee3f0522892055b313fd8f

Observation aa0e2dae-a3e0-4897-ad3b-cf91d50542d9 · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.612699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.612699Z digest=sha256:98fa1fbed7263f34b3942499687e07de57c990b5a988555a61a645e0457f9164

Observation 5391658d-cc28-48e7-8402-61fa8f74db9d · outbound

This paper cites Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:55.066549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.615330Z digest=sha256:788bad4dea984c863c8c915cfc92d11edbe8863e2f01897a6660654da94bfa20

Observation 999d48c4-4613-4f7c-8c55-3b4b207a152b · outbound

This paper cites Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:01:44.750369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.617854Z digest=sha256:230888a4de39984199ac3fdf5b50117f44d9e0328085e9236ca576061e4fe1d5

Observation 74699a4a-1fd0-4095-98af-9a2eed0326d4 · outbound

This paper cites Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.620769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.620769Z digest=sha256:2457f2887664ac81c418bb916afb956a6137336292317f6f147b5ec69fb61f54

Observation 99260707-88f6-4fe3-86fc-3d63e485bc8a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.623654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.623654Z digest=sha256:77040c10029a8f3bbc028d16f56a391360f571145e396113ab8680e87045fb04

Observation 21afcc6d-3ff9-4e5e-a30c-b57482d666d9 · outbound

This paper cites PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.626588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.626588Z digest=sha256:7a6b42d959f13b156b69bdbb882e74d769462ff6f4efd0670c1ea55dd72f486c

Observation 40454ea8-570d-4c8b-97ec-ca34c0314d78 · outbound

This paper cites Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.629219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.629219Z digest=sha256:7e57bfc0e1c65812b1ed7e13ab6e8177a7ec2ec2100420a7fa8d271c7d2ea8db

Observation 24f0135a-0377-40f5-bc64-063830293537 · outbound

This paper cites output as array.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute output as array

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:44.714610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.631887Z digest=sha256:a74f30da7275347405480641d6ac9d75088a722a527fcd38ba83e1a69277166b

Observation 0139a27c-89ff-4c33-9b54-781123c46005 · outbound

This paper cites Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.587591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.587591Z digest=sha256:d170eb492c77270222b3e7a13c8a54e090d4ec199933b72beaae8191c3b58956

Observation 9acfdb9f-c477-417a-b29c-08e840dc63b5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.604113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.604113Z digest=sha256:a0db779d92fe9bacd3a8da5fea4a6b8d6e238908cb27c86412f8732ae66d7b08

Observation 7c8d72bf-8bd1-451c-bc23-a2fae51aa21b · outbound

This paper cites Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:25.226343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.570968Z digest=sha256:b2b35a17db3684c4e0c3e12566f4dbf4166e5dfbecaf24942d5eb14036805bda

Observation 966d0eb5-9ada-45fe-b15f-ed69797d3ebc · outbound

This paper cites Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.559468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.559468Z digest=sha256:fd88ab216c062baf8b47b4d2c41bff19a52b0bd6bc215941ab15ecb5f0caff9a

Observation af04e517-c00e-417e-9224-a77163cee234 · outbound

This paper cites RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.554277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.554277Z digest=sha256:c824ed57e621cebe34b9d35f1b21c2c5438683215e22465b88a8af9e3bdd2db5

Observation c64b5ec6-e754-426b-a1af-6a6b80e2f5e4 · outbound

This paper cites Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,

Reference 2024

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:40.409119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.542183Z digest=sha256:4184cc8809cbeb18e18ced9f8f931491bdc0f44fdd6311c59abc7589564bbafa

Observation 84b3b5fa-2ca9-4e13-b78f-4a9887b3e700 · outbound

This paper cites Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.536143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.536143Z digest=sha256:265c394f450ebaf4cfa595e9a65228df15a125a93d6a86091a74ded1ae8ff772

Pith citing papers

No inbound Pith citation observations are available.