Pith. sign in

Paper Citation Record · LEDGER

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.09351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09351 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:29.631887Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact7
  • verified fuzzy1
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7faa10d-4099-41b5-a650-21c53b820358 · outbound

This paper cites Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.532259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.532259Z digest=sha256:f593e7bc4d0b6d9f324bfdc4f96ccab448c27f7dd5997f278ae07e1169835498

Observation 267c094b-32aa-4078-97f7-149f5ed65537 · outbound

This paper cites The Surprising Effectiveness of Test-Time Training for Few-Shot Learning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.539084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.539084Z digest=sha256:0b8543c04ef228ebbb89cf11b153de5141e2901c7e9c9746b5766a5497b2ac5d

Observation 274cfb02-528f-42a0-9697-9dcd0ab32c86 · outbound

This paper cites Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.545090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.545090Z digest=sha256:d29d270e7aa46ac0999bd86d366658384911cfde2133ac9da5af5585dbad8e90

Observation 9e6fa001-a216-4e73-9165-179998869704 · outbound

This paper cites Exploring LLM Reasoning Through Controlled Prompt Variations.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Exploring LLM Reasoning Through Controlled Prompt Variations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.548200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.548200Z digest=sha256:adbcd682f0c4ffd1b5becf4983a8087d2cba0ff944615e1caff8360975bb0e15

Observation e28a2820-d83a-4870-a3e4-ad5d68cb7093 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Teaching Large Language Models to Self-Debug

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.551448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.551448Z digest=sha256:acede1c3bb54fb55bd85cb9cb3c78b08752d6d14c2dae2c4445cce268342615c

Observation dba996a4-ac77-43ba-b281-90b9f444d1e3 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.556706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.556706Z digest=sha256:4cd47a587871a2f67869c1d2dbc2e8b1901fb45ff4b71965ad539d824feeb9df

Observation 138c76ef-b0c1-414d-9ba9-ae9aaba98965 · outbound

This paper cites Frustratingly Easy Test-Time Adaptation of Vision-Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Frustratingly Easy Test-Time Adaptation of Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.562315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.562315Z digest=sha256:f608360015898b028417dfa481b9a954b648edc7a3fb2479a96dea5aad4068ed

Observation 1d6734f8-74a2-4378-b1f6-cc53c7332200 · outbound

This paper cites M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:25.243674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.565203Z digest=sha256:32eb2e6ba70733c4a1ce60026c2ab81c46bfab912ba372df753359ae4f66b127

Observation a6e6da89-ec2d-4ea2-8465-ff9cf098849e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.568253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.568253Z digest=sha256:c8acba3882d5a7560a132687eb4796d5b2b377923cb49a3d8210e8afa98b39ce

Observation 3b7c528a-9b4b-4149-94d9-e81195fc1496 · outbound

This paper cites Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.573473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.573473Z digest=sha256:c0216363e6228e3eb02ae4d225108ed68b64de411b32c3dc56327f1139633931

Observation b3c73951-a3d0-45ea-9c3a-2043d58f2330 · outbound

This paper cites Calibrating language models via augmented prompt ensembles.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Calibrating language models via augmented prompt ensembles

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:40.437350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.576659Z digest=sha256:a87e67904854aa9c6686c30aba5cd7666f81c8becbcf2f1839a9bf1c0497e530

Observation cff92cfd-bdf2-47bf-a7a1-fc5c5e406a62 · outbound

This paper cites Test-time Augmentation for Factual Probing.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Test-time Augmentation for Factual Probing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:10.183656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.579337Z digest=sha256:cf74d86da86fee8580b5e3a5f98ddac9714f6d7e4e42667cc3077523fdaa93d3

Observation 4bd67d26-faff-47b6-bb9f-2fcad9165c9d · outbound

This paper cites Improved Text Classification via Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Improved Text Classification via Test-Time Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.582126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.582126Z digest=sha256:fc03303d032c184ef5175a8452b56471b21889262f79b3365ca054cd4d18f47b

Observation c6f71437-e619-4f1d-afcf-889046fc80fb · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.584800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.584800Z digest=sha256:692836b235b6e836ac3dc99cd7b12e758f97d7b29e0b0783b83a5ee0222021ba

Observation 1bbf7138-9f82-4521-ba43-69ffb2f84fcf · outbound

This paper cites Humanity's Last Exam.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.590449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.590449Z digest=sha256:d62382d5fcebae41ee407036d8ad373f888cf4949f9c3cc3fa5d3703f2602fe3

Observation 71adee13-e605-4423-b1c8-3a5ba9bb01dc · outbound

This paper cites Boosted Prompt Ensembles for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Boosted Prompt Ensembles for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.593127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.593127Z digest=sha256:8acd340c996d2eba23dae540922e2ce593f0f0f1dc91588deb25b894478722b7

Observation 925c9d0a-1c63-4e28-9703-13a0e88c3936 · outbound

This paper cites When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.595863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.595863Z digest=sha256:e54593b16747e58cbd90170dc340c5bdbaf621fcd51fe41a1c5fa67a46067a8e

Observation 7e0924ec-8f7b-4c4e-99a4-539415e9c089 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.598569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.598569Z digest=sha256:508c5dfd53a3b163708e8fb0b1ebc08641418603c9706e69bd1aa2bd42e82cb7

Observation 9ccfa0ef-9c32-4dcb-bb91-eb9c8f7731fa · outbound

This paper cites Better Aggregation in Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Better Aggregation in Test-Time Augmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.601164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.601164Z digest=sha256:6f3ea261018a16b753ac0778aae2f3ab43927e02491626ae38b08aee5b665514

Observation f4a00aed-f634-41ce-9bf7-0cc24d88e6c7 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.607177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.607177Z digest=sha256:b98f04963f620f32db4e2c3e87af5f4bfbdce7790f5f2e4f3d8d013cd92e7cd5

Observation f8e096ca-cdad-4c11-a066-6d3d0b38de34 · outbound

This paper cites Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.609878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.609878Z digest=sha256:8595f34f3a79ce1f5428ef437bd9e45921a4c9298abff39e0adc53d063c54fda

Observation aa0e2dae-a3e0-4897-ad3b-cf91d50542d9 · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.612699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.612699Z digest=sha256:40bf48b8844499546bab089c1a6c3a6244acd1d0384bddc59001796bdde8b5f5

Observation 5391658d-cc28-48e7-8402-61fa8f74db9d · outbound

This paper cites Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:55.066549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.615330Z digest=sha256:ec2e07f2995d7c7b5eb29cf21a42039e4f575377ab062100fdbf1fe829180f32

Observation 999d48c4-4613-4f7c-8c55-3b4b207a152b · outbound

This paper cites Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:01:44.750369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.617854Z digest=sha256:1339ff3105c2c1ff1bf097ee3feb0cf7778613f9d755625a02feb7a35f03b713

Observation 74699a4a-1fd0-4095-98af-9a2eed0326d4 · outbound

This paper cites Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.620769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.620769Z digest=sha256:8f40a0b64782c9671ee14b8f8d157763c371fc969faf152e4cbf309b3c94e70d

Observation 99260707-88f6-4fe3-86fc-3d63e485bc8a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.623654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.623654Z digest=sha256:ffaf60a6da72e097e57b9a4d3b4a793e7ff1eb2068880b77429d043e695e02b2

Observation 21afcc6d-3ff9-4e5e-a30c-b57482d666d9 · outbound

This paper cites PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.626588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.626588Z digest=sha256:3bba375308b21acfbaf0b3be8f86dc47c58025a17c424db5415717f9bc38dd1c

Observation 40454ea8-570d-4c8b-97ec-ca34c0314d78 · outbound

This paper cites Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.629219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.629219Z digest=sha256:454b92e4b91588c303ba23694c9cbc993ce413634fee73302994d1ec1c289a76

Observation 24f0135a-0377-40f5-bc64-063830293537 · outbound

This paper cites output as array.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute output as array

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:44.714610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.631887Z digest=sha256:70d341b4d5c230db5c914bb5e8c555f18b71869b7b543b2d157bb69bc0bc3098

Observation 0139a27c-89ff-4c33-9b54-781123c46005 · outbound

This paper cites Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.587591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.587591Z digest=sha256:380b18df498015bc297c9c4fb2e67145c3160de895e91694cb24f231725e1595

Observation 9acfdb9f-c477-417a-b29c-08e840dc63b5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.604113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.604113Z digest=sha256:f5ac8230a5103e12412b1b8b91adaf29f042eb2ce8b81dbf393941f5c45c0dfe

Observation 7c8d72bf-8bd1-451c-bc23-a2fae51aa21b · outbound

This paper cites Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:25.226343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.570968Z digest=sha256:25dc008a9f334b2b629820e1452a9c560db96b9c5ba99e29226661c1139caf1f

Observation 966d0eb5-9ada-45fe-b15f-ed69797d3ebc · outbound

This paper cites Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.559468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.559468Z digest=sha256:29510a5f12ddec43569ab911216aa8faa07e6d65e5a88ff4c8a71f5984643182

Observation af04e517-c00e-417e-9224-a77163cee234 · outbound

This paper cites RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.554277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.554277Z digest=sha256:6a29a0266c0d0deb7e885b317b0294c223ae8be5abaf8585614ef90a73d15d44

Observation c64b5ec6-e754-426b-a1af-6a6b80e2f5e4 · outbound

This paper cites Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,

Reference 2024

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:40.409119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.542183Z digest=sha256:2c66f463635e20fc641f8a64ed91dd96152b95ca34e5247876a6c458c035f386

Observation 84b3b5fa-2ca9-4e13-b78f-4a9887b3e700 · outbound

This paper cites Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.536143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.536143Z digest=sha256:5cc56a2d8e543789a4ee74c3de91a2859d4a1c3706360cff366d11e93dee0f94

Pith citing papers

No inbound Pith citation observations are available.