Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.21941.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21941 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:43.743614Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:51:32.303810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:51:34.729631Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41a1d5bd-d8e3-4943-95d5-8355e9cfffa8 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.332471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.332471Z digest=sha256:b979999b7a89356e26d7371a66aaf550870c535b1e59a7c26caa73df41a685c5

Observation 626451ef-7109-41b0-9c61-96cb3dc9d052 · outbound

This paper cites Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.440527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.440527Z digest=sha256:e723191de8eadffd79f7ed5e85270d2d657339a93b66779101238b2afb67df36

Observation 8ef982de-4648-49a6-ab6c-3a7bbb57e0f6 · outbound

This paper cites The Llama 3 Herd of Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.484105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.484105Z digest=sha256:dd376262aab8a96db27204e01cd7c462ce20fac8906d4aefd30d927b0982f829

Observation 95536eae-6b54-4d11-bdc8-9f8ea417005d · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.658486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.658486Z digest=sha256:94e3e4b519a67bb15e93e96cd23a67bb8ad4fb643ce70ca25bcf5e63223cff34

Observation 1c266c9c-4760-4735-802f-c143fd4fd2f5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.818840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.818840Z digest=sha256:fbe2dc12ccac5797f5ce5758458592b062e95c5a129fd81a1ca5481ac50dfc10

Observation 34f63ecf-aaac-429c-8775-beb645de9a9d · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.911718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.911718Z digest=sha256:747d826f32f56e9ae066eb394ac6ce6735f46d700415d285351052c76d401bb0

Observation 21bd6289-0180-41c4-8cd2-23152b44ea54 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.998884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.998884Z digest=sha256:8435b1618d58b108f0557678d61fe0ee6c3fac4bbdd22ff0bbe92cc655168aa2

Observation 726a627a-496b-4340-bccb-86ca399fcdfc · outbound

This paper cites s1: Simple test-time scaling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation s1: Simple test-time scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.121015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.121015Z digest=sha256:6b91dd20310c3c16179e27cb1bd3f297628726af11e85b66ef5dcdcc34ed1f11

Observation 863fd582-d7e9-48c3-ac1c-95901518b98c · outbound

This paper cites Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.281232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.281232Z digest=sha256:950cd5f3955af8477d179eaaca9fbec72e50b5a2b1b499c4bf727bc0c1391618

Observation 5a4ddb01-3c61-4e32-a477-e4e04a915e8d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.393276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.393276Z digest=sha256:fe11a1a0c3d6da4ee6354ffa8d73f461dca41954aad2d2392a38847dcc84be7d

Observation 822cb2f8-5903-45d1-8226-9d669659ca0f · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.499614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.499614Z digest=sha256:199689649161a9007db6506dc4a76549a76519a278bc3a632aa45cc633529d93

Observation 5115f042-f7de-44e9-8033-783d0859fe69 · outbound

This paper cites We abbreviate these names in figures to make them more readable.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation We abbreviate these names in figures to make them more readable

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:22:44.453195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:43.743614Z digest=sha256:1c321612efc7fde06831b214249a06ce741c15f3002514d617ce11a04d467a37

Observation ebcbb6ad-83cf-4c9f-8494-2c995f9ac964 · outbound

This paper cites A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:44.624366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:43.620849Z digest=sha256:fdfb79bee6831a78782db88d2204c2c0cc852a5836298468cd639680ad1679b8

Observation 1b23aeb3-3861-484d-8f0f-23ae9a909a79 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.270421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.270421Z digest=sha256:d3d2b9e9ee7c7677a3b1e8c3c0b42f12b2ef8283aff94514e4a5efe7ba652b67

Observation d4a2c687-b1d7-4733-8c38-14a52b70dc62 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.566750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.566750Z digest=sha256:1c9b7ad4ddc97ec3b578281bd9dd0f81bcc198addaf471a65c2ae6dbc8065c3a

Pith citing papers

Observation efef82c8-fc47-4fb1-b1ee-f585b8daaa35 · inbound

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs cites this paper.

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:51:34.734802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:51:32.303810Z digest=sha256:15b5bbd6dd936a53c6583556f8c6fa7fc5b97f5b60fdbee11a931291fde9ada6