Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:01.477082Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.28576.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:01.477082Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ce7f7f41-d48b-4e80-8c9d-4f2eacf33c24 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Graph of Thoughts: Solving Elaborate Problems with Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e21a5e66-a140-435b-8796-cf71034a8cf9 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c28381-4b71-42ca-9d6d-8a16ce472e79 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Debate or vote: Which yields better decisions in multi-agent large lan- guage models? InAdvances in Neural Information Processing Systems (NeurIPS), Spotlight,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99d3e645-2e73-4760-a8ba-f84dba51c93b · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cacb16a-b6cb-4228-93e2-c3c31f6b526e · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Improving Factuality and Reasoning in Language Models through Multiagent Debate
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86cb38bb-1d49-4a79-af60-a372adbf204d · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Bootstrap methods: Another look at the jackknife.The Annals of Statistics, 7(1):1–26, 1979
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f741c85-fea5-4a71-a7e4-f59f0155ce51 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B llama.cpp, 2026.https://github.com/ ggml-org/llama.cpp
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c599eef7-51cb-4aaa-8c3e-763040c9b49d · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Measuring Mathematical Problem Solving With the MATH Dataset
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e9b2d7-162d-4ef6-b3ca-5ae0840f0b3d · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B A simple sequentially rejective multiple test procedure.Scandinavian Journal of Statistics, 6(2):65–70, 1979
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb006de-9c05-46a9-85ee-f7926df310b8 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B V-STaR: Training Verifiers for Self-Taught Reasoners
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f3ed3f3-df60-41e2-9d5d-44f54f1414e6 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Models Cannot Self-Correct Reasoning Yet
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3fa1e7-b0e4-43ca-862e-76b7460d077c · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2fa40f4-9416-4a92-9d85-da7ad8d1a133 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3158fbd8-1a35-4f98-8b9b-1f9b132e1cef · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Models are Zero-Shot Reasoners
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d307dc6-a7a9-4bd8-a21d-fdc458ca8737 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Let's Verify Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73fc699e-70fd-4600-b3f9-f9ee9b52203e · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Self-Refine: Iterative Refinement with Self-Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41f381c-8c5c-41d9-8aef-74bdff93efee · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f3b467a-fdfb-4b76-a18f-a7ceeddc2b93 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Fair on the Surface: Transaction-Ordering Bias and MEV in Mysticeti DAG-based BFT Protocol
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e40ff60f-e0e0-4c29-9860-597da010867b · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B s1: Simple test-time scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09d33ea-403f-4738-85b1-ac4435c38fb5 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849ec5dc-5993-463b-aab9-21ca3305c913 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Qwen2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff981305-90ed-49c5-8259-b0a78b552d41 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B The sequential edge: Inverse-entropy voting beats parallel self-consistency at matched compute.arXiv preprint arXiv:2511.02309, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a636804-bd97-416c-afc8-45d29c695e58 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d58206-ae11-42d9-85a2-f42d2964c6e5 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e18785f-cc04-4bf4-b8eb-5619454463e1 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68fa443e-7841-459f-814a-55147701f9bf · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3773e0-13b6-4392-96de-f30f6bf94230 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ab401c6-35bf-435e-9d8b-717772884b58 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb48a105-620c-44d4-8ad4-a06fb29c3e07 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8701c473-0ea7-40c1-b40a-557549e12a40 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f502b14f-dbfa-4269-9b97-4981a9bb318b · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d636e9-6db7-40a2-88d0-9c5568733379 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19fe31c5-9bf0-46b5-b97d-5da5dc5a9d70 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Incentivizing llms to self-verify their answers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92ca9c9d-b4e6-48ea-8da8-f6c54236be20 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Progressive-Hint Prompting Improves Reasoning in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558abeb0-e40f-4fde-bca7-0c93e7a6f5f2 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f9128a1-3ea9-48a0-85e1-d1c939428a48 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bea29f-2a0c-4cc4-880a-5c2a2e311368 · outbound
Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.