Pith. sign in

Paper Citation Record · LEDGER

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

As of 10 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.28576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28576 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T03:19:01.477082Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce7f7f41-d48b-4e80-8c9d-4f2eacf33c24 · outbound

This paper cites Graph of Thoughts: Solving Elaborate Problems with Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Graph of Thoughts: Solving Elaborate Problems with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.572680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.572680Z digest=sha256:ae017d6cb16f905d2adb0abd33a02bcf8ed640a756234e8b71441aec8aefdf55

Observation e21a5e66-a140-435b-8796-cf71034a8cf9 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.616128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.616128Z digest=sha256:f92c69aae9a7314b9ff9f81687f5789c4e452b49a58c86c4f363f67ce5259fec

Observation f6c28381-4b71-42ca-9d6d-8a16ce472e79 · outbound

This paper cites Debate or vote: Which yields better decisions in multi-agent large lan- guage models? InAdvances in Neural Information Processing Systems (NeurIPS), Spotlight,.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Debate or vote: Which yields better decisions in multi-agent large lan- guage models? InAdvances in Neural Information Processing Systems (NeurIPS), Spotlight,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.652283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.652283Z digest=sha256:da9d5dc18fabdfb5d6bf1e508dc49250e08c1cb88fb0e125efb31d2180429ad6

Observation 99d3e645-2e73-4760-a8ba-f84dba51c93b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.690707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.690707Z digest=sha256:248efd0276f6c268381662dcb8ec82ee72cf3d70c352daca43b9451fdcfb214e

Observation 1cacb16a-b6cb-4228-93e2-c3c31f6b526e · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.728716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.728716Z digest=sha256:615d80eb3f857ec981b4f039a32edd5c49659a949ae7f9c94425227da02b5f54

Observation 86cb38bb-1d49-4a79-af60-a372adbf204d · outbound

This paper cites Bootstrap methods: Another look at the jackknife.The Annals of Statistics, 7(1):1–26, 1979.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Bootstrap methods: Another look at the jackknife.The Annals of Statistics, 7(1):1–26, 1979

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.769984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.769984Z digest=sha256:0def1ea1cacdc49e371c8fd3317138cc2e753fdf623b0508b50f7c470645faed

Observation 1f741c85-fea5-4a71-a7e4-f59f0155ce51 · outbound

This paper cites llama.cpp, 2026.https://github.com/ ggml-org/llama.cpp.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B llama.cpp, 2026.https://github.com/ ggml-org/llama.cpp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.809324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.809324Z digest=sha256:30745aab2b83af676b84c86c588c02166f9c5e111302649058e2139e3bb90cde

Observation c599eef7-51cb-4aaa-8c3e-763040c9b49d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Measuring Mathematical Problem Solving With the MATH Dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.819809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.819809Z digest=sha256:5df8efe9e086ba34d423048f602be95eb760b69e882959b1c8a7bb77b2f4edb4

Observation 34e9b2d7-162d-4ef6-b3ca-5ae0840f0b3d · outbound

This paper cites A simple sequentially rejective multiple test procedure.Scandinavian Journal of Statistics, 6(2):65–70, 1979.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B A simple sequentially rejective multiple test procedure.Scandinavian Journal of Statistics, 6(2):65–70, 1979

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.861811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.861811Z digest=sha256:3b526e41e89382b9093d39a2f4817ffda3f929004b6b357365383d5487b269ab

Observation cdb006de-9c05-46a9-85ee-f7926df310b8 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.902163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.902163Z digest=sha256:21da236f7a096c68629346709cd11adeb52d4a877378b02e68f478690a16ff14

Observation 4f3ed3f3-df60-41e2-9d5d-44f54f1414e6 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Models Cannot Self-Correct Reasoning Yet

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.943645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.943645Z digest=sha256:d7684ce5110426c7e2b3ea45ccfa3c13e21311ee51eb2140f78dbde7aa9dd410

Observation 7a3fa1e7-b0e4-43ca-862e-76b7460d077c · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T03:18:59.984759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:18:59.984759Z digest=sha256:9506e11613c53a05a096f3034c5708513824258713cf132026613784749c41ac

Observation d2fa40f4-9416-4a92-9d85-da7ad8d1a133 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581, 2025.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Scalable best-of-n selection for large language models via self-certainty.arXiv preprint arXiv:2502.18581, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.058675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.058675Z digest=sha256:3d4b4e7bc9fdb04b6d38d14f6b64b104b39524b98df007dc377c909c8feab043

Observation 3158fbd8-1a35-4f98-8b9b-1f9b132e1cef · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Large Language Models are Zero-Shot Reasoners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.116288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.116288Z digest=sha256:26984afd4e119c03fd19b591752a9355ec2813f20f472d54ee2590148131fc30

Observation 6d307dc6-a7a9-4bd8-a21d-fdc458ca8737 · outbound

This paper cites Let's Verify Step by Step.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.187681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.187681Z digest=sha256:6c4dcf8b35924c59c8edc852c3210542c903a235e5df75be98cd3d5ea873734b

Observation 73fc699e-70fd-4600-b3f9-f9ee9b52203e · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Self-Refine: Iterative Refinement with Self-Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.237462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.237462Z digest=sha256:2070d09b4d3ac8a4e877bd1c4dbf58d0957aefd5f2fb3d33c5958b8b7d686cd3

Observation b41f381c-8c5c-41d9-8aef-74bdff93efee · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.303898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.303898Z digest=sha256:e91b36d053276bb98ed9c28a95ea40823b7b4a33a9114b5fdd670a8eeda04232

Observation 3f3b467a-fdfb-4b76-a18f-a7ceeddc2b93 · outbound

This paper cites Fair on the Surface: Transaction-Ordering Bias and MEV in Mysticeti DAG-based BFT Protocol.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Fair on the Surface: Transaction-Ordering Bias and MEV in Mysticeti DAG-based BFT Protocol

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.364677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.364677Z digest=sha256:ebc2ceba582fe0fd14b7a1fe6e0c6803623ae90bed85aef2a4cb6fb179e67bf9

Observation e40ff60f-e0e0-4c29-9860-597da010867b · outbound

This paper cites s1: Simple test-time scaling.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B s1: Simple test-time scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.429895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.429895Z digest=sha256:39a8fec8c4c9939e1ef0892ba9efaa98c8b1fb4c3ec9ecae7a9f3d6c0c245979

Observation c09d33ea-403f-4738-85b1-ac4435c38fb5 · outbound

This paper cites Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.468509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.468509Z digest=sha256:076b8c2f0d06c5c6a05a2f736a02bd9dc2131035125013c2d3933dc84d6ee792

Observation 849ec5dc-5993-463b-aab9-21ca3305c913 · outbound

This paper cites Qwen2.5 Technical Report.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Qwen2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.517313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.517313Z digest=sha256:576673a0ddb9b43d93cb1075b2ffc12437b2877aaa9d2c6f516798901e87a48d

Observation ff981305-90ed-49c5-8259-b0a78b552d41 · outbound

This paper cites The sequential edge: Inverse-entropy voting beats parallel self-consistency at matched compute.arXiv preprint arXiv:2511.02309, 2025.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B The sequential edge: Inverse-entropy voting beats parallel self-consistency at matched compute.arXiv preprint arXiv:2511.02309, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.567426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.567426Z digest=sha256:0d0a1bdef75b10b5873198e701b5ef81f87096609564188e1be9b7634103cca7

Observation 2a636804-bd97-416c-afc8-45d29c695e58 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.608216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.608216Z digest=sha256:7bf9c32fed4994aa9c868952f98dbe15b40d8042067cc0f276019c1748c0e868

Observation b3d58206-ae11-42d9-85a2-f42d2964c6e5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.644343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.644343Z digest=sha256:50bcb5f1ce3ebc0e983f471714c1dc52cb9803b4bd1b19555f5cd2ffce2b33f8

Observation 3e18785f-cc04-4bf4-b8eb-5619454463e1 · outbound

This paper cites Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Inference scaling flaws: The limits of llm resampling with imperfect verifiers.arXiv preprint arXiv:2411.17501, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.700425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.700425Z digest=sha256:7c1ad4c27204632437d583e87aaa7d2a0cfcd3c194caf22039bcd5e4fe132e63

Observation 68fa443e-7841-459f-814a-55147701f9bf · outbound

This paper cites Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.768803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.768803Z digest=sha256:0c1e795672d3ffa39a2835d206e6e9aced75a53b9d329ec86c6cad20d5135fe6

Observation 9e3773e0-13b6-4392-96de-f30f6bf94230 · outbound

This paper cites Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.810695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.810695Z digest=sha256:50226d193f3a9cf42c003dc7b30a504d4dc21852e2f1b4d338053bf468fd3edf

Observation 1ab401c6-35bf-435e-9d8b-717772884b58 · outbound

This paper cites Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.887948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.887948Z digest=sha256:85b47322630b24d749b661953fe90f3bf028753bd47cbe891d65817b825071d9

Observation bb48a105-620c-44d4-8ad4-a06fb29c3e07 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.035096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.035096Z digest=sha256:b1b09cfbb9e54de12409a942fe394a0e4d1b9258559b077fca7db9095e13f640

Observation 8701c473-0ea7-40c1-b40a-557549e12a40 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.111637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.111637Z digest=sha256:df7cf7e445b98e3a7b837da006d44bc46e0e6a16fe4d05d2ad9f188348a25d6a

Observation f502b14f-dbfa-4269-9b97-4981a9bb318b · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.177346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.177346Z digest=sha256:a9eaeb66c2d1daee8e2cca6fd7ba4a44f8e6465b56ef348c76911c363ad2e1dc

Observation b3d636e9-6db7-40a2-88d0-9c5568733379 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.220557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.220557Z digest=sha256:816445d7b0f824cb1e8f6246fba7c8336e8b21885b9696dc4728d235cf6ed8e2

Observation 19fe31c5-9bf0-46b5-b97d-5da5dc5a9d70 · outbound

This paper cites Incentivizing llms to self-verify their answers.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Incentivizing llms to self-verify their answers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.285736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.285736Z digest=sha256:1c360507d51fdabe1b796f6aa7f5850706f4122ff7a1127ec5155b42b25cc91c

Observation 92ca9c9d-b4e6-48ea-8da8-f6c54236be20 · outbound

This paper cites Progressive-Hint Prompting Improves Reasoning in Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Progressive-Hint Prompting Improves Reasoning in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.329887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.329887Z digest=sha256:3966ffd48bbae8e102684f4c7ef5e06e8cc348f7ea655229f65d77fcb8eca440

Observation 558abeb0-e40f-4fde-bca7-0c93e7a6f5f2 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.409610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.409610Z digest=sha256:3063227f238462cf9c39e65a91c075a385576e841592790675514199476bc5e0

Observation 9f9128a1-3ea9-48a0-85e1-d1c939428a48 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:01.477082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:01.477082Z digest=sha256:7d3c7298d07b695f5121a8715a8e25dbdb2c6851e96413ebab6a1d5827034b19

Observation b1bea29f-2a0c-4cc4-880a-5c2a2e311368 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-07-31T03:19:00.973897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T03:19:00.973897Z digest=sha256:a87d2a5c6b33a3af4abf610f644cf6a92bade1826d806e8d35246553903c0e24

Pith citing papers

No inbound Pith citation observations are available.