Pith. sign in

Paper Citation Record · LEDGER

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

As of 23 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2505.23540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23540 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.646894Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:10:40.020460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:55:53.842268Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d0201b1-c789-4d40-a1b8-4531e2a509e7 · outbound

This paper cites URL: " 'urlintro :=.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.753612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.753612Z digest=sha256:fb16728e465731ee4414cc8e99c7562bdf879da52cc385f1fd77afc46d979679

Observation b9fa71cd-82ec-4690-a49a-03bb98b80ddd · outbound

This paper cites write newline.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.884044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.884044Z digest=sha256:995945ec283a72dd3aefbfc9e5f746b00ced3ed309f1fb62461405ca251c7801

Observation 2cc13414-55b7-4a9c-abf4-e85bea0b9a3f · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.051024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.051024Z digest=sha256:09f2a2c41428891091a257afc5aa874604a791c5615160c3b968b8de09222153

Observation f36448ab-9dbb-470e-bfb2-1ce6b9b67b6e · outbound

This paper cites PaLM 2 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.263325Z digest=sha256:aeb77ee538c01d64e8949cdcb35425cf9bc519808304d60328614d49910fbfdf

Observation ae85c740-cb08-4a98-8fe2-77b25cb86a95 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.435160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.435160Z digest=sha256:41e63e0f8f6fd70136f5e39e52379eeae8f0aab703bbb48de72d4cd765e75f17

Observation 0a5a5b99-02cb-4289-8626-cc777c298547 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.525774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.525774Z digest=sha256:5ef1ef098599800108ff3adec938e6665871a3af8f08a26bc85fe3f9a423f026

Observation c877b047-7045-4853-8417-22c2d296aa21 · outbound

This paper cites Qwen Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.699588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.699588Z digest=sha256:d9bab7d61b9c6d127da9855c4c4b63a556975d81060825098e06b73a179a42dc

Observation f0670dc0-94b8-41d2-b2bb-c56c94912132 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.872992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.872992Z digest=sha256:e37bd39c86f8899202f428a410ef6e685337d50e11806cb506bad1f540deadad

Observation 4e73070a-331a-475f-9fac-195c6406a8e9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.005985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.005985Z digest=sha256:368d04e17acbe7dae2daecc82ec981a4be56251ed98c2937043dc861d7bfa63f

Observation 789ab79c-abe2-4128-a481-15da00115460 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.128007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.128007Z digest=sha256:19911051396befaa35d8cfc89f866ea785b3839a9fa09daee879b1ce69bf3c88

Observation 04de3fd1-ae6a-4501-9a81-4e0544bd0d13 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.284922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.284922Z digest=sha256:1b6a58c71eb27963644b063d60ba6d03a0c78f393c19050a76dd9f33f01d4b95

Observation fe43d705-b098-4d91-80cc-20a738956ba7 · outbound

This paper cites The Llama 3 Herd of Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.497500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.497500Z digest=sha256:4867aab9647469bebd399878481d90b31caa2791199150b6125ec866dc0b5e9e

Observation 5d37c2cb-3193-40da-84b9-393519fa4a1a · outbound

This paper cites u rnkranz and Eyke H \.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning u rnkranz and Eyke H \

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.641896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:48.600730Z digest=sha256:784c89736ae7f201e83f29f001f319c497727ec37e78e077185c92e9de93232e

Observation 9ed398c0-88ae-4c62-87ff-d30bc9adbbb4 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.773375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.773375Z digest=sha256:d4c597f20df02f0c3b70011048ed297cf11add14940f19045cfd58986cfcbb8a

Observation 807d9afd-850a-474e-a795-f9013df45462 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.941524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.941524Z digest=sha256:348c60d10f6d76af289684ca9eea4e65cbb75906ad6e98ce554cd3546265c344

Observation 1d2b6096-c2c2-44ea-bb5f-f08a890ccb5d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.412047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.082974Z digest=sha256:f34b79ba832833a86a44d36b2d5221e035fddfb568f5f0741e5e37f53a7f0c19

Observation c0be4535-8a3f-4e1d-ba62-e0be2797985c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.267026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.267026Z digest=sha256:4b367e076356f7fd95e991023b7549a935a240d0e8a152e2d3ec21557c644c13

Observation 971a7edd-e295-4cbe-977c-2b9e0b5e304a · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.170201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.366558Z digest=sha256:abda616af7990a1472b0af44c67846eb8b2b1352582bdc2d3761badac5e21c0a

Observation 9a896b54-0054-41d1-b37b-f21852b92d9d · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Curious Case of Neural Text Degeneration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.542632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.542632Z digest=sha256:f0f6cfbc273102ec52f83e82fffdc9d5bb023d503ba6bd5f66562c5dcd0bef90

Observation aa60184e-e272-4d86-8765-02f23d4f1a18 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.764910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.764910Z digest=sha256:b4b031709ddbc8c52c66c9e9d08eaefec4ccffa3193151cd4aa7023c4c0802ad

Observation 8589f443-2d0a-49af-a999-e6da8d773234 · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.921998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.921998Z digest=sha256:2b701b79020abadae4af59901dfba8dbab8c0c274f7e86bbed4329a7e18e2401

Observation e39e65c4-b8a8-4410-93bd-1a9090af818b · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.073075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.073075Z digest=sha256:4a8f15cd2dec2c9f6776b1c6fb03c216fa86b8cd82da6c125d0bdd423e462b63

Observation a93cddcb-a3e8-4f57-ab22-35ccfae51949 · outbound

This paper cites Mixtral of Experts.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mixtral of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.236031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.236031Z digest=sha256:736528743e2d98b1c1b947d2f935610fd4fd48f89c70330d9f2257337088e4a5

Observation 5d64946e-9d68-4229-a38a-d70737774dfe · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.415748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.415748Z digest=sha256:3943c834502e5cc887894f862421466e14f761114dc2add0bc8e2c18ecd6770d

Observation c18b59a3-11d1-4f60-a310-2f278c5be249 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.896260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.526871Z digest=sha256:5b4c2c937771527f82ec8f737dbc0b5e9e7d5495c181e1d6e2b289bdf84382df

Observation 25372d40-7a3c-47c2-9c23-05b9270628b4 · outbound

This paper cites Let's Verify Step by Step.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.661025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.661025Z digest=sha256:6543afa52f598ddd8bee708f91e95a341226a007ef10603fcb3e2ae4ad1e0eba

Observation 4088bc24-c3d9-4787-8e90-48c51ab95d93 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.753591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.753591Z digest=sha256:38f68715b676906c6cb9213f19dd3b71567c8c977f6911e5c12094eb2f23defe

Observation 4dc331b4-4341-4494-8e12-08eb8254157a · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.846271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.846271Z digest=sha256:7095f50dfb150309b0e59d8abaa7a6f311c3aacbf09ce9036b2d43ab17904bb6

Observation 68230bdb-df1b-4f35-9f3e-b72edb08be27 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.654117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.956262Z digest=sha256:5eb78a7be13e9c908ce15bb2da3d2a3e12e8c191bd7bc7c717bed3b50cc012ac

Observation 3dc8d264-c4ac-4339-a805-f8f2b22d8a18 · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Inverse Scaling: When Bigger Isn't Better

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.113697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.113697Z digest=sha256:5b55e810c44e757351c063b293f920f8059277da72ef71b254d28cf8558e5eec

Observation 292d24a2-6db8-4036-bb00-29d468322805 · outbound

This paper cites Large Language Models: A Survey.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.344663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.344663Z digest=sha256:1bb840d80724042f32e96c5a798dedd5b45e4fda4fcaf7d273d995d5dcb026d2

Observation 84a2e209-521b-4d8d-8071-a0a6fea3e594 · outbound

This paper cites GPT-4 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.544455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.544455Z digest=sha256:091c1aed26e2c50dda2eebcb5717bd0bff0f723d87cba98aeef70be9a0d77792

Observation f24bd252-aa67-4027-9ffe-31d4bcb8ed9d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.715839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.715839Z digest=sha256:245fb61bc56b532e223f78a1d2e5ac5a95bf6956fc507e9564673416be73e7f0

Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.881145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.881145Z digest=sha256:1e04f8a324ed30cf0f92f6bbeee3cdd9e932b26a4cd7765bad8a205b4df6cd1b

Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · outbound

This paper cites Self-Consistency Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.028303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.028303Z digest=sha256:383a357231a1d8063b9f090b5625201c96b3a9966ec22458810a82b4055b3af7

Observation 63350242-06bc-4505-9b17-e3593cbb3996 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.236313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.236313Z digest=sha256:934893bd0e98c6f7dee1617da5cf214be0f26b573d7294db2070d822423432bf

Observation 5d901273-a9c0-4812-90d3-7e8131a2b228 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.442199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.442199Z digest=sha256:cab8c1e12ac9ade703456999bda5277ca493330b7d9f3763632b8f8ae4b7596b

Observation 2a5a65b9-8bbd-4185-9e6d-cf89e8a9d320 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.594267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.594267Z digest=sha256:fa383d8b4ed32de0a352068e4c54b6d155890066e966792488f80e55d145222f

Observation 5628600a-17b9-4f11-9666-e7cf21c8ed3f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.739102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.739102Z digest=sha256:925513c007049b456f6d1fb015be2610301abb758bea1e05cbed849a6b1c462a

Observation a53b2819-2362-4d2d-917a-ff8642221b71 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.883958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.883958Z digest=sha256:3e6d2c84f74a1d97d2f9da8ba9a370daf6cde5ec557f4560419a21b5dfb4a486

Observation e459c472-4f91-4d75-b669-3e9993087f5c · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.070330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.070330Z digest=sha256:deb9a77d28fe96ea16be425de49fb60f8e0006f160281086b08216e33897d419

Observation a118128f-1d83-4816-a98b-b1ee6b2c5513 · outbound

This paper cites LRHP: Learning Representations for Human Preferences via Preference Pairs.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LRHP: Learning Representations for Human Preferences via Preference Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.221253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.221253Z digest=sha256:7aaf14fc7f24047dd117cebf454f4b37f745be386fd1765d6649291b8a115d4e

Observation ccea8820-6185-4dab-846f-fe5503aacfb9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.362214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.362214Z digest=sha256:22de7e2996efe146f9f7e34b2dab3ea74decdb2068d6c18da5420f72e765cac3

Observation 7009d628-dde1-4970-b138-c3801022ed5d · outbound

This paper cites Consistency of a Recurrent Language Model With Respect to Incomplete Decoding.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Consistency of a Recurrent Language Model With Respect to Incomplete Decoding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:55.005187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.508094Z digest=sha256:ddfb229a681f060a71d9a0df39c5979f7bbd026c9c91f4f4fe89d9f97d4d0b5d

Observation 6261dd51-e4f9-4d11-89b5-626a1f7d92d9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.456689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.657068Z digest=sha256:9e296efd5e191df9641fe004400c907bc4b38bff199d1edbb0a7102f584c9a9e

Observation fcbbe315-6b07-4e8d-be48-ea0ba4728290 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.850737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.850737Z digest=sha256:b906dc322397c1a5ad1b41ca871de8bf49e66d474731b379e16bc969b08674ea

Observation 0cc5b074-990d-47eb-8731-df040fe7866f · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.971380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.971380Z digest=sha256:f0be1ff194eb84186b74d5c011ced4953bc192a6326c8bec4b420fec230fb85c

Observation 12fac24c-5d97-4132-8ce7-02079af468ed · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.069667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.069667Z digest=sha256:e65c376a86da97de9459b031349a5ddbf8bc284300f793495dd52bb8ef0701ff

Observation 0c0ef65f-e24d-4869-a4bc-854791757086 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.203982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.203982Z digest=sha256:67d637b60f327d2fd770ff1683e139ea868877a2f824b0d79615556746a47590

Observation 37060bab-bcc1-490a-83ca-f6405add8166 · outbound

This paper cites RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.380705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.380705Z digest=sha256:7e85f34565ca89b916d1b8f691f9baf445d3859c808deab828cf2e6c9b523755

Observation ca19ee2c-5fa0-44d7-a150-b0197044a3a3 · outbound

This paper cites Self-Rewarding Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Rewarding Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.498345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.498345Z digest=sha256:04a4322ba17ece5d4265442ca508f15381d8195658cfe95093750af11c1757a0

Observation 8cc37db7-4a5c-40ce-99ae-84eaceba409b · outbound

This paper cites Token-level Direct Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Token-level Direct Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.646894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.646894Z digest=sha256:7458a4b57c3faad784d9b998a80592713dba2f2f9eb44b07cb01dfd06643dac2

Pith citing papers

Observation 704f80b4-7721-4f40-80da-ebb3b85067ed · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.845454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:f0ac7191d04af6653778043e1b79c6037e5612225e7e612be5e9bc4311e9eb63