Pith. sign in

Paper Citation Record · LEDGER

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 1 inbound Pith citation observation for arXiv:2505.23540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23540 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:54.646894Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:10:40.020460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T03:55:53.842268Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d0201b1-c789-4d40-a1b8-4531e2a509e7 · outbound

This paper cites URL: " 'urlintro :=.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.753612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.753612Z digest=sha256:7733ecc7ed06b534e3a5ce363699a04c25018b422a0bb3344f889435e5aad61a

Observation b9fa71cd-82ec-4690-a49a-03bb98b80ddd · outbound

This paper cites write newline.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:46.884044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:46.884044Z digest=sha256:662faec24301d5bdfbe9fa0d84583d02b41a3ea663b164e11dfe698ea43bab33

Observation 2cc13414-55b7-4a9c-abf4-e85bea0b9a3f · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.051024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.051024Z digest=sha256:59a701cea4c964086b0fcfba69cd5f01fd4d0d5cc4335b863358abdcbf3c7d11

Observation f36448ab-9dbb-470e-bfb2-1ce6b9b67b6e · outbound

This paper cites PaLM 2 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.263325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.263325Z digest=sha256:f82d06947d942f20292f9e0c6f0cc6db3d4efafe5081c422d8cb947fa99269a4

Observation ae85c740-cb08-4a98-8fe2-77b25cb86a95 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.435160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.435160Z digest=sha256:cbba7bccbb97fa8daa79b7c728bb85f5aace1900af219d92eede8e5a9e075ec1

Observation 0a5a5b99-02cb-4289-8626-cc777c298547 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.525774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.525774Z digest=sha256:a223345932d2acd750c3c027f23c031bc7dcbe9389ae3f1e7216ae340e93d914

Observation c877b047-7045-4853-8417-22c2d296aa21 · outbound

This paper cites Qwen Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.699588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.699588Z digest=sha256:3a9da4e1e07e94d2fffda77ff09213874bdf8d168c09e7c74fdef4cf9702a217

Observation f0670dc0-94b8-41d2-b2bb-c56c94912132 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Constitutional AI: Harmlessness from AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:47.872992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:47.872992Z digest=sha256:628fd5ac8dfa0adac20b2061aa838f0e9dd7584b7bfad9569ed27693387e27f5

Observation 4e73070a-331a-475f-9fac-195c6406a8e9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.005985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.005985Z digest=sha256:145d210b4f16aa3e86467aee637d4bbd9aadbfc689641ad207e0deef948fdf2e

Observation 789ab79c-abe2-4128-a481-15da00115460 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.128007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.128007Z digest=sha256:7d35dbe6186db7fee9d06b9f2f49da531920b37bb48bdef5dc35f66fdedda8d2

Observation 04de3fd1-ae6a-4501-9a81-4e0544bd0d13 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.284922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.284922Z digest=sha256:aedf1f4fd0f1b8ba8e96826a155b840f0c04560c1efbc18834793d418b0b01f9

Observation fe43d705-b098-4d91-80cc-20a738956ba7 · outbound

This paper cites The Llama 3 Herd of Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.497500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.497500Z digest=sha256:7e7f72927102680fbcf58791e61a1ff4168598f853bf6ff8361a7a8ded0db5a9

Observation 5d37c2cb-3193-40da-84b9-393519fa4a1a · outbound

This paper cites u rnkranz and Eyke H \.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning u rnkranz and Eyke H \

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:48:56.641896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:48.600730Z digest=sha256:36cf29d9ea8a9e6724c7f44c4f57751696d32d2e095291e42f1253e9fd75df0d

Observation 9ed398c0-88ae-4c62-87ff-d30bc9adbbb4 · outbound

This paper cites Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Towards a Theoretical Understanding of Synthetic Data in LLM Post-Training: A Reverse-Bottleneck Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.773375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.773375Z digest=sha256:6c44e01d0a27f16e4a58fb5a175e01c7e7fb1f8536e93e8357a942b378933295

Observation 807d9afd-850a-474e-a795-f9013df45462 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:48.941524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:48.941524Z digest=sha256:f66f9b0d5dbc0eb7e510aaef23e9dad9e35b91b1fb89d734fbca052226e17970

Observation 1d2b6096-c2c2-44ea-bb5f-f08a890ccb5d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.412047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.082974Z digest=sha256:bfa35e9204a10b4ed734df231777ef69062c8f6b5bdf333e008d8e64f5d703af

Observation c0be4535-8a3f-4e1d-ba62-e0be2797985c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.267026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.267026Z digest=sha256:65426e8ad8f69ce25cc8ce099e3dcc5e064f063ee8080cddd2183131cadedb18

Observation 971a7edd-e295-4cbe-977c-2b9e0b5e304a · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:56.170201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:49.366558Z digest=sha256:5b57c23c614093ee1642a4a8d3fb90ba2c693a7bf57e5773b52403a8568d813c

Observation 9a896b54-0054-41d1-b37b-f21852b92d9d · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning The Curious Case of Neural Text Degeneration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.542632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.542632Z digest=sha256:e30678287a843064511d8f5221e5550b7efeeb999dd3d774b67ec16ae283b017

Observation aa60184e-e272-4d86-8765-02f23d4f1a18 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.764910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.764910Z digest=sha256:d265274499368a2ce01dab3f9caff5d49cff3c110335b763d770614063e86e09

Observation 8589f443-2d0a-49af-a999-e6da8d773234 · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:49.921998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:49.921998Z digest=sha256:1423fb251a7abee1e6aed4c7ade4957366abeb153f812113e23234bff6cf9b88

Observation e39e65c4-b8a8-4410-93bd-1a9090af818b · outbound

This paper cites Mistral 7B.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.073075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.073075Z digest=sha256:e307967769612bea7cd2d31c654f1264ccbbba8c8dcb273811da04829dd04079

Observation a93cddcb-a3e8-4f57-ab22-35ccfae51949 · outbound

This paper cites Mixtral of Experts.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Mixtral of Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.236031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.236031Z digest=sha256:70448ca76f630b2c0df2bee825562790dafc70e39101c1e1cd14bb85494cc019

Observation 5d64946e-9d68-4229-a38a-d70737774dfe · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.415748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.415748Z digest=sha256:82d5b28a976e36a05e4eb711c19092d974c911e6de98541d2e6e601786a1fb2d

Observation c18b59a3-11d1-4f60-a310-2f278c5be249 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.896260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.526871Z digest=sha256:7b4725823b80f8e4174c4c5df6b58ec97df697c8cccbb4be9498746b82e55198

Observation 25372d40-7a3c-47c2-9c23-05b9270628b4 · outbound

This paper cites Let's Verify Step by Step.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Let's Verify Step by Step

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.661025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.661025Z digest=sha256:409e3f30a00e2eab4a12efe1093b0ca2c7fa55b767962cdc023beef9ec4fbae0

Observation 4088bc24-c3d9-4787-8e90-48c51ab95d93 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.753591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.753591Z digest=sha256:13e1c028ddd9fbdbb0968595ce8d156d9feb5a108544a723adeb0c734776bfc2

Observation 4dc331b4-4341-4494-8e12-08eb8254157a · outbound

This paper cites Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:50.846271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:50.846271Z digest=sha256:ba4417afacf1183480dc3569962b81e74eb068bea702f2414db2e7c540a99bd9

Observation 68230bdb-df1b-4f35-9f3e-b72edb08be27 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.654117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:50.956262Z digest=sha256:4da6606ae8ff8875b8b734100d91d516be6d644c32180b20281e349c4419c221

Observation 3dc8d264-c4ac-4339-a805-f8f2b22d8a18 · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Inverse Scaling: When Bigger Isn't Better

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.113697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.113697Z digest=sha256:ecf151d5334c7ed5d2ba1b0efd4f4a5744c1b0e6ae5704250ee68ee36015c8df

Observation 292d24a2-6db8-4036-bb00-29d468322805 · outbound

This paper cites Large Language Models: A Survey.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Large Language Models: A Survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.344663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.344663Z digest=sha256:e4d3d3aa0ec433d2d28d52dd5be53fee404f9f6157eda906d4b0e7a3e5b685cf

Observation 84a2e209-521b-4d8d-8071-a0a6fea3e594 · outbound

This paper cites GPT-4 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.544455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.544455Z digest=sha256:c30fc5e7105d12062a0f8837fd3fdd8875978041ae7379c5018bc6ddb7150784

Observation f24bd252-aa67-4027-9ffe-31d4bcb8ed9d · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.715839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.715839Z digest=sha256:2047959d3d697f075b388c53b137f26c91e170d3858078a2b149e9dd03889de4

Observation 6b8c74b3-c7d2-4fd9-9fa0-e88b8aaf2d78 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Iterative Reasoning Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:51.881145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:51.881145Z digest=sha256:4d504788a546fd80e885eba1b68b58f5c24ca0414227c3c8177cff34313ff1ce

Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · outbound

This paper cites Self-Consistency Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.028303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.028303Z digest=sha256:a38095f4f3a2c01189c238a1045cc418841c569f5ee0a9ed975b4f935d22afe4

Observation 63350242-06bc-4505-9b17-e3593cbb3996 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.236313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.236313Z digest=sha256:d9cbdd2b0650f8af633e0d348d485700c614bb0dee3aabd372bf6615d2c8529c

Observation 5d901273-a9c0-4812-90d3-7e8131a2b228 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.442199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.442199Z digest=sha256:a061fc5c445cac61c8cd5425c307d477b49e3e4f5b2e2a5ff2f28e0b70bf7d9f

Observation 2a5a65b9-8bbd-4185-9e6d-cf89e8a9d320 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.594267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.594267Z digest=sha256:d5f4856fd841173172ff2f219793e901bcb1db1af1e5adceee7149fd68256065

Observation 5628600a-17b9-4f11-9666-e7cf21c8ed3f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.739102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.739102Z digest=sha256:f43453e7aeeca6a947a4d9b2e3d06d56c611304374d97d4c0f80da1d074c27b9

Observation a53b2819-2362-4d2d-917a-ff8642221b71 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.883958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.883958Z digest=sha256:94714bb6e7fe645137c8213938a279d2a31c967bd43b2d46d41bea751913deed

Observation e459c472-4f91-4d75-b669-3e9993087f5c · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.070330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.070330Z digest=sha256:fea25b63171a05e8883e8847d55237a4d0fc45aa6d0295301579afb8a3a45bd2

Observation a118128f-1d83-4816-a98b-b1ee6b2c5513 · outbound

This paper cites LRHP: Learning Representations for Human Preferences via Preference Pairs.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning LRHP: Learning Representations for Human Preferences via Preference Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.221253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.221253Z digest=sha256:d76972a2287d09a3888daced1104c1b7f0bb70f627a2fd6c5e044fdda13c9abc

Observation ccea8820-6185-4dab-846f-fe5503aacfb9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.362214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.362214Z digest=sha256:24d2c8ff2a81a05a939522e605bf0159689bec6bfbf05d5cb878533c9d890cdd

Observation 7009d628-dde1-4970-b138-c3801022ed5d · outbound

This paper cites Consistency of a Recurrent Language Model With Respect to Incomplete Decoding.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Consistency of a Recurrent Language Model With Respect to Incomplete Decoding

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:55.005187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.508094Z digest=sha256:0c2c64f89116bcceebd140b7ec60689c724409b6b864b00b698e39da9a6e1f9c

Observation 6261dd51-e4f9-4d11-89b5-626a1f7d92d9 · outbound

This paper cites an unresolved cited work.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:48:55.456689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:48:53.657068Z digest=sha256:1b6ff154b9aab4ae53d4fd10911302e6352deb60e86a1a8c1820aad36ac5e7a6

Observation fcbbe315-6b07-4e8d-be48-ea0ba4728290 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.850737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.850737Z digest=sha256:7fc00b64926ffa5253bea48ddd579029dc872a103fd82eb3f835542e7c7d4613

Observation 0cc5b074-990d-47eb-8731-df040fe7866f · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:53.971380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:53.971380Z digest=sha256:50535cdf2619a3d5cfa652d38f808cd405f6925e88946536a0fe1e74417137c2

Observation 12fac24c-5d97-4132-8ce7-02079af468ed · outbound

This paper cites Qwen2.5 Technical Report.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.069667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.069667Z digest=sha256:69d85e4beb70317ad965e22873c433d0b2f95b3d10ecd15f73cf01c9cb28f70c

Observation 0c0ef65f-e24d-4869-a4bc-854791757086 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.203982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.203982Z digest=sha256:f96ee7ad8ef4a6ef56aa7ebc86d35d608c27c9f523623871e1a985f928071a1b

Observation 37060bab-bcc1-490a-83ca-f6405add8166 · outbound

This paper cites RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.380705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.380705Z digest=sha256:e9e9f32807fc7dd3af92c1b1230a8b4d33854d92441256327c995138ea9bf5e0

Observation ca19ee2c-5fa0-44d7-a150-b0197044a3a3 · outbound

This paper cites Self-Rewarding Language Models.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Rewarding Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.498345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.498345Z digest=sha256:dec34ad1acab8394dfd48ee991f5607602a93eeddc4cf1749cd2d78a59bd944a

Observation 8cc37db7-4a5c-40ce-99ae-84eaceba409b · outbound

This paper cites Token-level Direct Preference Optimization.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Token-level Direct Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:54.646894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:54.646894Z digest=sha256:d21ac77bbded77fcc64de2f1b2ddb261c2aaf2734ffdee9a70d1944403adf3cc

Pith citing papers

Observation 704f80b4-7721-4f40-80da-ebb3b85067ed · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:53.845454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:a4c8573e5e15e930258ea7fea76124df784b8450a0be644b485ff38d338c8b0e