Pith. sign in

Paper Citation Record · LEDGER

Self-Consistency Preference Optimization

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2411.04109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04109 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:59:11.533287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:33:19.400166Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4cb839f-9b51-4a5b-b4e7-7eda387fc184 · inbound

Large Language Models Can Self-Improve in Long-context Reasoning cites this paper.

Large Language Models Can Self-Improve in Long-context Reasoning Self-Consistency Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T21:59:11.533287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:59:11.533287Z digest=sha256:167f63633c3b56a15a1ba5a4c57b4cb88a8e9237d2b480bbb70be862602a9d13

Observation 3a3ae537-c3f5-4657-af8a-3492ada82995 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models Self-Consistency Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.052464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.052464Z digest=sha256:abfd8d1b6e5ae3c24ea4e333bd93806f06035cd1414e53848ca44aa1049a7ffd

Observation b49c8399-1627-4437-b1af-e00e4c418b71 · inbound

Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel cites this paper.

Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel Self-Consistency Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:48:38.561488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:48:38.561488Z digest=sha256:bef484a5dafa1186732b17afc9d16cbc4bdfe59b72bf0505fd13a512db42059d

Observation ca1fc7a1-a447-468d-bf4a-3bd4d4669ab3 · inbound

Learning to Generate Unit Tests for Automated Debugging cites this paper.

Learning to Generate Unit Tests for Automated Debugging Self-Consistency Preference Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:52:11.716082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:52:11.716082Z digest=sha256:1433fccfe78b109ea1f882da3726e9a09ee66408bb0b1294b18d962b902fc770

Observation 191a122f-9361-4dbe-a7cd-8653068959d5 · inbound

Self-Training Large Language Models with Confident Reasoning cites this paper.

Self-Training Large Language Models with Confident Reasoning Self-Consistency Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:15.120818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:15.120818Z digest=sha256:27d4df2833255d4d44f48d329c52f435f135bfa187f4822a44771e35b50dd123

Observation 7c5d64e3-bd74-4310-87b3-46498df3ea7f · inbound

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning cites this paper.

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning Self-Consistency Preference Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:52.028303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:52.028303Z digest=sha256:413e385768a7681626437b3ddff5baba9eba0ebdb1d927a9ba85ee6cea6e2fcf

Observation 6ed41c72-6d8e-4c2f-b043-a25ab1fadf0e · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Consistency Preference Optimization

Reference 293

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.327311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.327311Z digest=sha256:e77468d216c497fe4ea73041892940e4ffa610d9151ca937a79eba1354fbaf6e

Observation 6e28f834-ee8a-4d65-81b9-eb47024abda8 · inbound

PRInTS: Reward Modeling for Long-Horizon Information Seeking cites this paper.

PRInTS: Reward Modeling for Long-Horizon Information Seeking Self-Consistency Preference Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:37:47.736132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:37:47.736132Z digest=sha256:4e37744e7eea2ceccfdbfe960566df52b371a97b7d12c27be72155199ab09dc0

Observation 09e4dea1-2b7c-4614-93da-5a41c3457393 · inbound

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model cites this paper.

VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model Self-Consistency Preference Optimization

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:46:06.312666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T15:10:16.533927Z digest=sha256:172824ab54aaabb73abdbf289eecc201ad0f2d5581176d37dde40905e6b05c57

Observation a25e04a3-5dd9-4280-bb24-9d877e87244c · inbound

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs cites this paper.

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs Self-Consistency Preference Optimization

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:45:59.758355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T16:58:10.013475Z digest=sha256:e6aa22b51a0011fa120d93491945469a34158f593f011e829f8babf77ee566a0

Observation 0cc65ef4-4238-4b39-8c76-c2cdaf7592f5 · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.281607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:dac661cce2d70668d31699e28c80f7cd7126102288fca66216c74eb0dd22a599

Observation 6bd566af-f63d-4dde-b37d-7a3dc918f11b · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Self-Consistency Preference Optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T17:52:43.510864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:52:43.510864Z digest=sha256:7986a9ad1d4565462db5d93de334806b771e370cdd2e4e549235331cf365fa65

Observation e4b6e54a-bb67-468a-a97d-a49592c36766 · inbound

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era cites this paper.

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era Self-Consistency Preference Optimization

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.402214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T13:29:11.509966Z digest=sha256:b09e3a7abf7fc005c8bc2ea11d29c083a9faa3ea0dccaf4bf52185f1e7174eb1

Observation 2d29a922-e717-4ef2-9cc0-0f2b45dcf30a · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:8c4ff45c1d45f8afb4cff3c813484838caa4cb29572107a45ca2e4f56661a73c

Observation c55cb8ae-dc47-45c7-9e27-fe2c1bb069fc · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Consistency Preference Optimization

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:42.168233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:42.168233Z digest=sha256:db22f091fa90ec756b8016ed7ce0fadf2dc5d3713188426701a0cb1511fca531

Observation 818058ae-800e-4e9e-b1d8-8a1fc22eb194 · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Self-Consistency Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:35:01.444948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:35:01.444948Z digest=sha256:849fab0e587ce5e3b4bddf8e904a804b79674f65cb9f3a5300e4d5fbf737b844

Observation 4576b5c3-1f36-49ca-87d9-7a8bfed8dd8c · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Self-Consistency Preference Optimization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:42.347486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:42.347486Z digest=sha256:9f9f406a1b1356941a0d3a3def45e5320eedd3960a4633bfbda059c878f5989e

Observation 4caa0439-287f-4296-bd16-6aa4e452b899 · inbound

On-Policy Self-Distillation without Any Supervision cites this paper.

On-Policy Self-Distillation without Any Supervision Self-Consistency Preference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:42.234006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:42.234006Z digest=sha256:299328f8233c9881c1884511dedec56b601c97a334ae8fb74310998d8965889c