Pith. sign in

Paper Citation Record · LEDGER

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2506.02519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02519 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:31.204976Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:15:25.731014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:20:17.443152Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa548e25-64b5-4bac-8189-c931f8c6ec3d · outbound

This paper cites 15 samples ran- domly from the test sets of each of the 5 task datasets.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning 15 samples ran- domly from the test sets of each of the 5 task datasets

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.017486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.326291Z digest=sha256:d3f553bbe7f5173ee757bb4a7fef757072a235d940ebcf2a340e7247c8360573

Observation 55482794-62bb-4290-bbfa-574208ceb042 · outbound

This paper cites InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.098866Z digest=sha256:5e1b112e972d828bd03ce9b6f497e6b2ceb6efe5ee44045207038df73468f2e2

Observation 5def50aa-9528-476c-b9af-c84f39a50e25 · outbound

This paper cites Self-Rewarding Language Models.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Self-Rewarding Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.183317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.183317Z digest=sha256:acba0f2c4d4230317524a133a69c66bddde9025c916b6db1f8b4b6cde55e26a8

Observation c6faabc2-b02d-4592-a40a-437e7a2c97ab · outbound

This paper cites D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:26:33.194898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.255519Z digest=sha256:eec55f7e4a9a535f6182ce7e380ed021a29802d896ac1dc0f630329b0e94b9e3

Observation a67ad4df-5358-4527-a057-e52970b2fe60 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:32.858269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.394283Z digest=sha256:e485847f5e3bcbe97815be80be5f80e821ed2b4487ea2985e09884aa034e24a8

Observation 857e342e-3ed7-41aa-998a-f5fe4e62ad91 · outbound

This paper cites Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.695497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.472695Z digest=sha256:9879e6b87ea4daa5930abe0e25ee8d690462ede751ef09d7a289836b1c98f689

Observation 3f781dfc-39cc-462b-a3fa-3447b8f7e9db · outbound

This paper cites Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.526400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.579647Z digest=sha256:f5e57c1992c25b3f93f21bda0f0fb37e6854a00e5789a7fb9ec8867f722582dc

Observation a21e6bc8-409d-4c90-9361-936e5e07d537 · outbound

This paper cites Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.352408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.650907Z digest=sha256:4136ea99f6e987af31420102ad9d9b70f56f77bc053b25c596be99bc11bd8e47

Observation ccb69efe-3f00-48f1-a32e-97fbc31b16b3 · outbound

This paper cites totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning totally correct

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.157523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.727717Z digest=sha256:beb1583a11cf775ff3ed1e79fc030c8864b798c8df5a1abe8e80690e829982a6

Observation d3c17616-0ec0-479d-9c18-81c0cd731193 · outbound

This paper cites cases where one of the two ratio- nales is better than the other (label 1).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning cases where one of the two ratio- nales is better than the other (label 1)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.017497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.823842Z digest=sha256:366b52ebe5a76bff89008564bf44aa699143585f4a483713e75802a1441a3b77

Observation f39cce1e-c00a-4e15-bf07-aa85169b0e84 · outbound

This paper cites one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.864042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.886039Z digest=sha256:6d53ae96c284186d306c5b50d32b950b17489a2e57f47adad5489178bc0d0200

Observation 285d7483-7c87-492a-94f8-0bb4cd7fbde5 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:31.751699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:30.960338Z digest=sha256:825f9d6c12a85fc4536be890dab507854c7870ec0d766ccfc26454c04ab030d5

Observation dd7c32f8-8c38-42ce-9d29-3f7570953fbb · outbound

This paper cites This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:31.061800Z digest=sha256:fc7dc57aa09831458aae5e49bb4cdb54db39e7d5a30288c66f04bc0cc265fa82

Observation 66785ae3-7dcc-4c78-8803-f292447672f6 · outbound

This paper cites This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.500289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:31.137030Z digest=sha256:3dc183e0d408af0614253cb5ff8c21eb2501d1cba81e9bdb43917a55611cc3cb

Observation 8dd4c841-36eb-4c66-8130-9ba0aab6223a · outbound

This paper cites The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.408499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:26:31.204976Z digest=sha256:6949681c1b5a953a2309cc18673164e5a7f103c41023a5ca3228b700d67dea72

Observation ae6e3fa3-400b-46ea-91df-c96822058fb8 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.050480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.050480Z digest=sha256:874017eac3b7f29dd7ae31e612f7c83d8f83f5dd58a38616ae37bbf1213139e1

Pith citing papers

Observation d244099a-b9e7-4af6-930d-9eead7aded40 · inbound

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework cites this paper.

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:20:17.445202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T15:15:25.731014Z digest=sha256:6774def65c65dcc7353b2722c78d614aba6961c4e9dbc2da40fc7e0ae6132462

Observation 2bdac217-2b84-47fc-b5f7-c896d9f0a724 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.060991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:3f5b5591fc2cd3b64e880a3e4ce4d934e6b7294f3808916d03aea6787b111c2f