Pith. sign in

Paper Citation Record · LEDGER

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2506.02519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02519 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:31.204976Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:15:25.731014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:20:17.443152Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa548e25-64b5-4bac-8189-c931f8c6ec3d · outbound

This paper cites 15 samples ran- domly from the test sets of each of the 5 task datasets.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning 15 samples ran- domly from the test sets of each of the 5 task datasets

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.017486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.326291Z digest=sha256:1829a5336e87828a26260a779a90d7e54c1e6ef3ec857b98bddc3d96bccc30ec

Observation 55482794-62bb-4290-bbfa-574208ceb042 · outbound

This paper cites InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.098866Z digest=sha256:b9f318397af1b6cedc0b7c14a9d36acb30a19edc6c04373d6cb7369fc453ceab

Observation 5def50aa-9528-476c-b9af-c84f39a50e25 · outbound

This paper cites Self-Rewarding Language Models.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Self-Rewarding Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.183317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.183317Z digest=sha256:acba0f2c4d4230317524a133a69c66bddde9025c916b6db1f8b4b6cde55e26a8

Observation c6faabc2-b02d-4592-a40a-437e7a2c97ab · outbound

This paper cites D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:26:33.194898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.255519Z digest=sha256:1aeeb5029326f38b669ef57b546b70a7f5852ebabe3ea85ffc41fde5a5757ae8

Observation a67ad4df-5358-4527-a057-e52970b2fe60 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:32.858269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.394283Z digest=sha256:6771f3ee2bf0f59d536e24886e93136212abafeb77d7d73db04a9bb51e25fb11

Observation 857e342e-3ed7-41aa-998a-f5fe4e62ad91 · outbound

This paper cites Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.695497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.472695Z digest=sha256:6691e15ee2a15fa22f36f6be390cd8108d61483ceac1778d96db9fe42c528831

Observation 3f781dfc-39cc-462b-a3fa-3447b8f7e9db · outbound

This paper cites Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.526400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.579647Z digest=sha256:dfd6e9d1bb89c8342b810e31be0921ebbb6391f267d6cbf6f72c1836695d3d75

Observation a21e6bc8-409d-4c90-9361-936e5e07d537 · outbound

This paper cites Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.352408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.650907Z digest=sha256:afefd0341e795de7b27583e78b5761fcc83f20bdabdb5fc3e3d39d177e8b3431

Observation ccb69efe-3f00-48f1-a32e-97fbc31b16b3 · outbound

This paper cites totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning totally correct

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.157523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.727717Z digest=sha256:a23929395eb908df392760ca29da207e24675a98a89da399ac3bfc50e2d183f9

Observation d3c17616-0ec0-479d-9c18-81c0cd731193 · outbound

This paper cites cases where one of the two ratio- nales is better than the other (label 1).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning cases where one of the two ratio- nales is better than the other (label 1)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.017497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.823842Z digest=sha256:2faf8a4b8e3e47f30022ac2bafb386fc41fa3d5bb2dc21a82ca1444b453999bd

Observation f39cce1e-c00a-4e15-bf07-aa85169b0e84 · outbound

This paper cites one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.864042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.886039Z digest=sha256:8d448c6595112c6f6d2284ed7ca672089bda17ce57c52ed2e1e7c8664f2c76bc

Observation 285d7483-7c87-492a-94f8-0bb4cd7fbde5 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:31.751699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:30.960338Z digest=sha256:0d667978f44a181814cfd77beff2a085c0ff7199b692f663655f8dd927dc46e2

Observation dd7c32f8-8c38-42ce-9d29-3f7570953fbb · outbound

This paper cites This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:31.061800Z digest=sha256:6416b7262f91cdd8ea07eeaf0db2c52a250948689614684550ea6d1426674a85

Observation 66785ae3-7dcc-4c78-8803-f292447672f6 · outbound

This paper cites This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.500289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:31.137030Z digest=sha256:d1f026b4b9f62fc472f615338be576177ee69d677364ca00727233866892f9e4

Observation 8dd4c841-36eb-4c66-8130-9ba0aab6223a · outbound

This paper cites The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.408499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:26:31.204976Z digest=sha256:1f31cd78f56a272aa1daaabd828f9547a371d391dbe722f81296d0a4d4f2ab94

Observation ae6e3fa3-400b-46ea-91df-c96822058fb8 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.050480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.050480Z digest=sha256:874017eac3b7f29dd7ae31e612f7c83d8f83f5dd58a38616ae37bbf1213139e1

Pith citing papers

Observation d244099a-b9e7-4af6-930d-9eead7aded40 · inbound

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework cites this paper.

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:20:17.445202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T15:15:25.731014Z digest=sha256:4963961e246f6b1c3d3b1ee1e2fdfa2bd875dfbcebe436e0ae426bb9b80d550a

Observation 2bdac217-2b84-47fc-b5f7-c896d9f0a724 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.060991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:eaf64c11c8f115ec891b69367d58cca704e67462e341eaeca1498ae6e9ea0f5a