Pith. sign in

Paper Citation Record · LEDGER

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

As of 19 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2506.02519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02519 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:31.204976Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T15:15:25.731014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:20:17.443152Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa548e25-64b5-4bac-8189-c931f8c6ec3d · outbound

This paper cites 15 samples ran- domly from the test sets of each of the 5 task datasets.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning 15 samples ran- domly from the test sets of each of the 5 task datasets

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.017486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.326291Z digest=sha256:cdfa676bddbe7c0b2f3fd51a03faed255b618755e5bc77e66217351f2c247bd0

Observation 55482794-62bb-4290-bbfa-574208ceb042 · outbound

This paper cites InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:33.315717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.098866Z digest=sha256:7aa0cbfa3da4b55c9b5a75aa72ea2a6e3a9d59fb361268464003ce1f84b80175

Observation 5def50aa-9528-476c-b9af-c84f39a50e25 · outbound

This paper cites Self-Rewarding Language Models.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Self-Rewarding Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.183317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.183317Z digest=sha256:902dbe2dff0e89e095dc1cbc5304cf9a7c22f1b9a9c9ac29e85f6ec59ad3ed91

Observation c6faabc2-b02d-4592-a40a-437e7a2c97ab · outbound

This paper cites D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:26:33.194898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.255519Z digest=sha256:b80fb6a84ddf83d75fcb4f6e6ed298ec49c3c1fc43e47b176f82a5d76bdb9b4c

Observation a67ad4df-5358-4527-a057-e52970b2fe60 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:32.858269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.394283Z digest=sha256:7801d0f117f8b2d20c73de2106d1d644dc6bff9c6277c913c00449e5868d5e84

Observation 857e342e-3ed7-41aa-998a-f5fe4e62ad91 · outbound

This paper cites Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.695497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.472695Z digest=sha256:76506f680a65dde7214e645fb938a554a8ffe016bed5d7fc047140efa0642dbe

Observation 3f781dfc-39cc-462b-a3fa-3447b8f7e9db · outbound

This paper cites Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.526400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.579647Z digest=sha256:162a29c326df4fc4a5c64da9190f275d91d744d7f74bb172043d8d5b56219d3e

Observation a21e6bc8-409d-4c90-9361-936e5e07d537 · outbound

This paper cites Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.352408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.650907Z digest=sha256:cce4e35e3c2ded8250aef186495352bdb9293801699544debe0d224127f4b256

Observation ccb69efe-3f00-48f1-a32e-97fbc31b16b3 · outbound

This paper cites totally correct.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning totally correct

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.157523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.727717Z digest=sha256:db7bcc45804d9fac1cdf60d78d3d8c0189c044a5fa08bb45fde04a3bb4070397

Observation d3c17616-0ec0-479d-9c18-81c0cd731193 · outbound

This paper cites cases where one of the two ratio- nales is better than the other (label 1).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning cases where one of the two ratio- nales is better than the other (label 1)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:32.017497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.823842Z digest=sha256:30c8124e0f88e9e27b50804a6ba74ec54561d39e3053bb92e41409bc06f6ebd7

Observation f39cce1e-c00a-4e15-bf07-aa85169b0e84 · outbound

This paper cites one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g).

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.864042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.886039Z digest=sha256:d7f135aa8c05ff2e56acd380d1065d4eecb0c86317412c8be833f56adf3582a9

Observation 285d7483-7c87-492a-94f8-0bb4cd7fbde5 · outbound

This paper cites an unresolved cited work.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:26:31.751699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:30.960338Z digest=sha256:1afbae62d01401473832e19f2363cee098956009e7825eb4767ca7668415beb3

Observation dd7c32f8-8c38-42ce-9d29-3f7570953fbb · outbound

This paper cites This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.606453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:31.061800Z digest=sha256:c284b3f4f831d6ec27f33e94567d9cc212a04d8f9dc1ce7769c935b72224a330

Observation 66785ae3-7dcc-4c78-8803-f292447672f6 · outbound

This paper cites This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.500289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:31.137030Z digest=sha256:a268b40ee47ade38adc780287fb9edd27de43e50d2dff000acfbd9212b84f9c0

Observation 8dd4c841-36eb-4c66-8130-9ba0aab6223a · outbound

This paper cites The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:26:31.408499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:26:31.204976Z digest=sha256:f7b5a89f06bd33288afa1cb2fea951377292e44c7008149bfb237a730b8a4629

Observation ae6e3fa3-400b-46ea-91df-c96822058fb8 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:30.050480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:30.050480Z digest=sha256:5fda2cf6abc8e0580d09fc95191726ca9ed56c06b0ae5e512e630cee65c34d17

Pith citing papers

Observation d244099a-b9e7-4af6-930d-9eead7aded40 · inbound

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework cites this paper.

"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:20:17.445202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T15:15:25.731014Z digest=sha256:0c8b85b18a6c5024a3bdd8f21af869205c91425411722d3981b87c59f21619ad

Observation 2bdac217-2b84-47fc-b5f7-c896d9f0a724 · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.060991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:f676e71f7c0fdeae6f344224e204916f913fb01d9dbe6f56c6c4bb1a9830dce1