Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:31.204976Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2506.02519.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:26:31.204976Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T15:15:25.731014Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T15:20:17.443152Z
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa548e25-64b5-4bac-8189-c931f8c6ec3d · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning 15 samples ran- domly from the test sets of each of the 5 task datasets
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 55482794-62bb-4290-bbfa-574208ceb042 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning InAdvances in Neural Information Processing Systems, volume 36, pages 11809–11822
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5def50aa-9528-476c-b9af-c84f39a50e25 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Self-Rewarding Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6faabc2-b02d-4592-a40a-437e7a2c97ab · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning D Dataset Samples Details of datasets were discussed in the ‘Experi- ments and Evaluation’ section in the main paper
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a67ad4df-5358-4527-a057-e52970b2fe60 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 857e342e-3ed7-41aa-998a-f5fe4e62ad91 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label out of 0 or 1 such that 0 means that the final rationale is totally wrong; and 1 means that the final rationale is totally correct
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3f781dfc-39cc-462b-a3fa-3447b8f7e9db · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Provide a label of 0 or 1 where 0 means that none of the rationales is better than the other and 1 means that one rationale is better than the other
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a21e6bc8-409d-4c90-9361-936e5e07d537 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Definition of Metrics Estimated from Human Labels Different rationales were presented to human evaluators in jumbled order to avoid biases while comparing rationales
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ccb69efe-3f00-48f1-a32e-97fbc31b16b3 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning totally correct
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3c17616-0ec0-479d-9c18-81c0cd731193 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning cases where one of the two ratio- nales is better than the other (label 1)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f39cce1e-c00a-4e15-bf07-aa85169b0e84 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning one of the generated rationales is judged better than the other generated rationale (comparing R1g and R2g)
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 285d7483-7c87-492a-94f8-0bb4cd7fbde5 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd7c32f8-8c38-42ce-9d29-3f7570953fbb · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This means that employing two variants of same LLM is useful to obtain distinct and diverse rationales which are useful to improve quality of preference data for DPO
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66785ae3-7dcc-4c78-8803-f292447672f6 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning This shows that our choice of using likelihood of final GT answer for selecting winner ra- tionale aligns with human preferences and is suitable to obtain the preference data
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8dd4c841-36eb-4c66-8130-9ba0aab6223a · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning The results are summarized in Table 14, where Table 14: Performance comparison of COLLATE with SPIN on additional benchmarks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae6e3fa3-400b-46ea-91df-c96822058fb8 · outbound
Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d244099a-b9e7-4af6-930d-9eead7aded40 · inbound
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2bdac217-2b84-47fc-b5f7-c896d9f0a724 · inbound
Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.