Pith. sign in

Paper Citation Record · LEDGER

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2411.16797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16797 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:28.151808Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:12:14.146302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:12:24.245674Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95bc0cae-529f-4883-8a40-ca69200b221a · outbound

This paper cites GPT-4 Technical Report.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.011334Z digest=sha256:1c2d42bf49436bf68f4fe983f15de132fad51e6e89156c84de0c1024efc5f870

Observation a87d9bb1-b8f4-412e-a86e-1c7728b940ef · outbound

This paper cites LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.045862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.045862Z digest=sha256:2f004a94188eaa1c289c3af418b4d6d403783fabbe63ae2d1bbac47adea8b10a

Observation 19300bb6-215f-4525-8c1b-dfba43ff98a7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.057753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.057753Z digest=sha256:c9fcea60037d59eec1081be808bc6c60a1a7e8b5343395810f2c9cb9fccfa26f

Observation 61f6a845-ddc3-4742-a7ea-4ee0279e2fe5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.069139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.069139Z digest=sha256:43b10de18a88d2bdadf3cd771b7ddba2a446499f0fa7b0ce9cb5b045ec55d514

Observation cae75261-3045-4195-a7a9-874e2055e0ce · outbound

This paper cites Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI conference on human factors in computing systems, pages 1–16,.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI conference on human factors in computing systems, pages 1–16,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.583008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.074657Z digest=sha256:cb44b7438b4312e72384e4029dc3ad67725aaf6362c6513b551674b2c9ba6938

Observation 53caf2f2-4260-4faa-9023-f7f2b949041c · outbound

This paper cites Harnessing the wisdom of crowds in wikipedia: quality through coordination.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Harnessing the wisdom of crowds in wikipedia: quality through coordination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.565545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.086707Z digest=sha256:5c628bcfe5c5db0c72435e05032e8df0871694fe16f2860d732153e0e7b1bae9

Observation 4e4c0ce4-9ec1-4d48-8447-56aa6fd17594 · outbound

This paper cites Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.096521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.096521Z digest=sha256:633582b52a917659e93e04342e208412d68b38f10bf59e300d5f5d9a8c34448a

Observation ba4b6d1c-96ff-45b9-b14a-67e25f3d780f · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Language Models are Few-Shot Learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.101763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.101763Z digest=sha256:33f253584d860ac2071327f9687c79ed7d304cf8044011fdc3b50fc71427f9a1

Observation 27e43cd9-f0eb-4806-acb8-f83fa111fd6c · outbound

This paper cites Brent Mittelstadt.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Brent Mittelstadt

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.547773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.107527Z digest=sha256:709b550ae00d0f9b0837c48c034de443f9efbe23521aa55f2e7278b9e767e734

Observation 34c92536-b659-430a-97b9-ef7c8ab81621 · outbound

This paper cites Mitigating bias in algorithmic hiring: Evaluating claims and practices.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Mitigating bias in algorithmic hiring: Evaluating claims and practices

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.530508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.117355Z digest=sha256:4d12e02c3f5698ff4f1b82f7cb0b38e953c34d8c4a4aa8183c15323aedd7b4fa

Observation 503f515d-7c8f-469f-9462-2aea3a1d1f8f · outbound

This paper cites Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.122091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.122091Z digest=sha256:1fc936e3d9e606e49aeb41a02ddff4b1baff3260c0ea65ab2bdc6b1db35df1cb

Observation 2601fffa-15fa-4834-b521-b3049b48088b · outbound

This paper cites Galactica: A Large Language Model for Science.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Galactica: A Large Language Model for Science

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.127021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.127021Z digest=sha256:6478b4e231f252825c73733e1017d7f99511a6abc2ec80e903f6ae67b26be505

Observation e6b5a435-7c4b-4712-a09f-a082c53473a9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.131826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.131826Z digest=sha256:7c437a4c7180fa2ef06cf2d60b03984450a4d32f9119d0c0f5a12d15d70d494c

Observation eb19abfe-7bb2-4ca8-933e-5ed0ada9733e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.136780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.136780Z digest=sha256:7158453f2e29f5700f452213c61a2cd214189db979af4923542df63324dd9142

Observation 54595ea8-8143-4981-a0fb-8dd4dd0d188d · outbound

This paper cites The collective intelligence of random small crowds: A partial replication of kosinski et al.(2012).

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models The collective intelligence of random small crowds: A partial replication of kosinski et al.(2012)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.513679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.142188Z digest=sha256:193ca7780a6a6c714081719676bf03779f6ea0928752201e2efed4dd3075d58b

Observation 6fe34682-5c77-4a20-96ab-403d2b2a1928 · outbound

This paper cites Deep learn- ing for computer vision: A brief review.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Deep learn- ing for computer vision: A brief review

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.495559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.146899Z digest=sha256:4a146c36fb73eff923b3a67094b9d77df48339323422e072df37a5efd3f1bb03

Observation 12b2009c-b670-415f-9e3a-f3a2d16cb20f · outbound

This paper cites Large Language Models and Causal Inference in Collaboration: A Survey.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Large Language Models and Causal Inference in Collaboration: A Survey

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.091534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.091534Z digest=sha256:cf9bc721a972362ea3050a8bc19e243f567655ac9f296583012f7795a5c91e8c

Observation d569f57f-47d3-42e4-8597-0465ff4f3ff7 · outbound

This paper cites Accountability of AI Under the Law: The Role of Explanation.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Accountability of AI Under the Law: The Role of Explanation

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.051711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.051711Z digest=sha256:1fd1462a8f826f665f3562a7c5fa570444ef38d451d927f0f0e9b312fdc9a501

Observation 0854f7bf-941b-457e-9c80-7013be8bdfc7 · outbound

This paper cites Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.081674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.081674Z digest=sha256:e66935540704af91fe35aec9d45c017eea772e6fe948e8326fc521e66a6a4910

Observation 3c86ee35-6580-4e3e-a5d9-06b33a76a49f · outbound

This paper cites Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.151808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.151808Z digest=sha256:eaba60fb79409fe2350947be1a236e6816dfc54f59d40a4774b548b64911651d

Observation 9d7ca5ef-e005-41ae-b3dc-9c4b0dc066d4 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.028565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.028565Z digest=sha256:9ff3a355f7a0d0e81f0c28bcbc7c1922ab18ff4a2908f2902781105a8f8d4cc4

Observation f8771898-a9fa-4b81-966d-fd942435a37a · outbound

This paper cites Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.112281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.112281Z digest=sha256:522db6534d5b3347e084d2b8c27cface1992261e4e19051353d7b747a3a6071c

Observation 48530819-7c65-4fbe-ab5d-a983f3184455 · outbound

This paper cites Open Problems in Cooperative AI.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Open Problems in Cooperative AI

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.040010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.040010Z digest=sha256:ccb1ac05cb987b32e375ee71243bf59cbe0f96a978f5261181ff05ddd72336cf

Observation 05d23a2e-e037-42cb-94c6-a462d9c0d66b · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models On the Opportunities and Risks of Foundation Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.034743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.034743Z digest=sha256:ec410ac9b69666e89f564915ed5a4668a04769c81c45c94ae42525df3c73b9f4

Observation 09915226-6215-4d05-90b2-57f47b6f486b · outbound

This paper cites Fairness without demograph- ics in repeated loss minimization.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Fairness without demograph- ics in repeated loss minimization

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.599437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.063661Z digest=sha256:f11c2039261f67efa6e96a0ca0e17c39df3525cbb90638ccc7a7dbed0f2ace5a

Observation 6d595faf-8b1c-4b46-bc1c-6a69ebc3f796 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.017481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.017481Z digest=sha256:1f3375688b37614fb88cb5cd4d8e7f12955509e88c39be4f20ac07f05ebbde0a

Observation 42229a8a-2355-4e89-9290-239f8a47b5af · outbound

This paper cites Anthropic.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Anthropic

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.626772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:21:28.022885Z digest=sha256:d8e0b17d45974b80f49575c7517b1dba635d6321fd9758e2369ebc7b55e74ce3

Pith citing papers

Observation b6c747e7-9df5-4b07-b6d4-2880cad55f2f · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.735769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:0672b299ffca81449e7e6240da6eb46b16a36af4f42de56f27869c4ddbd1b303

Observation aed2077d-01a0-4a6c-ae0b-56bd1a8115af · inbound

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment cites this paper.

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.247068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T17:12:14.146302Z digest=sha256:8d2b5103b4d76d8d7e4d452e432672cfc88022cd333a76e005777aeeab5d5490