Pith. sign in

Paper Citation Record · LEDGER

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

As of 6 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2604.07650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.07650 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:13:04.305435Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:29:13.728227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T09:00:59.794538Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact16
  • verified fuzzy9
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45ccfa34-725b-414e-8136-939cb84314d4 · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles gpt-oss-120b & gpt-oss-20b Model Card

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.913602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:808b8191ae6fafac8fb2d1cf0a4af304f5ed69ba27a9ca035628c2b4cc00b46b

Observation be98a279-555f-4c17-ac85-36eacbdc39c2 · outbound

This paper cites anthropic.com/news/claude-3-family.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles anthropic.com/news/claude-3-family

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.606013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:7e41f4e13d4d7c3f363efa44c524644c72ca0cc7c73935d7cafd0b2c0c6e3b08

Observation 2cb65387-d213-4111-bdaa-2acbc2c3cffc · outbound

This paper cites Qwen Technical Report.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Qwen Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.955929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:cee9735829a880ea511eb82535025b88817cf6e16398ff020b19620b4001046c

Observation 05dfea4b-ad56-4640-9520-b4d45d004aa2 · outbound

This paper cites Beyond the surface: Measuring self-preference in llm judgments.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Beyond the surface: Measuring self-preference in llm judgments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.612265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:df667e0c2469a632867c90fa162c8168a5b34153e256f3d53470fafc56e14744

Observation 5e083a0e-f225-4d4c-abcc-945affb1a190 · outbound

This paper cites Investigat- ing data contamination in modern benchmarks for large language models.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Investigat- ing data contamination in modern benchmarks for large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.619461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:5fd182aced6f95183b744d70dee53f3512d3812a7b2e939ef32733b57cfef062

Observation 08c57516-465f-4a13-9bd6-b1dd58f6c11d · outbound

This paper cites Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.589427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:7581da3fe235f626a04e12209b974211095718fe0d1be37ffe72d7086b1fa169

Observation 9b32f6f8-9314-4fa8-b5da-cb107f2ce88e · outbound

This paper cites The Llama 3 Herd of Models.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles The Llama 3 Herd of Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.959742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:67c566c9fd52a42042b595c130a57dd5fa4e7dc0f3f89458e9c9d2ff669534a6

Observation d6d876d2-44ea-4f21-ab4c-5c99360345aa · outbound

This paper cites GPT-4o System Card.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles GPT-4o System Card

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.857547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:f203524ac31666d1b596aa9cd4e310b71915514318fac4de93039f7953c46e97

Observation aba21ea2-547c-4c03-946f-c625043e6dd7 · outbound

This paper cites VicunaNER: Zero/Few-shot Named Entity Recognition using Vicuna.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles VicunaNER: Zero/Few-shot Named Entity Recognition using Vicuna

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.963826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:2873ff5feb343195810b226920aaac7c3081e125e89f0510df5b587d0ad3d217

Observation 135829f9-52a7-4128-a77b-d841e8affac8 · outbound

This paper cites arXiv preprint arXiv:2502.01534 , year=.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles arXiv preprint arXiv:2502.01534 , year=

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.892444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:d28272cd26ab47b2b31cebf2e1e9a049a57def817013d43bcb9ffa675327822a

Observation 0e1387c6-a6b4-487d-82fa-df1282b8a272 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.887342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:2f6f0ca14aed69f0508d576bfe4b47a86354e416f3fe323493c529b536b62be4

Observation f93f71fb-5673-46c0-80a8-bc6a68fe120f · outbound

This paper cites Accessed: 2026-03-31.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Accessed: 2026-03-31

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.615373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:a1af8825dc1f0e9c3de81b0cf3d117877eb66eb7b660ee874cbdfec11be4de7a

Observation 5e3121e0-dd79-4c48-b006-410650061054 · outbound

This paper cites Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.579752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:5d9c5886ed92423b53173c37b0f9585903e794d56e426b2a5a2dc88e67eab7ad

Observation a8509044-21a0-43fe-acbc-a89901447a6d · outbound

This paper cites Instruction Tuning with GPT-4.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Instruction Tuning with GPT-4

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:04:18.193148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:18d0f1604aa72643bc59743875584fb5def80d1522654cd88b912cc86e799971

Observation ea98a32b-8e6f-4555-8ef5-a0dc014494af · outbound

This paper cites Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.583568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:6d42f3cd05433490354f2e87d5251bc0ebbc5adda74af3023a5ffced1691b27a

Observation e6a2227d-c024-43ef-906f-6b345c260a27 · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Detecting Pretraining Data from Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:07:25.061786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:6556598b31b0bc19ec26f451ac9b3a23459b19b55dd5be192f6fa07f639ab776

Observation 4cbde1b0-05ca-4a33-8170-13613f8e2f96 · outbound

This paper cites OpenAI GPT-5 System Card.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles OpenAI GPT-5 System Card

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.948246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:1e6cf97b15df7cf98605c5e70e29458365f0248018ad0c57f1d544ab641b16b2

Observation 5e5771af-2a38-4203-88b1-f67c0b69f26d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.872920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:8cdbef6e315dc004bcc31d909cb6a9f1591618f26fce269db1ac4070c30b8b09

Observation bc23bf87-020e-4918-b012-c9bee0b681a5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:20:58.944832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:8a9cb6eed102bc55af543c296ce4c84eb1b97e011777810cc3abb0576c43c016

Observation c22ab2d3-dc15-4fa6-8b31-c76c8a472981 · outbound

This paper cites Who taught you that? tracing teachers in model distillation.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Who taught you that? tracing teachers in model distillation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.602467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:c2168e254a065262ddc85ddc650fb4a36e76af30da361b5bb1920e226a09dc97

Observation dec08c7f-4854-4589-bb93-55abe983f543 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:20:58.903293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:c54e88ba7d7659fbef25a3c91c1e7c6d757fbfba384f97acd6920d80865a93e3

Observation ddf47d33-7cd9-4f04-934d-022046108481 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Self-Preference Bias in LLM-as-a-Judge

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:46:32.590484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:fd7386c8f97cf20f9f3b3fc18d962380bfd2d629fdc483f6d1851c07da2775e0

Observation ed014e5e-1fe5-4460-9cdd-94128788d928 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Benchmark Data Contamination of Large Language Models: A Survey

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.376209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:e4c63031a1c1c167b689b81b8cca427d96bca998177085a6edc3a06686d9366c

Observation 2f074eef-00b6-4cf4-88dd-c85cc68b3796 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:00:24.757706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:7b33f9da1df546202e7f273f44d0c6ec54862d2131cf191becc56a0ec7d6aaba

Observation 227c3cac-82a2-40f0-be47-0d92e3231ded · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.932401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:9983adbc2928a0ec7844d177c8e6a10df99533fcf87f291270b9e82b7cb940d8

Observation dda0d095-5306-4d66-8f7a-4072d4d5d2e6 · outbound

This paper cites Instruct.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Instruct

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T11:31:46.597778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:0fd57a8fdfb16f4707d77fdd4e574eaf6924f0a1819492ea91c5b900a6533250

Pith citing papers

Observation a544d1b3-1055-439b-ac59-126ffae119ee · inbound

Knowledge Is Not Static: Order-Aware Hypergraph RAG for Language Models cites this paper.

Knowledge Is Not Static: Order-Aware Hypergraph RAG for Language Models How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:00:59.799891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:20:31.981340Z digest=sha256:6931e6cd8f477bc9d64e2202823933898a2bc21644d588dfe1dc1c17da6173cf

Observation 4e42e32a-aaa7-4861-80d4-4c3e1a07cccd · inbound

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage cites this paper.

Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:29:13.728227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:29:13.728227Z digest=sha256:822d5f102e55f91f2dd53409db72289cf9d710cdb995dfea1cc4f12ea62a5b52

Observation 22513b24-f95b-439b-8871-aae44cf9808d · inbound

Grounding latent algorithm routing in transformer reasoning cites this paper.

Grounding latent algorithm routing in transformer reasoning How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:25.242442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:25.242442Z digest=sha256:ae520c63065f201c8c411376465fb6ab82691799f2566dfb5f30ddcfb4981d33