Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T17:13:04.305435Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2604.07650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T17:13:04.305435Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:29:13.728227Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-11T09:00:59.794538Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 45ccfa34-725b-414e-8136-939cb84314d4 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles gpt-oss-120b & gpt-oss-20b Model Card
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be98a279-555f-4c17-ac85-36eacbdc39c2 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles anthropic.com/news/claude-3-family
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2cb65387-d213-4111-bdaa-2acbc2c3cffc · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Qwen Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05dfea4b-ad56-4640-9520-b4d45d004aa2 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Beyond the surface: Measuring self-preference in llm judgments
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e083a0e-f225-4d4c-abcc-945affb1a190 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Investigat- ing data contamination in modern benchmarks for large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 08c57516-465f-4a13-9bd6-b1dd58f6c11d · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b32f6f8-9314-4fa8-b5da-cb107f2ce88e · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6d876d2-44ea-4f21-ab4c-5c99360345aa · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles GPT-4o System Card
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aba21ea2-547c-4c03-946f-c625043e6dd7 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles VicunaNER: Zero/Few-shot Named Entity Recognition using Vicuna
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 135829f9-52a7-4128-a77b-d841e8affac8 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles arXiv preprint arXiv:2502.01534 , year=
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e1387c6-a6b4-487d-82fa-df1282b8a272 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f93f71fb-5673-46c0-80a8-bc6a68fe120f · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Accessed: 2026-03-31
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e3121e0-dd79-4c48-b006-410650061054 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8509044-21a0-43fe-acbc-a89901447a6d · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Instruction Tuning with GPT-4
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea98a32b-8e6f-4555-8ef5-a0dc014494af · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e6a2227d-c024-43ef-906f-6b345c260a27 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Detecting Pretraining Data from Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4cbde1b0-05ca-4a33-8170-13613f8e2f96 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles OpenAI GPT-5 System Card
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e5771af-2a38-4203-88b1-f67c0b69f26d · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc23bf87-020e-4918-b012-c9bee0b681a5 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c22ab2d3-dc15-4fa6-8b31-c76c8a472981 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Who taught you that? tracing teachers in model distillation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dec08c7f-4854-4589-bb93-55abe983f543 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddf47d33-7cd9-4f04-934d-022046108481 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Self-Preference Bias in LLM-as-a-Judge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed014e5e-1fe5-4460-9cdd-94128788d928 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Benchmark Data Contamination of Large Language Models: A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f074eef-00b6-4cf4-88dd-c85cc68b3796 · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 227c3cac-82a2-40f0-be47-0d92e3231ded · outbound
How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Don't Make Your LLM an Evaluation Benchmark Cheater
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dda0d095-5306-4d66-8f7a-4072d4d5d2e6 · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a544d1b3-1055-439b-ac59-126ffae119ee · inbound
Knowledge Is Not Static: Order-Aware Hypergraph RAG for Language Models How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e42e32a-aaa7-4861-80d4-4c3e1a07cccd · inbound
Harnessing Disagreement: Detecting Correlated Agreement Blindness in Multi-Agent Triage How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22513b24-f95b-439b-8871-aae44cf9808d · inbound
Grounding latent algorithm routing in transformer reasoning How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.