Pith. sign in

Paper Citation Record · LEDGER

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2607.13347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13347 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:30:04.755723Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T09:29:49.561922Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6970fad3-7f24-4f5d-97be-ec2017a8c22c · outbound

This paper cites Surikuchi, Ece Takmaz, and Alberto Testoni.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Surikuchi, Ece Takmaz, and Alberto Testoni

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:29:59.460677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:29:59.460677Z digest=sha256:451469a2b816b5f77baf6b3b18805ee092345d1439366fa0e01822b49d4599e2

Observation 63ed05f1-1f15-4de9-86e2-86ab607fc275 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:29:59.668652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:29:59.668652Z digest=sha256:f8bff5ba2e9523abd15e5fad6a059545c231a1ee24fa7b13861a0e15c412b96c

Observation c4fdd783-9c94-48e1-8f14-50796b705ed2 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Training Verifiers to Solve Math Word Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:29:59.811281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:29:59.811281Z digest=sha256:90bf5eb48dbdf9c1531d4ff971991ac1690fe47679aae1daeaae80c2b0914c0b

Observation 063cdf7d-ef8a-472c-9907-89a12180aee8 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Scaling Laws for Reward Model Overoptimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:29:59.895501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:29:59.895501Z digest=sha256:2f7c89b495079334425d979b0f4971ca0c74edc6e9d79a6d64e8cbf44236f31d

Observation 85e6d35e-0e8f-41e9-b60d-136d2941a2ac · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.081154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.081154Z digest=sha256:4e73659516d62d93da697efd2920264aef81553229ba5c27000175d1b699eb2e

Observation 67a1e79d-721f-4446-92fc-9eae1e6e064b · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.181274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.181274Z digest=sha256:786c750a3a454941ef9f2aeb0225f93a3cb68c77d544c5d825fbcf4f909d5c9d

Observation 8ac2bf8b-29a6-473b-b460-be714b2779b0 · outbound

This paper cites Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.268539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.268539Z digest=sha256:dd6d14ae3755a0b068160f05ccbb41ff50ee5b94f6ab11cabe85e148533073ee

Observation 9ee16917-487c-46ed-9f5b-f1df8f0db0db · outbound

This paper cites Beyond String Matching: Semantic Evaluation of PDF Table Extraction.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Beyond String Matching: Semantic Evaluation of PDF Table Extraction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.376334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.376334Z digest=sha256:eeb8bb020bcd2a3a09b44407633193d2e183b921ef45b88edc0297728edf0fb2

Observation bd79bc58-d3bd-4eb0-8bed-8afeccf7dde5 · outbound

This paper cites An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.492775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.492775Z digest=sha256:aa30b1989d91fe0ae3ea100e100540917a154af92ffafa4af223fba962c59368

Observation e25c95bb-291f-45e8-bb72-6fcbf4f87306 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Large Language Models Cannot Self-Correct Reasoning Yet

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.658188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.658188Z digest=sha256:ebb5a8e9233ddbdbca1ed240713d4cb3080868e79b590d3ff5c5437646515e41

Observation 06ab8d62-8c7f-4833-b970-fc3287d07c32 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.788948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.788948Z digest=sha256:8e92e9b969a0e984bac2717473acbd5b2a4b0cf700d8e08c66f375a8ee9ea3a8

Observation eb4fb963-1868-4733-af0d-eaa11a3e8fa9 · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.881954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.881954Z digest=sha256:2d6a103119b3e1a5b222c3931efcf7326bc552432c7f4b044a7dd666ecf126f4

Observation 51fc8f21-76a6-47ad-8ce0-91858d9c8e51 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:00.967761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:00.967761Z digest=sha256:1835dc1a3760787931895bc64312973f9cc0a62cce69784070388334639ee8ce

Observation dd641cae-5a74-4c96-907d-06d8423e45df · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.190743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.190743Z digest=sha256:06aaf90196805dc49f20072737a8a1f07b3c0939c9fbaf2c028e96048e58aea3

Observation 27d41115-94a2-4d06-9ebb-9d17b1f05689 · outbound

This paper cites Let's Verify Step by Step.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.341930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.341930Z digest=sha256:62dc8e576e86ca558df91392fdf3489236292aa30592e530526acdda3a878628

Observation 8fbdd9ef-86d9-44ae-ba40-b94326a2a7c1 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.434969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.434969Z digest=sha256:9e3ecfb3ee1ad4277b736fb7adb56ddc73fffb6dc46cf0179feb0789becda612

Observation 4be6b76c-3030-494e-91d8-6ee89653d025 · outbound

This paper cites When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition When Does Verification Pay Off? A Closer Look at LLMs as Solution Verifiers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.547918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.547918Z digest=sha256:bcfc697929d4504e0d70996f93598e3d6fee8cbb34593cba5c2d667867e0ffc4

Observation 0228b135-2bd8-4f0d-908c-0136c0249b06 · outbound

This paper cites Optimized Table Tokenization for Table Structure Recognition.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Optimized Table Tokenization for Table Structure Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.657101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.657101Z digest=sha256:4b957c56ca02e63cb1c7000dbfb49764a736a4b2e2111885770fec62ba5290ab

Observation 696412b8-f9c7-45ae-9825-851d42cedbc2 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Self-Refine: Iterative Refinement with Self-Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.760982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.760982Z digest=sha256:5cd504e127adcaace521d8c81ea00e49a3a20f2ebea27fcfd56754eb9f5d6b10

Observation 42eb22b1-5ca2-4d90-a77b-4ca2e0303008 · outbound

This paper cites TEN: Table Explicitization, Neurosymbolically.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition TEN: Table Explicitization, Neurosymbolically

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:01.912409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:01.912409Z digest=sha256:3e865f9d71e410bdd74b447a54b73d08f37310ff3c7a28c55e5c7e004d3dd8d3

Observation c11037a8-27f6-4085-b147-a88ac08cd889 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.022622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.022622Z digest=sha256:9a5ac3d7fe9678d8d5812095a912a02978fba762cd9c09430eb3e81160342a32

Observation 919ad2e9-237f-467d-b443-adfeafc70dc6 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.232688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.232688Z digest=sha256:93bcc604533647cf2d90c1da233561f39f88bc526dbba55577cc1d4659c1e162

Observation 925c1d8d-086f-46ad-9bee-321ec6fd14d2 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.285651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.285651Z digest=sha256:b963945a60d7e6186e1b6b9db56b0e34f30620c78dbbbc6de5423b2e63c58699

Observation 65e8981e-e48d-4047-9536-e65929466467 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.376067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.376067Z digest=sha256:103d48a4fde2325f100cd55b05e0a01840aad3ac01fb4a8e2e20e698d50619d8

Observation dadb9c46-bfb1-4024-b796-34f6b99a304c · outbound

This paper cites Kelly Buchanan, Mayee F.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Kelly Buchanan, Mayee F

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.445402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.445402Z digest=sha256:f2d9283ca56a782e10c52d1c3e53d361957c733d59fe2c5abb4d913866ce8aa7

Observation 4af724ca-23c2-4dc9-9237-42e65ab813a8 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.545477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.545477Z digest=sha256:fe4ac99872d67cbbd32893405c193dda8ca20eb4f620cb22ecd72025e1be3ac2

Observation 6b8cb060-e586-4a46-b9dc-4bec9fcb6873 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-02T05:33:34.906795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-02T05:30:02.632168Z digest=sha256:0de92f7c077b09bf3d97f8237fc9f1afbf4ae8d56124427ae160faa31269e1f0

Observation 935af909-d024-4f37-9004-dd04e791bcc5 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.804383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.804383Z digest=sha256:810548ed933b5615435fe6aed4e2444d0de0411c8ed88443ff4b1dbc1074ee6f

Observation 0506cba1-8b80-47f1-9224-9a7ea00e7d3c · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:02.880383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:02.880383Z digest=sha256:a2dd426620ba7961cb283f830da1b651f55fd1f06590721078254e305add5fd3

Observation 5dbd6d8c-715f-4d96-9be3-dec537a15f69 · outbound

This paper cites Aligning benchmark datasets for table structure recognition.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Aligning benchmark datasets for table structure recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.007689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.007689Z digest=sha256:d3a91d3cf07ea18a87975dea2ba9d9d2b5d4f1c910989fb7292a07094df45892

Observation 62eb7816-4eb1-4673-b774-54d36db3be00 · outbound

This paper cites GriTS: Grid table similarity metric for table structure recognition.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition GriTS: Grid table similarity metric for table structure recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.111252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.111252Z digest=sha256:a933e16c5ce564abdf5850fd71dd058071fba351540664e592e031195652dd79

Observation 8c3e718e-7925-41d8-af20-9d62bc090ce7 · outbound

This paper cites SynCode: LLM Generation with Grammar Augmentation.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition SynCode: LLM Generation with Grammar Augmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.266859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.266859Z digest=sha256:8a81b40f22b109be888f50aec6e9c25465d8b4d9249f396ede3325145c75a4c9

Observation 470e3904-5b98-497d-9b01-e2604a09c953 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.415974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.415974Z digest=sha256:f65ce16d378e099cef246e83d8d78f31d866e17f2b21501bc1360b94582c9358

Observation ac5739a1-78f4-4a1a-8fad-a6da8a0a24c9 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 34

Resolution
verified exact
doi, observed 2026-08-02T05:33:34.669074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-02T05:30:03.480135Z digest=sha256:143586d024a323f404b58a73d44d8dd3fb9c73be847a7076e3d9817cea348848

Observation 50da801b-6989-44fc-ada2-191f393a3399 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.626698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.626698Z digest=sha256:a65e339bec316e3d25791378b490a4cceefc1e300343c2a39ac918bc26092533

Observation 7b48686a-ac41-4dd9-bdd7-be10960e65fa · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.778082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.778082Z digest=sha256:3045051540e1c3e6acbc9abf407027ffbaa6750c1c03b4416dbbf7869e3359f3

Observation 48c4dd79-fbef-4919-818a-d6be750161bf · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 37

Resolution
verified exact
doi, observed 2026-08-02T05:33:34.420662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-02T05:30:03.853188Z digest=sha256:7f8e298dedc1c0188fa808315e7f161a029db3661d1bc938eee38b8abe4dc1a1

Observation 27366696-8c64-4c81-aa1d-18eb06cad908 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:03.952979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:03.952979Z digest=sha256:c608653b24e3ed48007be278f510aa7501f6d60ad5f233bc30d9412bceabab02

Observation c8230fa8-9dc8-46e1-ac56-1cd08f0239e0 · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.029139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.029139Z digest=sha256:e633db17aad39af7e6f057610a2b585b0ab2397d28b7c978e21bda2ea1cb64bd

Observation 696615e4-bca2-44fe-bc36-216bb0779fef · outbound

This paper cites Nanoscale Cathodoluminescence Spectroscopy Probing the Nitride Quantum Wells in an Electron Microcope.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Nanoscale Cathodoluminescence Spectroscopy Probing the Nitride Quantum Wells in an Electron Microcope

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.099899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.099899Z digest=sha256:4f9ba9400018449cda793a09323941b40a8a5eb02edf503eb41bc137b89550f3

Observation 5b10c82e-cd9d-4394-a16e-ee2ee5983872 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.190121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.190121Z digest=sha256:9c597fcecc1c294d2550c7f925879ad33f072f2d8f2ec94e45c41572b1af6271

Observation 9addce6a-7735-411c-8cd4-eda301a81d2b · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.258677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.258677Z digest=sha256:caebbe87dae7fe446c7fe2ce22f0d6491ac9848b2ced1f70b2a161753aedf3b8

Observation 4d5f947c-e91d-4acb-88b4-6eec6aa322ce · outbound

This paper cites an unresolved cited work.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.341034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.341034Z digest=sha256:b508fce6c59281199f157a6909b836d55585aa0e89022d452bba2e16005a0458

Observation ae1d7950-092f-426b-baee-9221d1bed747 · outbound

This paper cites Image-based table recognition: data, model, and evaluation.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Image-based table recognition: data, model, and evaluation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.435825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.435825Z digest=sha256:82e35f9d9ff49ef1fdd179b6d583b291a1a8b1f7318b784be34014c4f2124269

Observation c00c37c5-f72f-4540-aa54-4c93eb9684bd · outbound

This paper cites More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.547511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.547511Z digest=sha256:fb3b426bfe3e97bd9e7e99b6e6f28dbadcc923feb5919d86dfecebfe63a356eb

Observation 41700b40-250f-4351-a40b-518bf1cb5984 · outbound

This paper cites Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.657199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.657199Z digest=sha256:8e9ed0266c041a98b60ee9b47c1b1182443f810881f9e18493e3efa195217d52

Observation 7660ba61-ab08-49eb-80ce-c415e1ddf896 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:30:04.755723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T05:30:04.755723Z digest=sha256:6da61de9dfc98f6cba786aa3e84f27e3eb066ea8426101274c020176c36505f3

Pith citing papers

Observation 7be6c274-ce03-4e81-9ee7-f1e52c8bdeff · inbound

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles cites this paper.

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T09:29:49.561922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:29:49.561922Z digest=sha256:3dba3e5a8d385112c8dc243bc4f483d9ef559506a7716f42bd61ebf11959f5ac