Pith. sign in

Paper Citation Record · LEDGER

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.00777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00777 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:47.526084Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:20:33.617934Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc6415e-83f8-4705-a460-ea733bb3876e · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.279759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.279759Z digest=sha256:f22738d07224ec2f338807a3794b1d38bb375a8e3953a7685d85543c6254fbe8

Observation 7106c982-2c18-40da-b7fc-c50c17e8645b · outbound

This paper cites GPT-4 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.374841Z digest=sha256:e4da540f8814bf81f7421478b0db2a000151c6553a94b06d50abc7299b0771ef

Observation 77d10d3b-a33b-47cb-84f2-7a33f045acd9 · outbound

This paper cites PaLM 2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge PaLM 2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.455923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.455923Z digest=sha256:6cf085a71502aa41904215ace455fcd9bf16d08fadfe255a16d8fdc71289ac92

Observation c8b80b04-3649-4f7b-a8e1-866fb883057d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:52.076468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.506723Z digest=sha256:6deb236fdbf49b00e4a5c4143cd5f92819887f2cb2bc651b36164f89223cdcd6

Observation 2b483b88-b595-4566-a055-62d9d73d343a · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.570013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.570013Z digest=sha256:033ca853885caa7e7b5ccbe5c447eed194581ae78ccdccb89602d4a32c68a17d

Observation ae6a5b49-c271-4d04-bbc5-f107955debdb · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Do, Yan Xu, and Pascale Fung

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.667380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.667380Z digest=sha256:25cb902923f615c3a64b92c98d800b8aa8b1f5015238d5291dd3c0099d5a7cae

Observation 14591c23-b534-4fe1-a311-27e02139c2e3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.774395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.774395Z digest=sha256:12bb72d5b6d1c1179a8fbad4647f73392a9c7c9f1fcc82ee2cb6cfe4ae4f854f

Observation dbe47d2d-c4e9-4e09-94f4-5926e32a1963 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.811421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.811421Z digest=sha256:8d17922563c2a16518413482c08e25ef8f026e371a3c3ade90d987465e8ef8b1

Observation fdc7ff7e-8cd7-4932-bb5f-d44032f25b2e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.799121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.908302Z digest=sha256:31b277f654780c6d80e5b95c3225615eb43308b31eeabf0b73b9ded4bb0197cc

Observation e46f93ed-c5e8-4406-8463-fbfd8a3d8f3f · outbound

This paper cites Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.971142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.971142Z digest=sha256:a63ab11e020f83adbe8a39f2c0898ef86b5c8380b9441fd4454cfa1e76eaa59e

Observation 6ec81de2-7bac-4235-9bde-7f605eb4f692 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.515495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.045874Z digest=sha256:d86af22298ebdcd2b733a8c512aa21d629402b290e946c3cf3f2917e11cd9609

Observation c16909de-1bcd-45b2-a1ba-703ef69ca796 · outbound

This paper cites The Llama 3 Herd of Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.092977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.092977Z digest=sha256:c7802b724b6215f80f92beaba6fe8bc7026f3b9bfbd0a161267e250006d96e41

Observation 50cf506a-f7c1-4a37-a1bf-2367ccb2b5fb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey on LLM-as-a-Judge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.191039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.191039Z digest=sha256:cf7155f062608a73266c81796ce6c140b38e2fa970f4d1ca77ac1d57e5948bf0

Observation 5c4fc371-c7ee-459a-a98d-c2389ef9501d · outbound

This paper cites Textbooks Are All You Need.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Textbooks Are All You Need

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.278334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.278334Z digest=sha256:6a52507aaabfbcb67a6c5ec520e0198afba168690f5d7249c213e2e763bf3891

Observation 273682e8-0e02-4749-83a5-286b9146a2da · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.337098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.343234Z digest=sha256:5830661e2b9fad655afb18a3b18d566201a099dc05cb2690fe162802ffcbe5d4

Observation cbae3a98-ca16-4dd6-bbdd-44dca82ce701 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.158245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.450715Z digest=sha256:aef61a24933e1762c7fde83367b458231260a2e1a0ad1557dfcba63b5194ea51

Observation 318fbf0a-8b9e-46ba-b30e-1bc1ba27c0f1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.973844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.579781Z digest=sha256:7aed8850b40d947374d22c8fc6b26a84ba6ed1f2f423c11d5b5ad7439f50f957

Observation 0acd8fa1-ae60-42a3-a3ab-4e4fcb1837a8 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.819956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.624950Z digest=sha256:baf0b7c2c6147f0b6f252e1c667e0e643e9a746e3f44b6821b636f532913fe72

Observation 750697ef-7b14-416b-8c6e-a9e46d38f0a4 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.701200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.680467Z digest=sha256:382f7f30c6de7535438f60eebe82c177d5bf75b99ede0c8429c74048e80b155d

Observation a37e638d-10f7-4dbd-a7d6-7fc50ad630ae · outbound

This paper cites Mistral 7B.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.774585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.774585Z digest=sha256:0cd96c9eb265494db8d52d3e4b65864e80c9ed4cd1badbbfa0fd8ab2fd652c2a

Observation a5c456ea-1544-4d02-b8f0-e6aad9adf624 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.846828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.846828Z digest=sha256:708c761b5f908fa7eb2fcb75705183fcde0802eb8b889dddf6dbd458d3b7fde6

Observation 9435f877-6007-4322-a81e-c5810a4c9458 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.551364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.901982Z digest=sha256:a6da006e8ba289d8c17e122fee433ff5061b0e87b31838ce1cab92cb3a01e89e

Observation 81544950-e252-4a74-9359-af0383cc8fb3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.386188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.985167Z digest=sha256:78cd26ecaef3f0eb69460e7c588d6154283d73c4210e25e2aaea84515bcd427b

Observation 26eb1209-bb7b-4000-a877-b69af6dcc843 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.234920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.060211Z digest=sha256:feb3bb68e5cfcbc044db149e4d2fc1b92d3a929ec04b0a1bfd76f5c2dd093389

Observation 2560fe2a-3da9-4493-8ab7-659330931aed · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.062140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.123893Z digest=sha256:b55d6146702546cf9a10ccfe5e3b1c53b2b74cde9096e4e65fd3fbb4a20c5979

Observation 298d0589-20bc-4e25-85c1-96296dd9b862 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.164069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.164069Z digest=sha256:e69b444ae3e2085df29f8d905a4a41fb3670f22e6159d9d6dc63a475b55b4f41

Observation 06bf94d4-9c2f-48d7-9bff-76396c9df471 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.844122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.199548Z digest=sha256:943f2af2cf5d34e298fcd5889c763841067a9295b9ff68f87f27f5fa28a32efa

Observation 4a2ef83a-53bf-4b90-80a8-e7d056c11299 · outbound

This paper cites Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:48.093541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.290988Z digest=sha256:ce38f7c7b14fdc4cf65de39ce4e2f3dd4956eca8cad6cb80ed32d9530f0f8df9

Observation 5cadf5ed-49d7-46e9-bd80-4030446042c0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.699229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.386523Z digest=sha256:b37d1b42b0405d5be0ba46957e309cc3c085e6c0459eaa7ddc852f1ac969e4ff

Observation aff110c5-0e2c-4e02-a339-daba80ea5fbf · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.447467Z digest=sha256:377e89a881a9b804ba495d823995096d2e5e057ff8b299a3307dc7020aa6b331

Observation 83333015-ee06-4a8d-b60c-dfcf78e74003 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.488172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.522250Z digest=sha256:a1232c465c44ebcfa8e1538c5c4253884d1049167642cd993055ccda678e5ab7

Observation 42e80012-10ba-4cc7-8113-db1af6a125ff · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.662507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.662507Z digest=sha256:1e1dbec8cbfd2a5033046b37eb4e8b5735a3346f7b9935825ac322329939b17c

Observation 2095bc52-8ce1-4095-8456-d2006553892e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.258384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.735422Z digest=sha256:528adecfa373cae7e22d3c5962a8a66927602a797b00149d54324ff299007377

Observation 92117b81-12ff-469f-8f40-7013e7db6479 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.070708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.811541Z digest=sha256:46e6558ed8f931a9af21c453700910351a370591985311e2d8eac6246ba3c43b

Observation ce3edb10-ee22-421e-b374-a70e25b9e77d · outbound

This paper cites Large Language Models: A Survey.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Large Language Models: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.923985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.923985Z digest=sha256:383d66eed7f2ddb97ecf57a69180b6098ed00788131fca381a0640d006449b85

Observation 97fa6cba-6fe9-4112-b633-bd51a6587bc9 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.028316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.028316Z digest=sha256:be9c8681b29abde7d6c7401ca1946cbc85ab5dd90c097936ab07912dcb0da851

Observation 833e2177-030b-4af2-bc95-7632073d094d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.903055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.066169Z digest=sha256:132c3a386056496bd817d14cc542a7533c561ece235c8d2c4aa2fff575bfe3b0

Observation 26328e55-814b-46e0-b17c-7f2f17807602 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Capabilities of Gemini Models in Medicine

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.105692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.105692Z digest=sha256:5dde64b6605529f5730d21851b61fcc0bee18d9f1a9cf50ebd050f893d5629d7

Observation e1eb2724-5824-4238-ae81-2151f47b44b0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.696990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.175013Z digest=sha256:b433bfe06651db0db289e31825b919c025e0b01ec8582cbdfac6a4269f16dc3b

Observation a1c00059-8291-4ba1-8775-412fa7dbda72 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.272794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.272794Z digest=sha256:e517d01f263c8af08fa5e1f4726b22fc1c81e37c1b00944bd230592152b9e3a5

Observation ae4979be-f501-4ceb-9281-b6eb49557a47 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.559413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.351173Z digest=sha256:ca5ccd01470c3707284b391aaef2f7150c51772b06cb87e9a1d447e9890f6ff2

Observation 314fb2e5-19f0-4af3-9c87-50d91fa7f7fd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.447687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.447687Z digest=sha256:44154845fe7ebd47c26fc6c1f4ca53c7cb4bbf22c5e4709ae24a86ddcecc8d92

Observation e742cba8-03c6-46fa-b04f-7a509c554f3c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.559562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.559562Z digest=sha256:a4d0aa37ae8c9dd750d2ba58b29dd9489ea8080efe7a372781c1030dfb2aeeed

Observation 68bca6ec-97b6-4e23-84de-f71a635099d1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.635872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.635872Z digest=sha256:b3da3bbc20a1fa302116d332602294e14ba89a5c91f921aa80d87bf4fe44cb6f

Observation a894cab8-c0f5-4d16-a157-c7161b72cab6 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.689377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.689377Z digest=sha256:94a55ca6065a98ce57e3b7f6175608511fe65b5f09fe8f3f7bacfd86436abb8b

Observation 4663eb2b-2686-445f-a5d7-a8a7b0df7032 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T12:02:47.762727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.767706Z digest=sha256:9344c8025127779d29ca8879c363779aaf1150d4a8b838a17485e5c6151975b2

Observation 30c32547-5fa8-4a28-bec9-fa48dd432be1 · outbound

This paper cites Qwen2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.842724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.842724Z digest=sha256:3e2d6b9c4e5c5863e80e7a196b8be82da7726960872da5c88684835209129e5c

Observation 95adc132-2783-4092-a826-74124f853fc3 · outbound

This paper cites Qwen2.5 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.912360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.912360Z digest=sha256:bd2e77f4c28b4d9b0ab76d26de9838bd0721cf63647c7567c9524d3611edd97a

Observation 58ebfe33-87a9-4cf2-acb7-27659c5495b8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.010722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.010722Z digest=sha256:889c1ee52e533c0d641f2ef6755821ddf2526fc254aa7baa5b9061565cadd862

Observation 04287b94-8cbe-4d13-a478-4cbb1f7b122d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.084587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.084587Z digest=sha256:f5af7b44abd2398b7d0d76c9468d058533d963c99ddf3cd12361d10e1a5e2790

Observation 82a3b566-caff-4747-837c-9698baec91c2 · outbound

This paper cites A Survey of Large Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.192473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.192473Z digest=sha256:b126a3401aa5fe19db82a50e79ba0a919191633ec907b1e0adef48ea92a2b6ca

Observation b3160186-47ab-43f5-aa75-3981d65cb121 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.241849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.241849Z digest=sha256:95104450eb254fe29b2c049044c14c0b016a289767db9ec91a68b66d3bb0e380

Observation 29b03e43-fffa-4182-a8a2-54e8643a8ea0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.359311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T12:02:47.356554Z digest=sha256:316f4681606efe7309780503d2cbc73e199f13e6e3d6c584c2e181053d5b448e

Observation fa99093f-eff0-48e5-8bd8-ad20e2821bd3 · outbound

This paper cites online" 'onlinestring :=.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge online" 'onlinestring :=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.459081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.459081Z digest=sha256:0e317004038567c310e7e9a4bc9d3f0563a5b42ae319726f60a21c670a1f83f7

Observation 65b484b4-22af-470f-9f3e-ec5db528a1a3 · outbound

This paper cites write newline.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.526084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.526084Z digest=sha256:6204edd5cb7316191ee409be48ca8e675938263ea64d5b62710c5c79cfa26f60

Pith citing papers

Observation b2697847-5dea-482a-82da-5a066fb655a6 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:3d68ae22dc3e09f5eaada312b884c80dada66eed9bf5395643e72f75946de3c0

Observation 3d01d9e4-23ce-496d-8e1e-f52f7e92bc81 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:33.617934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:33.617934Z digest=sha256:e91792461f8ff8b4b33b7cff5d5a28e73fddaf8f3d69e4cfc1bbd76274aef789