Pith. sign in

Paper Citation Record · LEDGER

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2505.11887.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11887 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:52:16.027375Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe99e6b2-9ad4-4833-9ce6-5a156b5afac0 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.660025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.826594Z digest=sha256:ad9a04380b7638901d9451cbe83573d49601988603468351e84c6f80dd46106a

Observation 736cacf7-1b7a-4eea-961a-98d75003af34 · outbound

This paper cites A Survey on Evaluation of Large Language Models.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation A Survey on Evaluation of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.832274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.832274Z digest=sha256:b1f768eebf984c80842aad94e3991b31b7fc562f5be183b16b694b1071a02b88

Observation 3ae1315f-b2e0-4ea2-be39-cdbecc09e727 · outbound

This paper cites Introspective Tips: Large Language Model for In-Context Decision Making.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Introspective Tips: Large Language Model for In-Context Decision Making

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.839099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.839099Z digest=sha256:7c0f922a2c732276a5b097d1857e23415cf6026b6f35caa7605baf1645e880ec

Observation d3a4a363-d82e-462d-9b14-f15d2b378a85 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.844477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.844477Z digest=sha256:28528095ec261c9e8bc369f93861a17bf861b74137a6a7f532de09ded72a98a8

Observation 063b6d11-0eb9-4b03-9afa-a12882c2c04f · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.849900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.849900Z digest=sha256:2095a6fe2fe2f9473f39968df1bd82d901fddbbfbd1c0ac194dcc8dde5392660

Observation 8d7314dc-5800-4c7e-9114-d2d66514e84f · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Fu, Stefano Ermon, Atri Rudra, and Christopher R \'e

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.855086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.855086Z digest=sha256:0751c974f4ebb063f800536ed681d8abf607c05bc176646030d787294cef0285

Observation 36d53d3c-ce5d-4da7-9ce9-5c06e55ddab4 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.860881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.860881Z digest=sha256:f04803209e3a9f631d174802055297f7ad02b83ac2691fa680c7d9c804756d19

Observation 5ae0853d-ebca-4618-beea-738a08a39a98 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.865639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.865639Z digest=sha256:305ea4cf01aae9d1b8cea19347ab3d739aff678be7eabac2732f2de7e2ef520b

Observation f0ff8caf-385a-48e6-b8e6-99532332d0a0 · outbound

This paper cites Toward a Formal Model of Cognitive Synergy.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Toward a Formal Model of Cognitive Synergy

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:52:16.366965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.871422Z digest=sha256:f0056314f0c94d7d2d7c8fe020bd195c41089cdbde08895f595b683c6ca2d951

Observation bfa08233-90d0-4f6c-b2fe-e0a910ef035a · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.876404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.876404Z digest=sha256:869f8292c48824143576d39537b2df0aaffafa61751c8e5cb8451eb7b45a0695

Observation 89f45ce1-1142-4618-b202-8b989b4a4413 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.881968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.881968Z digest=sha256:e6a7fd8d1d1e79f339dfc072e20001767d5f7262ab357900ae461275b3a42f53

Observation fcc36fad-f0dd-4f5d-92a2-66b04e719431 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.599814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.886953Z digest=sha256:d722356d7ce8f03a08f93e0fa9355db0cedece5d4409a853881a52f788c4fb72

Observation 75244dc7-19e0-4da3-833e-598a2453cda8 · outbound

This paper cites Generative Judge for Evaluating Alignment.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Generative Judge for Evaluating Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.891381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.891381Z digest=sha256:6c9d4cb8fd9d2c999e7bccf91a6ba3ed27bdb42c090e432ac61d5dbcc2c35411

Observation d0546bd6-e589-410c-8fa1-fd9973c40731 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.895949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.895949Z digest=sha256:6dbeb76688e48376faf2e606f50bfb41ae7b5f6ee6b7b7de819b318617a89ec7

Observation 60031fa8-d1b2-4d9e-ac79-746eb84c8c87 · outbound

This paper cites Leveraging Large Language Models for NLG Evaluation: Advances and Challenges.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Leveraging Large Language Models for NLG Evaluation: Advances and Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.900326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.900326Z digest=sha256:4ec7c67ca293fcbdd886d4995fdf382b955e79db066a3f07a1d444682279e11a

Observation 4ce9db86-1999-470d-8155-96053a3c25ec · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.904615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.904615Z digest=sha256:90170b55d58ee9d9b440758a745bf72eff828e095fe96399c6ea55bb2f3e4acf

Observation 7cec2532-cd5a-4414-8135-6becd8038079 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.908759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.908759Z digest=sha256:2649ce59727b49a36ee6d26a85e969aa75b74d1e76d467b13b99fc12e12010e6

Observation 85bb0d07-65c8-4179-8b85-d7dbabe780bd · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.561310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.913094Z digest=sha256:ac740251066aa2f9cce5a49d0ef9543bd17360452c2d316e09c881c161248bc7

Observation 99621ce5-1fcb-42a1-be4e-5ba4f055cfb8 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.544356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.917947Z digest=sha256:7e459dae4cd985ef16dfb0f48427e633f19f009a943a4050e43daeecd0a65415

Observation 27a5002f-8dd5-471b-99bd-417cee32d824 · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Capabilities of GPT-4 on Medical Challenge Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.923080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.923080Z digest=sha256:58419b0a08a448e153d7fdd32d22db72e0aca9ed9b21948b2ec46a147fd63ca5

Observation 1bf04891-3e13-4240-be05-c998aae1a61a · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.928985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.928985Z digest=sha256:45574b682d9718d2accd08ac3a76a17a13dba39bfb1124abb248bd9b43d5bcd5

Observation ec796b18-65fd-4cc5-a77a-8f7563445835 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.515007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.934256Z digest=sha256:f7e8b9a7ff168105a43222370ed343edd2e9e252e30991226c0c4a7a948d0139

Observation fd885281-6373-4f5d-9fe6-2571e34e8ebb · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.939911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.939911Z digest=sha256:4dbe246160c5304a4fa57097f2f74fbc1b8b22cea1b5f06f4a8fb310d251f9b0

Observation 26999ed6-402f-43ee-8b98-d0ecba9448f0 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.945505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.945505Z digest=sha256:731735754d73cf3d339ee4414a3cb7a510c81c0936b7a63f6e48131d3146b2ad

Observation ec019236-0aeb-45d0-9fbb-c2e864adf274 · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:52:16.485216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:52:15.950355Z digest=sha256:9b5f0d78c2534fd08535234d9e8352f5971d86e7738cc969f2a972dd7074ed5f

Observation d52e11df-71b6-4140-b3e9-f58d670e1e82 · outbound

This paper cites Towards Expert-Level Medical Question Answering with Large Language Models.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Towards Expert-Level Medical Question Answering with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.955259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.955259Z digest=sha256:3b75db06a654c492b08b142821b54684a5821a1042cee6affd693da990c25d50

Observation c502c932-4b3f-4e1b-8c9b-c0d10e69fae9 · outbound

This paper cites Towards Generalist Biomedical AI.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Towards Generalist Biomedical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.960802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.960802Z digest=sha256:0d740ed11bf68bf9b82d29567e15e6c5e6d9e4fabbf5b89489d3c8305c4cc9f9

Observation 62e19b43-25cc-4929-bff4-926201a354b8 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.966499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.966499Z digest=sha256:445f589c3d3c5036d02b7474d7bc82e4a24e32b4ddbc53a0d55bc4c731c6d62b

Observation 31e83056-c601-4f4b-9e4c-0b4f6bc2ea90 · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.971613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.971613Z digest=sha256:59d79cb9e15b38c84cd3182d50304cda97bad02d8d76d4aa197a741f4545e426

Observation 0ed5c82e-5a2f-4981-85b3-0e38ec60b5f2 · outbound

This paper cites Metacognitive Prompting Improves Understanding in Large Language Models.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Metacognitive Prompting Improves Understanding in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.976781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.976781Z digest=sha256:0fe1c7612f2d8c0ccb8950f548c566ae8fec02eda358359a65f351537e2bec71

Observation d6a6421d-a7c4-4106-9d8f-69d28feb8bf5 · outbound

This paper cites PMC-LLaMA: Towards Building Open-source Language Models for Medicine.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation PMC-LLaMA: Towards Building Open-source Language Models for Medicine

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.981975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.981975Z digest=sha256:941359b9d427651a8e7127d67c0b436d0eb68d0a01455b78dfab2fb10cdf2ced

Observation 44085475-2b97-4358-981d-f60b41e82d89 · outbound

This paper cites DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.986548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.986548Z digest=sha256:58d48a40117bfe52fd0e64c8dce0354c18c5984444bc70f816c4a0eb7a84092a

Observation 6f264726-f235-421e-a510-d611fd71af91 · outbound

This paper cites Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.993415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.993415Z digest=sha256:8ac3553eab6b615604024fa6f9573fd5cdcbb64b3dd0faed98a6ef51e9d82bb9

Observation 7bcb0c7c-1f6a-42b9-8737-f99deb595d32 · outbound

This paper cites MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation MedGPTEval: A Dataset and Benchmark to Evaluate Responses of Large Language Models in Medicine

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:15.999735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:15.999735Z digest=sha256:04993bb0dfa5887095216aba695afd2f4c99d0f69bb8bffa408a0a8719a21081

Observation 9a442ca2-7af7-4e24-b452-500945c5925b · outbound

This paper cites an unresolved cited work.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:16.006632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:16.006632Z digest=sha256:04a3dca88da6ba42ff1d68e2ea44f32a7e2f3f34fbf5a382bb2bcd0fbc14cdb6

Observation f3741cb3-ea52-4b82-ac10-7fd2b015fbae · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation BERTScore: Evaluating Text Generation with BERT

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:16.011526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:16.011526Z digest=sha256:104c9fde3b1e866e7ba4fcc452fb0d65afac99fe6661cc6a633d0912e3544ec8

Observation a460ea7c-0bef-4bbc-860a-30cb27ad795b · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:16.016392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:16.016392Z digest=sha256:27dde25473f4eb84d25b5e7787bb09dc97f0dcdaa4bc04db5418a2deee09f323

Observation f9f921f0-261d-437f-a04e-24c482c895c0 · outbound

This paper cites online" 'onlinestring :=.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:16.021831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:16.021831Z digest=sha256:32dba85eeeb8a3bff8e30c8dbeb82e2cc338c6deefa449a88a399dbc7bfa9113

Observation 1da77e8c-2915-4da3-9140-e64a5116ad75 · outbound

This paper cites write newline.

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:52:16.027375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:52:16.027375Z digest=sha256:7db0519dd5334547128be4dce831575ab3fb7910150e89e7e7933dcfba4511da

Pith citing papers

No inbound Pith citation observations are available.