Pith. sign in

Paper Citation Record · LEDGER

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

As of 13 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2607.09322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09322 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T03:54:11.805387Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:05.046139Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:18:05.490656Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2dc38c9-bdc1-4e7b-8961-534e6ede89a9 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:052f3e3c1303386aa827be55e5370ba87a929f67b44e8d5921e498b5943d0c18

Observation 28d8a0e3-5c05-4c54-9ba4-d1888d6d36a7 · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:be171b9b838cfc0bb734f6bb2e862165710bf6fdc21c6e98649930f87e776bf7

Observation 9b1a0c8c-5a39-4331-aa4d-3e96d3ba3be8 · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:44c5b5cadad037980c685b2ee6dedcec935db93774dcd07bf15c197a3d042faa

Observation 71f76bde-a72e-49e4-8407-3fab2c8dcce8 · outbound

This paper cites Journal of medical Internet research 27, e84120 (2025).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Journal of medical Internet research 27, e84120 (2025)

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:bfaac31ebea3fb140522e2d0523158c85e64a34887edeae46cd2d3a52d0e90e4

Observation 0bf90271-06f2-4de6-b820-f1e188789564 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:9d74b278066c267b9d7bfc0c7245bba34bb07940025c43366058189a3c9f54c5

Observation f0f38a46-759d-4f96-8a0b-d64f15ac5029 · outbound

This paper cites NEJM AI p.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making NEJM AI p

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:ff0c62d64d53b9e12af9eb7e6b7b64b50e9ead1245afe36bd3d79d0a8ebed9dd

Observation 82dda204-1e89-4c0c-b6a7-cf6f494e0323 · outbound

This paper cites PhysioNet (Oct 2024).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making PhysioNet (Oct 2024)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:f5a57b511071be83dce68f84faa9e889d9d1f40174b94c3b467d1cbec3602fb2

Observation 9b5a0c37-3660-428a-87d4-636e8b001b28 · outbound

This paper cites The NarrativeQA Reading Comprehension Challenge.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making The NarrativeQA Reading Comprehension Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:4386f664a0a023885b9bec47f14b66cd9b982cae328039c9651b3b16b715d63a

Observation 7d106ae4-3c01-43c2-b565-11a61e6529d8 · outbound

This paper cites Advances in Neural Information Processing Systems 35, 15589–15601 (2022).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Advances in Neural Information Processing Systems 35, 15589–15601 (2022)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:fd5820b79fe18d904469176a3e9fa17a9d480c9db3ee8e9d8fd8d7f8e244769d

Observation ccbd14d6-5ced-4988-890d-f5e5a86af649 · outbound

This paper cites ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:e5432fb7ae76dd5634345b0f411cc934a55571f87d28a70115e696b17e92c4f9

Observation 12c92b9d-138e-4f9e-989f-b1932c4b5d59 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Lost in the Middle: How Language Models Use Long Contexts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:2c0d32b47e4a9bcc605c4fea9724c854f8928467b80dd1ea8a67091f75529d87

Observation ab4aa5d5-7d36-45c2-98df-64383f0f16bf · outbound

This paper cites Evaluating Very Long-Term Conversational Memory of LLM Agents.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Evaluating Very Long-Term Conversational Memory of LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:02a7b2d571b039856866d046b1bdc1d60545efd6fd6fc6ef097732af9d8e61bf

Observation faaa1bf4-56e8-4a62-a837-9eafcb49ecec · outbound

This paper cites https://developers.openai.com/api/docs/models/gpt- 5-mini (2026), accessed 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://developers.openai.com/api/docs/models/gpt- 5-mini (2026), accessed 22 Feb 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:78a83c1a166d9fbf856b018a9517e758bfaff9b6205bd3278c1b1579a323f609

Observation 70e9150b-08d4-491c-b4dd-941c922a4ddb · outbound

This paper cites https://developers.openai.com/api/docs/models/text- embedding-3-small (2026), accessed 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://developers.openai.com/api/docs/models/text- embedding-3-small (2026), accessed 22 Feb 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:3a3e3b864cdd5a81d4198c26b59b14da58714e4a2b5b7d841222b4336aa4db4c

Observation d8dcca31-e913-427f-ab89-399fad0e1c7f · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:4d52af25e38d2bfa856bd491dbf8a1edcf4d4c4bd78f49ae5ac1671038170547

Observation 936a5c28-b295-4068-b876-b2adf1a71291 · outbound

This paper cites https://qwen.ai/blog?id=qwen3 (2024), accessed: 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://qwen.ai/blog?id=qwen3 (2024), accessed: 22 Feb 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:4eae4b9cc050099da2ebf4527cc4c3e244cd0950c029ce42d8c3f30ee8f09927

Observation dc345ca6-188b-488c-81de-58a74508ceb7 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:84518c6647cc05293237a3c033e05dfb90aae28dff1f055184eaf95a654378fd

Observation 570233fe-6d1d-48d5-94fb-843e3950cecf · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:4f765ffeedfbb8de0e3437ae411c827e58c6a824536914cc876b5dba99cf4ba7

Observation feaafc4f-d5aa-46e7-a84b-b8207ae98d3f · outbound

This paper cites LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:69fa4fc53e3df19ffd84a23f364ac390b14c41862a66816e29526edbac986f34

Observation 0777cc87-ae87-4d29-bf74-8ca3aa481b63 · outbound

This paper cites In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:40b030e50e95c4204ec59e3ce9e2378dbdbcaf6ab4cfc58d19b99bc095ca5f1a

Pith citing papers

Observation 6b953d02-81ef-47eb-b079-2255cc91aa7c · inbound

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR cites this paper.

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:18:05.497141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:18:05.046139Z digest=sha256:03a1ec4e11b2281a40968b12663d2b6bcade1a80c57990db13952de5a3d75ffc