Pith. sign in

Paper Citation Record · LEDGER

Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 51 inbound Pith citation observations for arXiv:2311.16452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.16452 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 51 of 51 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:19:33.202015Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

169
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fd1fe60a-68e2-4196-a587-1c877d8a65b3 · inbound

Can an LLM Learn Preferences from Choice Data? cites this paper.

Can an LLM Learn Preferences from Choice Data? Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:56:01.019110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T04:54:12.367946Z digest=sha256:d16ab2ec47e6d90eaae456e65725dbfb2ba5599eed42d710479cd8e4f71acadc

Observation 40cf7b2b-c48c-4b16-bb3f-232de0c584ce · inbound

Capabilities of Gemini Models in Medicine cites this paper.

Capabilities of Gemini Models in Medicine Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 179

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:13:22.971803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T17:13:22.759147Z digest=sha256:7e087de1d255d82b3e63dd3276a6d61f2ba8971bd6178c2342c4afd8469ced8d

Observation a519c327-9818-40e0-b5f5-d09500c33bca · inbound

AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments cites this paper.

AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:53:19.586587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:53:19.535482Z digest=sha256:83ee8eec5b83d334d9fafd89f4407dd6955acda25bc5b1bf7624b13a2354ceef

Observation 009b931c-2c08-40bb-a714-12ac3ad3bc94 · inbound

TextGrad: Automatic "Differentiation" via Text cites this paper.

TextGrad: Automatic "Differentiation" via Text Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T11:27:58.162073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T11:27:58.098484Z digest=sha256:d9327bf904e606d22db115947d3f5e5fb8b4c9b7f1fac2d9d3a4e8b62e2b340d

Observation 4b101104-0530-4786-99c8-f4f0c95200ff · inbound

GPT-4o System Card cites this paper.

GPT-4o System Card Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:43:19.077599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T18:40:14.847351Z digest=sha256:a55e12455d7a56f540945b91eb4571616cfd73925715548a86a223b379ce3a19

Observation 2c9d2596-c233-4a81-91f0-a9f73bb0d42f · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:36:50.275415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:0c5393f985b9ac747cf62ab6a00819fe5f67c14bc687d1380dcc9f8631461e55

Observation 590815fc-94c2-46c3-b16f-481ce68d607f · inbound

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning cites this paper.

Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T03:37:27.794889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T03:35:56.594095Z digest=sha256:9ecedd84b1ddf93f58fd814a40b7275a0ac7cb026e7f1a913b4fbcee48a4412a

Observation f8520b19-0c50-45f4-94db-ffbdc7d9dc38 · inbound

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems cites this paper.

Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 183

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:11.068056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:39:49.832151Z digest=sha256:c912c4217e91e96f9abb9437a2f1eaee4a894efe79eccb4c21c6422e067326ac

Observation 3cd2eaad-b15f-42f2-bf48-c9c9468ef675 · inbound

Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving cites this paper.

Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:33.202015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:33.202015Z digest=sha256:bf151e9a9c19b64a84405598322dcda1a3fede3c1e2d399435222620d7d3f8f1

Observation c6a7b146-9077-46f3-ac73-8b87b1a0b477 · inbound

MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning cites this paper.

MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:15:46.275717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:15:46.275717Z digest=sha256:0f8b941383ed79926d8b5a5e6ff9c18bbd93f7b248bcc6e0ca9d65fd021dceaf

Observation d5bf19ba-eb2d-412d-a731-3b80054b4a6a · inbound

Towards Effective Complementary Security Analysis using Large Language Models cites this paper.

Towards Effective Complementary Security Analysis using Large Language Models Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:46.615026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:46.615026Z digest=sha256:e4b177b9e0f36384e46dff678673aa954051aa10b39c458f2273ff9415d0f32d

Observation e6cd44f3-eb7c-4740-8564-9cdf7933fda2 · inbound

Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems cites this paper.

Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:59.632348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:59.632348Z digest=sha256:7e7461b9f8cc720815847db41eb7ee9176dd6d01f630a9cc4a89fd70ac560dbf

Observation 11300047-0100-41d6-8187-d94d83fae0bd · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:36.328613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:36.328613Z digest=sha256:20e64de27d635d70264f875e6784f611834ea3d2948bfc076b24b08b26fe5124

Observation da4f5f8d-d53a-454c-811c-5521ea9f913b · inbound

ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making cites this paper.

ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:17.584586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:17.584586Z digest=sha256:0033c012fd5f279fbf2298f44c2c1c6f7d8540c37eafa177e5017eb10d80aead

Observation 9850c560-7d20-4d05-8d5e-bf9bf6e0f933 · inbound

HIVMedQA: Benchmarking large language models for HIV medical decision support cites this paper.

HIVMedQA: Benchmarking large language models for HIV medical decision support Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:18.370147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:18.370147Z digest=sha256:96e831ef977df3655386b2564f4889ab3bc969f5ae5d0c86ca4047a518bb2476

Observation ec171c18-daf7-4b0b-8ba4-81621992f1ce · inbound

Making Prompts First-Class Citizens for Adaptive LLM Pipelines cites this paper.

Making Prompts First-Class Citizens for Adaptive LLM Pipelines Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:01:56.445463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T00:58:13.697256Z digest=sha256:1ff60189af6c7a3da9741ca4ad5c12a7a9febc4b3964d6b712446f92639bbb33

Observation 32e84c90-d21c-415e-9b3a-d8d4390ea1a3 · inbound

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks cites this paper.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.982290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.982290Z digest=sha256:b4ee16ee166e0402974503224e83a58be404fc559a1c4e88305803ee98c9bc7f

Observation 797b5eb6-1460-48b0-b778-10461e3117c3 · inbound

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations cites this paper.

Testing for LLM response differences: the case of a composite null consisting of semantically irrelevant query perturbations Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:38.354877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:26:38.354877Z digest=sha256:a1e895dd322d2ede1ad357139e91d8a4ad241fdd20072ec6a8e505e81a340921

Observation 59938cd2-eb3b-481c-9f91-ce3c1fd65b46 · inbound

Teaching large language models to reason like expert diagnosticians cites this paper.

Teaching large language models to reason like expert diagnosticians Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T16:43:19.371532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:43:19.371532Z digest=sha256:cc6cdbc6549678ce23901f0cc2cb0acd64aba0318ed987e9dc1f5b9e2768c207

Observation a36f5af5-f04a-407c-a62f-048d90e1344c · inbound

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks cites this paper.

Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:16:24.110939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T13:16:19.744864Z digest=sha256:bfffd3832d31c0c6e31cd85e90098ca249a372cd0a30317fbf8fcfb41af5c64f

Observation acac3c1a-a841-4656-8c0a-1cf74bfbb339 · inbound

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation cites this paper.

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:51:49.123850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:51:49.123850Z digest=sha256:4f71ca4f060c32a105aaa6ef77cb31712d095fe4ceb7722efae4c6305f1dd3b1

Observation ca6da547-36a8-400f-8806-751ac54a6d42 · inbound

Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox cites this paper.

Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:01:05.823992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:01:05.823992Z digest=sha256:e0a5e3b7798b3659666d5dcc41e498df2cfb836b60849e4c5f4b2c73a254c06a

Observation 29678fc4-f2d5-4061-beee-70719ec01730 · inbound

Group Selection as a Safeguard Against AI Substitution cites this paper.

Group Selection as a Safeguard Against AI Substitution Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 3440

Resolution
unresolved
no resolver link, observed 2026-08-04T06:15:35.881009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:15:35.881009Z digest=sha256:387d55bae7757e041a94dac543d53abd01e93c80bd577367703fd0dfcaa33265

Observation 3dd5afd2-3ac7-426d-96df-443e2c8b4ed8 · inbound

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution cites this paper.

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:44:09.290435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T11:41:26.274423Z digest=sha256:6edb0c4827fe13b5485338a5451adbfc7cfc10c989f441215507950567c32c88

Observation cee22403-0557-4751-bce5-60d25e451921 · inbound

Medical Reasoning with Large Language Models: A Survey and MR-Bench cites this paper.

Medical Reasoning with Large Language Models: A Survey and MR-Bench Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:25:26.711903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T10:21:39.892271Z digest=sha256:5f68a1a5e48c6fb550eec2c563c8144018eb2a456fca9ce9dc2c7783ffacdbba

Observation cf618542-ba1f-4615-af0c-2cef29eebcfd · inbound

Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework cites this paper.

Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:00.355002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:10:26.427585Z digest=sha256:8e04c5438d49d4c0ba289c975858ed88aae5512b27dab0813fd5d50f4942a25d

Observation 170f340b-0b20-49c0-867d-b02514f1934b · inbound

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events cites this paper.

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:20:10.025162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:19:15.920545Z digest=sha256:341efd1b593f318b0f0c9c0b68e16d55e4aa3903d2dbbedac97afdb0882a1e0e

Observation 3dfccda0-0ba3-4105-aaee-76073463d261 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:46:53.492347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T16:17:39.923337Z digest=sha256:4eca6ff50b65f579ca9d9369846500d17db01480b07c1481125da35f346b7ff6

Observation c4e771f2-2391-4501-9761-faeee07a3991 · inbound

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment cites this paper.

SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:42.849502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T02:21:16.825681Z digest=sha256:17e2b6ef63030108d8c28e5b474d819708fd6be06742d8d758f30a9eabfc900d

Observation 956e374c-8716-464b-93d1-484ce7d98ddd · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:10.702844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:43c0654904659ecccb77c090f52ab5897bb715cd759777d42d2d3f7eed204611

Observation b770c0a4-751a-40e3-ad07-ae5342690514 · inbound

AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems cites this paper.

AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:46:26.957906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:57:38.617652Z digest=sha256:c17496a80e96023a0e20c005a0cfddd706ff3eab4bdde85a40efa008a9af0fb4

Observation ec767918-802c-497d-a451-f65b1ab606e5 · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:11:21.948174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:28bfab0506041dcb8ac517700b9aa6a6ca2b9b23a4d28637d68284ae5394516c

Observation ed19ff2b-9d15-4ee3-8e36-240d3f797754 · inbound

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks cites this paper.

AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.879478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T05:10:02.941396Z digest=sha256:8a2d761f02b349e298e39fb5335c8cba8667f3d178b45b965078067683d45b5a

Observation d9ca0eea-1b14-4e43-95a8-113184e51f75 · inbound

PrivScope: Task-scoped Disclosure Control for Hybrid Agentic Systems cites this paper.

PrivScope: Task-scoped Disclosure Control for Hybrid Agentic Systems Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:18:37.473848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T16:17:37.824542Z digest=sha256:75f7ff393140a913919d8670704dfa4d0dd7613fcc0a9c4fe2335a14cdc225a9

Observation e7d11e7c-7a1a-4370-b91e-59d5084f1092 · inbound

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows cites this paper.

AgentCo-op: Retrieval-Based Synthesis of Interoperable Multi-Agent Workflows Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.418476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:05:28.416552Z digest=sha256:4ea6fd58d6220194dac048cd7c99ebf135ab13fe237bc0efd4621b7f99d980aa

Observation 3677f857-2551-467d-8867-a9b0271e0562 · inbound

NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding cites this paper.

NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.811730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:54:55.254082Z digest=sha256:fd513fb05f674bd41a84f674befba4ec49d797c75b415401a9f5a95acc89fb7e

Observation d8b688e7-c21b-4f51-bcb4-ea292fe8dc8f · inbound

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation cites this paper.

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:33:54.378091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T19:23:57.130409Z digest=sha256:64a8d1748856a7fe0c7347bb4e1ea4ba5747ec2edfb904f116d9790ee2d24699

Observation 2ade2156-6d2f-4fea-89c3-df43243b4886 · inbound

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation cites this paper.

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T13:12:52.371448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:12:52.371448Z digest=sha256:933e9d1c8776fadfdfe6e03f48dc3e7d7e73fb6ed5b5d6ca7db055ccd9a00fcd

Observation e7045d5b-d427-4604-8f75-4b2dd5c3f277 · inbound

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow cites this paper.

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.680465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:15:14.484574Z digest=sha256:95b7d1a18f71495756ba38c8a6e4a00384c2ec85106973f96e13a0c2dc85ccef

Observation 29bac44a-75a6-472b-b992-b5726d270afa · inbound

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning cites this paper.

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:36:09.250257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:12:28.812691Z digest=sha256:4f77cf62abc60903a4916e9a4ea63c2531e3d47b921039f2508189d363c3d7ad

Observation 81fec4d2-6705-4bd2-8830-503868681192 · inbound

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation cites this paper.

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:47.679305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T06:16:54.162410Z digest=sha256:091f2dc0abe3c22df645d612d6c40aa0869f8ed15c22847200f2d74abd8826dd

Observation 98cd100f-046e-4cab-8e8d-401729f7e5c9 · inbound

Towards Unified and Data-Efficient Prognostics and Health Management with Tabular Foundation Models cites this paper.

Towards Unified and Data-Efficient Prognostics and Health Management with Tabular Foundation Models Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.098161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:42:59.183555Z digest=sha256:d3a988b0caded958a68cd1d210ca48f937400ae7f0bfc466c65bb4bccab3627e

Observation 6fcbd93a-2fe8-4267-a2a0-dd2e2d0d2d03 · inbound

Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization cites this paper.

Small LLMs for Biomedical Claim Verification: Cost-Effective Fine-Tuning, Structural Dataset Shortcuts, and Cross-Domain Generalization Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.085245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T06:59:52.082717Z digest=sha256:255b07991d497cbb095e43a7ccd7fcf54770185bd3b22d38671682b2ea3a0b37

Observation cc07f78c-1766-4260-ba4f-012bf3996a37 · inbound

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries cites this paper.

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:04:37.476215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-30T11:03:52.230759Z digest=sha256:4ffd9d403382e0e83418eeec10291697de3cabe42a32fe534461fccebebf7b5b

Observation 76eaf936-269e-4997-9b52-d97e308b3a72 · inbound

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes cites this paper.

MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:34:34.377344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:33:50.358159Z digest=sha256:4b3a00cdb72ae7360b26dcd65e4b9769cf3c82eac73413c698298299a1eedb82

Observation e23af757-cfea-49b7-a47a-ff1b4e5124e0 · inbound

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning cites this paper.

FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.478910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T21:08:02.793814Z digest=sha256:5d298a92b18fc2136f846a05af6675e0128953aa3522c7cf3711e1473db80ca0

Observation 292b8c9a-bbfd-4729-81a0-eeeca18aea4f · inbound

Toward Trustworthy Large Language Model Agents in Healthcare cites this paper.

Toward Trustworthy Large Language Model Agents in Healthcare Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T09:29:31.521623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:29:31.521623Z digest=sha256:b1b4fe144cd0dc1d97489a276aa0977f7141a04d46efdd327697dcfaeb2d138e

Observation 66436146-ad99-4c73-8af2-4dc13a4f617a · inbound

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability cites this paper.

LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T15:05:03.512017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-08T15:03:14.483228Z digest=sha256:47de264e3dd3281b725be6d519900cd6ae234f4bfaa13a131ab5744ee7c50198

Observation aab801bd-2527-413d-9b33-3157153fa581 · inbound

In-Context Learning for Wound Classification with Small Multimodal Language Models cites this paper.

In-Context Learning for Wound Classification with Small Multimodal Language Models Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T14:16:01.779880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:16:01.779880Z digest=sha256:304e01e72d599e701e189661f256bb89aa9239e6bb0cdf05aa3aa77df722bf7f

Observation ea9593bb-f474-4e0e-9d9d-568a23ab9391 · inbound

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs cites this paper.

DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:38:11.801917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:38:11.801917Z digest=sha256:3e7cd4712e1b4d0edcfa5381f1f8157bca4cf9d26874f636b3d99233914f066d

Observation 39c08e43-e7f2-4c98-bf2b-6245d031a19a · inbound

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents cites this paper.

PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:21:20.359573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:21:20.359573Z digest=sha256:26257979845c01863caaa632b6b5a8b0ba49a47e2bdc3c8af174b6bdbfec084f