Pith. sign in

Paper Citation Record · LEDGER

MageBench: Bridging Large Multimodal Models to Agents

As of 24 August 2026, this Paper Citation Record lists 100 of 158 outbound references and 2 inbound Pith citation observations for arXiv:2412.04531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04531 v1

Coverage vector

measured 100 of 158 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:36:09.424513Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T23:20:01.094730Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T23:21:21.327759Z

Reference resolution

100 of 158 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40456118-d064-4593-b5d3-ab873dba978a · outbound

This paper cites an unresolved cited work.

MageBench: Bridging Large Multimodal Models to Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.961145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.961145Z digest=sha256:a44419f37f5f41ae92e1b032a32a13092a78ebcef2ed036669e35b79a04cb559

Observation 97901410-a563-4c21-b18e-35c298c9e870 · outbound

This paper cites an unresolved cited work.

MageBench: Bridging Large Multimodal Models to Agents Unresolved cited work

Reference 2

Resolution
parse uncertain
no resolver link, observed 2026-08-11T21:36:08.966044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.966044Z digest=sha256:7adac1d5b6ed92521489da56fa457ee4505c3adac356abeb44d0541f077ebe73

Observation 902f6efe-dab9-4eb1-a34f-d9e3ad83d9b2 · outbound

This paper cites an unresolved cited work.

MageBench: Bridging Large Multimodal Models to Agents Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.970373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.970373Z digest=sha256:c91bf66c7d7d3c026881ad99aa07699ce60e3ab196d3d375c3b231383ba8a9cd

Observation e1d66722-2686-425f-a0db-677b3b1169cd · outbound

This paper cites an unresolved cited work.

MageBench: Bridging Large Multimodal Models to Agents Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.974839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.974839Z digest=sha256:b33b8d829b05dcccbc84c5d1ef7d6e3471b7e57e0be6b9e4957251f88bd00f0d

Observation 902e6b67-c71c-41af-ae4c-505c2a57d8a5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MageBench: Bridging Large Multimodal Models to Agents Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.979514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.979514Z digest=sha256:10e306ac212e1b8f3c6039501346f13ed8727c21f38fa77de4ee3bf3b584cf09

Observation cea70a80-6be0-4e24-aa4b-54373ef680ab · outbound

This paper cites GPT-4 Technical Report.

MageBench: Bridging Large Multimodal Models to Agents GPT-4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.984727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.984727Z digest=sha256:354271fc38612e6315528e4cadc9eb94143d27aaba7543d5076b3d19934caddd

Observation 005d2eda-234a-44a0-a2a9-44d1bd53ef27 · outbound

This paper cites Few- shot training llms for project-specific code- summarization.

MageBench: Bridging Large Multimodal Models to Agents Few- shot training llms for project-specific code- summarization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.990262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.990262Z digest=sha256:1c26ceb08b1686e4d8b823aad3dc310b6b5976c6cf1c7746fa72c3bf4ea28e44

Observation 1a3d79e5-6ee6-4ce5-860a-b4752cdf010e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MageBench: Bridging Large Multimodal Models to Agents Flamingo: a visual language model for few-shot learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.994857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.994857Z digest=sha256:a1d9e6d96c0dc5b955db003db0da8f125d4e886efab6b2ab80f8bfb024cf5a4a

Observation 9091295d-0d1d-4ece-94ba-247f0f58667a · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al.

MageBench: Bridging Large Multimodal Models to Agents Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:08.999604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:08.999604Z digest=sha256:a86aff35b129dff9e04e7ae5b3e2909746ab0f1e7025c095cf4a2f0c843363c8

Observation bce4e702-2e47-4f6c-bb89-2be41fab4c2f · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

MageBench: Bridging Large Multimodal Models to Agents ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.004778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.004778Z digest=sha256:2d11a19df878bc825f18b1055be0924a4f18a1879963d9fa422454bea9275821

Observation 5c0e8103-feac-462d-aead-ccfe03d157fb · outbound

This paper cites Q-ground: Image qual- ity grounding with large multi-modality models.

MageBench: Bridging Large Multimodal Models to Agents Q-ground: Image qual- ity grounding with large multi-modality models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.010615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.010615Z digest=sha256:fcb84bf927160cac35e751b3e3e581a5f480a300ffc16f3c93e49cdafa025953

Observation 9e5d83fe-ecea-42f5-ad0f-69a49855b123 · outbound

This paper cites Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case.

MageBench: Bridging Large Multimodal Models to Agents Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.015223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.015223Z digest=sha256:5f137b69c85d15d8ddc0dfbffd6da767f71ea7d11b208ae55712c030e99979cb

Observation fe679658-6798-414e-8add-13e20a0eb35d · outbound

This paper cites Internvl: Scaling up vision founda- tion models and aligning for generic visual-linguistic tasks.

MageBench: Bridging Large Multimodal Models to Agents Internvl: Scaling up vision founda- tion models and aligning for generic visual-linguistic tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.020036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.020036Z digest=sha256:c617bae25c04b4c85028277499824f74bf819f6bf2de9191216364f17b128664

Observation 6d1f2977-25f3-49e6-b19a-f42d4af55cd0 · outbound

This paper cites Palm: Scaling language modeling with pathways.

MageBench: Bridging Large Multimodal Models to Agents Palm: Scaling language modeling with pathways

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.024766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.024766Z digest=sha256:a7a05bf9c7a4a987927b9280b9bc3390449229d78d7d5662ee90d0adad876109

Observation 4bd26a9b-63a2-4fb6-b15a-87b4279f65bf · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MageBench: Bridging Large Multimodal Models to Agents NVLM: Open Frontier-Class Multimodal LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.029864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.029864Z digest=sha256:2a3f3d9e6cb7543694196572704412437adba359f72f75dafd157865bcef2cb4

Observation 4f8b613a-27ec-4f33-9737-9797d26c6605 · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

MageBench: Bridging Large Multimodal Models to Agents Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.035083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.035083Z digest=sha256:c38d5d14c65b4ea830d3b980d82875a7e4fbc14652cb87171ed9ca36a6022c9c

Observation 633acbdb-57ae-4717-b15c-eca4f6993f37 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

MageBench: Bridging Large Multimodal Models to Agents Mind2web: Towards a generalist agent for the web

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.040150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.040150Z digest=sha256:466a3105b78c59a436b44526b7daa063edaee4c50427fc03b9a2be2eb0a85ee9

Observation 5fd9bf3b-a67c-4606-8b2d-d8252081d00e · outbound

This paper cites A survey on in-context learning.

MageBench: Bridging Large Multimodal Models to Agents A survey on in-context learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.044747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.044747Z digest=sha256:111a100aef25d356ce2d1d090fda27b9d7cb2b99568e68adaa237e401101a721

Observation d4253818-9249-49b7-88b4-3e366a6d9ce8 · outbound

This paper cites The Llama 3 Herd of Models.

MageBench: Bridging Large Multimodal Models to Agents The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.049827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.049827Z digest=sha256:bf26746337917c169dc7547a32995bd5eeb5e143094711fa1f67edbfde193d67

Observation 08ea16b2-5219-45f0-a256-1985a1867276 · outbound

This paper cites How Far Are We From AGI: Are LLMs All We Need?.

MageBench: Bridging Large Multimodal Models to Agents How Far Are We From AGI: Are LLMs All We Need?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.054368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.054368Z digest=sha256:29fe682fd113d417828515eb60e92f699009eeac43539cb033d1644088a8313f

Observation 7921fe95-972e-483f-b52d-a833ce1a2b35 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

MageBench: Bridging Large Multimodal Models to Agents BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.059138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.059138Z digest=sha256:6916ce5801e2ead58eab12d64468a3806c426494cf027cdadf56edc97114fc4d

Observation 96d6fedc-cc6c-4400-bab2-5df327626722 · outbound

This paper cites Multimodal Web Navigation with Instruction-Finetuned Foundation Models.

MageBench: Bridging Large Multimodal Models to Agents Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.063893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.063893Z digest=sha256:4fe5d4bab03fa9a96319af1bc0f697146f7f4158078d2c2d73d111e5e86454b4

Observation a565a4a8-a75b-4c09-99e6-70d5c01339f7 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.

MageBench: Bridging Large Multimodal Models to Agents Cantor: Inspiring multimodal chain-of-thought of mllm

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.068767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.068767Z digest=sha256:9ff9c9243cc0ac61d8231e46d902301d24bb4a7b2f0c53e1ad11aede8bd88da8

Observation 48f79f57-fc5d-4c95-bee6-7ce99cbf3a0d · outbound

This paper cites Formalizing properties of agents.

MageBench: Bridging Large Multimodal Models to Agents Formalizing properties of agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.073113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.073113Z digest=sha256:324417e5cefbd1fbadb9771802dfa233c2adc545bf54bff5dc7fca95a07a59c3

Observation 10e642a6-db10-4a42-865b-dad654e9a8cc · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

MageBench: Bridging Large Multimodal Models to Agents Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.077498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.077498Z digest=sha256:3c8d0ae562fc619b1c167bdfce592ac06d3dcf578bd87c5957a329b7dfa9830d

Observation 81e486d0-0a6c-42dc-8800-a609c8ea691c · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

MageBench: Bridging Large Multimodal Models to Agents Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.081789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.081789Z digest=sha256:98a419799fbebf977c055fb9c4960081f36939f9feb851c7dbaa1a436a370fa8

Observation 4f77febb-9195-413d-9e61-7c229870d8d3 · outbound

This paper cites ChatLLM Network: More brains, More intelligence.

MageBench: Bridging Large Multimodal Models to Agents ChatLLM Network: More brains, More intelligence

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.086409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.086409Z digest=sha256:ce6f09e61cae13f04032cf239ed7069ec25b17a44be620ba3b28d8b53c01c835

Observation e702223a-ed6e-42bf-a9e9-aa96e5930330 · outbound

This paper cites Sapien: affective virtual agents pow- ered by large language models.

MageBench: Bridging Large Multimodal Models to Agents Sapien: affective virtual agents pow- ered by large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.090563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.090563Z digest=sha256:061eec0a6724b069641ec5754c2862969f1d375201b092e84439495f7594ee99

Observation d0842d35-d8e4-47a3-a992-f390f8b0ff1a · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

MageBench: Bridging Large Multimodal Models to Agents OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.094459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.094459Z digest=sha256:1d7b0ebdbf218154761d991229a3fb1a4eb7ec6dcd1f11514d7500f9bff7d88a

Observation dd7a9de0-fd77-42df-9268-83dbc0e70f59 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

MageBench: Bridging Large Multimodal Models to Agents MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.098831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.098831Z digest=sha256:f3f30bc2ec896dcf563d8a9df0524466a13f81fd9341b6cb6b0a5b0d5d0ce88f

Observation e4b8eed5-48df-4373-873a-7fe68babdda0 · outbound

This paper cites Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning.

MageBench: Bridging Large Multimodal Models to Agents Look Before You Leap: Unveiling the Power of GPT-4V in Robotic Vision-Language Planning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.103146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.103146Z digest=sha256:750042ce8a5196a736b7dab5ebd112f2036ce92fab87c95c9e649a02ad6845e2

Observation 67ee2fa6-f894-47e4-b932-b6e5549c10b2 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

MageBench: Bridging Large Multimodal Models to Agents Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.107791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.107791Z digest=sha256:c9408c00df91d82204b3e420a6e10ff8d3d65f9345bd90a7bf4a672a508fd0a7

Observation 327fafea-72df-4157-ad40-296f00e54f14 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

MageBench: Bridging Large Multimodal Models to Agents Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.112585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.112585Z digest=sha256:a9ddb5a3925b7b8d8e753e542bc6c36728cee6a08cc0121b6eaa341c87d69726

Observation 4bab5268-be5d-429b-803a-7f47669020ae · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

MageBench: Bridging Large Multimodal Models to Agents TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.116911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.116911Z digest=sha256:9adb57c8fd12ea409828beda75a13e97a4379a12376892927dc0ca7485ba367e

Observation fa5cf90d-bce3-44ec-9a94-c1160a187e45 · outbound

This paper cites Large lan- guage models are zero-shot reasoners.

MageBench: Bridging Large Multimodal Models to Agents Large lan- guage models are zero-shot reasoners

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.121635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.121635Z digest=sha256:8c8faa50d2dafa269de2fc30b1960b4f30a2c57588d249dc4a287fcc9cd88ab8

Observation b730a4c6-e814-419b-9b2b-13321c3e62f0 · outbound

This paper cites Google re- search football: A novel reinforcement learning en- vironment.

MageBench: Bridging Large Multimodal Models to Agents Google re- search football: A novel reinforcement learning en- vironment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.126287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.126287Z digest=sha256:24118aee632376969aabf13b32cbef002e7215520ecdc9d81c72b9b8363cd562

Observation 73578bc7-8659-4232-b307-3668ec28137c · outbound

This paper cites Learning the user’s deeper preferences for multi-modal recommendation systems.

MageBench: Bridging Large Multimodal Models to Agents Learning the user’s deeper preferences for multi-modal recommendation systems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.130786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.130786Z digest=sha256:a584b8dee8701a782e93768cb1f06a5f1cd07e770d2528da05f052116a137cba

Observation deb0d743-c846-44d6-a3de-7eae714c4d94 · outbound

This paper cites Seed- bench: Benchmarking multimodal large language models.

MageBench: Bridging Large Multimodal Models to Agents Seed- bench: Benchmarking multimodal large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.135407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.135407Z digest=sha256:441b4c92d854911696826a02661f5a02788ac97dcc4c88153803dec60e1aa77f

Observation d22346d8-ad10-4c5d-99ad-ef32bada2684 · outbound

This paper cites Camel: Commu- nicative agents for" mind" exploration of large lan- guage model society.

MageBench: Bridging Large Multimodal Models to Agents Camel: Commu- nicative agents for" mind" exploration of large lan- guage model society

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.140554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.140554Z digest=sha256:e73cca7266e1ce5ccc20d98ce586729b143247c2c9dd9f7a61e4ebff40b1687a

Observation 316034b2-5d76-4439-a7fe-c5e8468987d4 · outbound

This paper cites MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?.

MageBench: Bridging Large Multimodal Models to Agents MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.145274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.145274Z digest=sha256:2a5b4e6a0a9e86086489fe02de84a4d99a5beada6be6eebf8095e31fb1cc734e

Observation e80899af-0633-4570-9005-b185d676b7ea · outbound

This paper cites Ap- pagent v2: Advanced agent for flexible mobile inter- actions.

MageBench: Bridging Large Multimodal Models to Agents Ap- pagent v2: Advanced agent for flexible mobile inter- actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.150148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.150148Z digest=sha256:23ed8ee7faca46dbde777108c2ec7f62c65321aeb8f6cd690679a10d971f2db5

Observation 7f337c62-d729-47d3-bc42-0c42ad7653aa · outbound

This paper cites MCU: An Evaluation Framework for Open-Ended Game Agents.

MageBench: Bridging Large Multimodal Models to Agents MCU: An Evaluation Framework for Open-Ended Game Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.154856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.154856Z digest=sha256:835df6710823973556cf241de61c918ad1aebaee570193e403038605f268c931

Observation a9a4a946-812b-4f44-9c06-f4ff7272d754 · outbound

This paper cites Visual spatial reasoning.

MageBench: Bridging Large Multimodal Models to Agents Visual spatial reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.159517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.159517Z digest=sha256:52ff6e5afe45e80858171e2ea5a7a289b79139e8e21c2fc4ba14f10cfb02d3ef

Observation 7999d476-e67e-483f-8244-cae3bb70df51 · outbound

This paper cites Improved baselines with visual instruction tun- ing.

MageBench: Bridging Large Multimodal Models to Agents Improved baselines with visual instruction tun- ing

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.164077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.164077Z digest=sha256:e4e1d1d29bde207142ac75c5b6168c98fbcc65cecb3af595876008e2f0015255

Observation 31e93e55-1923-46f8-af39-b93642bacee7 · outbound

This paper cites Llava- next: Improved reasoning, ocr, and world knowl- edge, 2024.

MageBench: Bridging Large Multimodal Models to Agents Llava- next: Improved reasoning, ocr, and world knowl- edge, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.168614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.168614Z digest=sha256:f566574c2a1f83f8c5737537e38b591f3976f1eeebb4dee3ff81592189c6c4ad

Observation 0335a167-cb95-43b6-939e-f37cbfd44e9a · outbound

This paper cites Visual instruction tuning.

MageBench: Bridging Large Multimodal Models to Agents Visual instruction tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.173371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.173371Z digest=sha256:e9c39965ffc81696998c0cdda370a690f99320e996b7f0b6d7faa0ebb36c1259

Observation 7d59b52b-546f-4ea8-a7e3-83f6f8565691 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MageBench: Bridging Large Multimodal Models to Agents AgentBench: Evaluating LLMs as Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.178312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.178312Z digest=sha256:bc086eeab9107a263791a92ca54ae53ca244be07a30e73187599a5ceccf564d7

Observation 336864a7-8cf6-4a1b-8532-329500031531 · outbound

This paper cites VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents.

MageBench: Bridging Large Multimodal Models to Agents VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.183086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.183086Z digest=sha256:45879d0c05e511b94685b25e84be2a53314e259d60aa19ba334a7de760754086

Observation 6ef50f4c-4269-4fec-9829-e9aff34586f5 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, 2025.

MageBench: Bridging Large Multimodal Models to Agents Mmbench: Is your multi-modal model an all-around player? In ECCV, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.187731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.187731Z digest=sha256:156d8bdceb581a4ee9c4957bd62850632fff3a2be13e5f699130778ebe0863f0

Observation 63352101-e00f-4da0-9aa6-a91ada5d8918 · outbound

This paper cites Artificial empathy in marketing inter- actions: Bridging the human-ai gap in affective and social customer experience.

MageBench: Bridging Large Multimodal Models to Agents Artificial empathy in marketing inter- actions: Bridging the human-ai gap in affective and social customer experience

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.192487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.192487Z digest=sha256:84ef23ce52bebc402ea651c2a17a9c20987297df6d92c65bdb3394522896f135

Observation 67b72738-45d9-411f-be7d-a0db5bbdb06f · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

MageBench: Bridging Large Multimodal Models to Agents DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.197848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.197848Z digest=sha256:b93e8c7141f0ed48c94598c18cd41b657786a7ace86358f4a155d567b5b86593

Observation c75b8675-174e-462e-8dc6-2a510442d302 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

MageBench: Bridging Large Multimodal Models to Agents Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.202647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.202647Z digest=sha256:378397896e6505330d1986e515afe06ed6de2a736b259046e4605c2198fd3e67

Observation c41ebba7-946e-4008-acf2-34b7d0207535 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MageBench: Bridging Large Multimodal Models to Agents MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.207178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.207178Z digest=sha256:4cde4dd515f653530a10043beff45759576f8c731c7e8c7277eb17f48964387c

Observation a3dc9997-63c0-46f1-997f-778ba9605f0b · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

MageBench: Bridging Large Multimodal Models to Agents ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.212407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.212407Z digest=sha256:d521d89c2fb569128de779a0a801ac9f0641914f3247fbdf220aa029751a7693

Observation 499b8b21-c2eb-4479-810f-63fd3d028345 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

MageBench: Bridging Large Multimodal Models to Agents Compositional chain-of-thought prompting for large multimodal models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.216645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.216645Z digest=sha256:2407ce835223c0a6b65fc01a06c4872192bf449b7afb09b2f29b28d82c60e15f

Observation 25952052-3eb7-4d32-b429-d9d4208edae6 · outbound

This paper cites Adaptive Machine Translation with Large Language Models.

MageBench: Bridging Large Multimodal Models to Agents Adaptive Machine Translation with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.220579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.220579Z digest=sha256:510d230aeed292a3ed3b4a24f06910e83230ff93be7f1c11efeb18f95acf2f02

Observation 90494260-fcad-4515-803b-d0e2be654d97 · outbound

This paper cites MobileFlow: A Multimodal LLM For Mobile GUI Agent.

MageBench: Bridging Large Multimodal Models to Agents MobileFlow: A Multimodal LLM For Mobile GUI Agent

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.224772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.224772Z digest=sha256:e3bd608d66057124dd1e89ac35724b27887b818ee2bcdbcfdfc159d17171fee3

Observation cecbdb22-3761-423a-9c69-a6f89c5e5aca · outbound

This paper cites LIVE: Learnable In-Context Vector for Visual Question Answering.

MageBench: Bridging Large Multimodal Models to Agents LIVE: Learnable In-Context Vector for Visual Question Answering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.228770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.228770Z digest=sha256:68efadacf373ba8a736fed22b235bad4140c7df9e42ebd92a7eae0a966118f95

Observation 83695d78-4b20-42fd-96b7-8f7342f54219 · outbound

This paper cites Summarization is (Almost) Dead.

MageBench: Bridging Large Multimodal Models to Agents Summarization is (Almost) Dead

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.232916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.232916Z digest=sha256:4d486a582e6bb7b40ffdf82c85c946148038a5444d516768175d22dab739a7ac

Observation f844f83e-3012-4f85-8484-bf1bc9f5f493 · outbound

This paper cites Virtualhome: Simulating household activities via programs.

MageBench: Bridging Large Multimodal Models to Agents Virtualhome: Simulating household activities via programs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.237547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.237547Z digest=sha256:c3c2ab066d4114783df8aa95321264c6695a05bf9b9ba5d03ed3fc2e62455004

Observation 27da8b06-43e6-4aa3-8790-012945fdab72 · outbound

This paper cites Imagination-augmented agents for deep reinforcement learning.

MageBench: Bridging Large Multimodal Models to Agents Imagination-augmented agents for deep reinforcement learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.242109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.242109Z digest=sha256:ed2b89004600a11fb3f1deb05d11211b5f42beb809caef10f26295d83305780c

Observation 2c67fb62-87a1-46be-879b-2fc57af3e811 · outbound

This paper cites Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering.

MageBench: Bridging Large Multimodal Models to Agents Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.246562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.246562Z digest=sha256:b7de5b602bc3cb4f127ccba8966d58b6833f2ca41fd35411d6417631b8b36e47

Observation 52f1354d-da9b-443f-ae2d-6b102b19bc5b · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

MageBench: Bridging Large Multimodal Models to Agents A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.251713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.251713Z digest=sha256:ecd7c5a2f68db4bc4a9b7c21cc241fded207275fecc954e2782e8ad7d6f3d415

Observation 6e8f7bf9-41cb-45f4-a278-2119fa5ece3e · outbound

This paper cites Image captioning for effective use of language models in knowledge-based visual question answering.

MageBench: Bridging Large Multimodal Models to Agents Image captioning for effective use of language models in knowledge-based visual question answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.256537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.256537Z digest=sha256:595b095fe29c736850355030ae75e2188dfce67dd5fd00cc7130565ac4085d38

Observation fef4287c-0f20-4bac-9ddf-9334fd047860 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.

MageBench: Bridging Large Multimodal Models to Agents Reflexion: Language agents with verbal reinforcement learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.261253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.261253Z digest=sha256:37088ddc11d383e5bd4a7b410a1ed8813104d53b52d22982ce81509c561e6049

Observation 09cab709-0ac9-4184-9102-f83246593e71 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

MageBench: Bridging Large Multimodal Models to Agents Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.265836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.265836Z digest=sha256:8c813d72916348721acffc2e7af598f031ef9b4ebb9c5637a6a786b4143ab2d0

Observation 31b9b961-7e0f-48f4-8808-ca3fd46ace3d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MageBench: Bridging Large Multimodal Models to Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.270532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.270532Z digest=sha256:a81f4134af8ecee7da63107350d5e4f80513baa8e1c7a28baee80b67c7525f23

Observation eb495731-91bc-43af-b8d4-5f92929a7089 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MageBench: Bridging Large Multimodal Models to Agents LLaMA: Open and Efficient Foundation Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.275304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.275304Z digest=sha256:92ff37b9985a7af61d6616c7cba1c74f7f3cc4b4567d36736e32f306f5947870

Observation 97909eb9-2c68-4c75-a47d-9da59aae5e37 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MageBench: Bridging Large Multimodal Models to Agents Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.279938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.279938Z digest=sha256:2a0c04390f234fc184f800cfa4fb773f54dd9919db8af1b5827b30907dd5be69

Observation 8d6b3a57-8554-4111-b53d-c96a3d1ec53b · outbound

This paper cites MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning.

MageBench: Bridging Large Multimodal Models to Agents MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.284564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.284564Z digest=sha256:4cbc50b3e281416cec9efd4bbd0f36434c3cd6d6c7e5bea1dc773065b0b11a5f

Observation 1f732bde-9bb5-4ef0-b662-28f412b1e363 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

MageBench: Bridging Large Multimodal Models to Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.289737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.289737Z digest=sha256:468ce404af1e6278ac97a285d14a8d95945270b2c8376ba0792350838e1dfa74

Observation c4c391d9-0656-4cf4-9e75-85d99589764e · outbound

This paper cites Qcap- tion: Video captioning and q&a through fusion of large multimodal models.

MageBench: Bridging Large Multimodal Models to Agents Qcap- tion: Video captioning and q&a through fusion of large multimodal models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.294670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.294670Z digest=sha256:b1aceb9bd8bde4dd04449fed05fb5331f6ade5a56dd02dfd529b87517fb69f8b

Observation eed7a586-31b0-4664-8169-54e925c88454 · outbound

This paper cites Large Language Models for Robotics: Opportunities, Challenges, and Perspectives.

MageBench: Bridging Large Multimodal Models to Agents Large Language Models for Robotics: Opportunities, Challenges, and Perspectives

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.299117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.299117Z digest=sha256:670c673675f07cdc1ac10afc7ab8154807a70e5baa3a0715c7c1c8b7459d0420

Observation 791d9f35-9e26-4ce0-b59a-7a2061563c66 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

MageBench: Bridging Large Multimodal Models to Agents Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.304053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.304053Z digest=sha256:bc859381057301d51614fae94670a8792f3d4fb822c6fb702380e079f618e7d2

Observation ed2eade3-0dd5-41a5-add2-a10acba6b602 · outbound

This paper cites Document-Level Machine Translation with Large Language Models.

MageBench: Bridging Large Multimodal Models to Agents Document-Level Machine Translation with Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.308985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.308985Z digest=sha256:531a10f9c20e7b0770b356bc05f83f6d5b90941474eac46fdf4fa5c2bac601d2

Observation 34bb36df-8a40-4271-8eef-6c6b6ce9b09c · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

MageBench: Bridging Large Multimodal Models to Agents MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.313830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.313830Z digest=sha256:75c21e1a52bddc804751242b42d1b5c3ab06785235e5f185964fc38c5193fd05

Observation cecd93f5-2e93-4ed7-bc6c-723fc870eefb · outbound

This paper cites A survey on large language model based autonomous agents.

MageBench: Bridging Large Multimodal Models to Agents A survey on large language model based autonomous agents

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.318773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.318773Z digest=sha256:1977cfa6c645eaf7ec72829bb9c90363f19e8dc471e45cfdee06f83bf3d3c072

Observation d4a277c0-1e0a-472a-bb89-90abb73f5fa4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MageBench: Bridging Large Multimodal Models to Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.323230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.323230Z digest=sha256:c50177f1848ba7595157f3fb691468e0e762a42ab2b2030f3f6abb5b818b873a

Observation 1b530835-45ee-4dcd-aa3e-9ecc663f501b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

MageBench: Bridging Large Multimodal Models to Agents Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.328415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.328415Z digest=sha256:c57f809c0fa744e1d0415b4d738c10f545ff8dc068ab7ea15db5695dabc0e973

Observation beb7e97b-11a2-42cf-bfeb-ae13103f55c2 · outbound

This paper cites Chain-of-thought prompting elicits rea- soning in large language models.

MageBench: Bridging Large Multimodal Models to Agents Chain-of-thought prompting elicits rea- soning in large language models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.333258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.333258Z digest=sha256:b3228cce9fb673269649ee7c0e15b7a172a870173fae5d43224745802b75a252

Observation 0260d9fe-3979-416a-b8b2-b8aa6210a55f · outbound

This paper cites mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning.

MageBench: Bridging Large Multimodal Models to Agents mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.337291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.337291Z digest=sha256:20ef616d2b7b898196a684d254da9d1b3590ae0d8d6877ba4f628cfc58cf4d54

Observation 204ab3bc-dceb-49e3-9fbd-a7ee16241952 · outbound

This paper cites Intel- ligent agents: Theory and practice.

MageBench: Bridging Large Multimodal Models to Agents Intel- ligent agents: Theory and practice

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.341754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.341754Z digest=sha256:366c77f3a4e1d7a86eb18b6b3b200df9c98a3451e5b8f9be50605de10387a992

Observation 0bd74fb2-6ac9-405a-83a5-48349851bb94 · outbound

This paper cites A glance at in-context learning.

MageBench: Bridging Large Multimodal Models to Agents A glance at in-context learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.345648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.345648Z digest=sha256:93df09490ab3e4d7b712bed5e1f76481d5c0dd61cabae66498aaee7c53188156

Observation 991490d4-96fd-4929-9e5d-a24d6058744e · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

MageBench: Bridging Large Multimodal Models to Agents DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.349963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.349963Z digest=sha256:e8ba8322d303c22b75cbb2c813ae61755845556aa4ced19540f994db4dadcacc

Observation 3c7483d7-c1cc-4879-aa9f-b1e625617948 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

MageBench: Bridging Large Multimodal Models to Agents The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.354143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.354143Z digest=sha256:3fae0d0dc2363f488c2876bbd1bad9fb4d3acf3885e94bf732594f9d9a763c70

Observation 6b75a0a2-f551-4abb-a8fa-5dfcb60e6847 · outbound

This paper cites Large Multimodal Agents: A Survey.

MageBench: Bridging Large Multimodal Models to Agents Large Multimodal Agents: A Survey

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.358806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.358806Z digest=sha256:d4843d970b54465183036da45dd9d7dd6ec7e8631d47d2336913a652691f5719

Observation 76555a79-5944-4ded-86ef-ddc863aec2e1 · outbound

This paper cites A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models.

MageBench: Bridging Large Multimodal Models to Agents A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.363607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.363607Z digest=sha256:6fe06168949e65f5ed75b93d7a136a41b4f2b9da5d2257ca3d177994d5e4df93

Observation d8487213-def1-4fcb-aa64-0299172b944b · outbound

This paper cites Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions.

MageBench: Bridging Large Multimodal Models to Agents Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.368161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.368161Z digest=sha256:30f1377a6eecc4dfbdd7f812dcca59ef89dfbc7434150aa2cee362b948d55110

Observation c918bfe3-d530-45be-b6ed-94781a47be75 · outbound

This paper cites Exploring diverse in-context configurations for image captioning.

MageBench: Bridging Large Multimodal Models to Agents Exploring diverse in-context configurations for image captioning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.372857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.372857Z digest=sha256:4812d2414992e76526a5338ce34ae929ba4632a37eeba9e137629f46fe2e07cb

Observation 8fa5f103-92e9-471e-a042-882f92a9761e · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

MageBench: Bridging Large Multimodal Models to Agents Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.377298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.377298Z digest=sha256:e0a632743518c9475e37577e383340945a2e06c6bea3ea5c8fb8e3e1c859b196

Observation 8bf58541-440e-4cec-9cb8-25b28ab2e68f · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

MageBench: Bridging Large Multimodal Models to Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.381777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.381777Z digest=sha256:f038fd33b349df3600e252ee1b856acab1654d2140fd99117f31de6872b5b6f8

Observation d299f22b-4948-47b0-b907-188d7fb51566 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MageBench: Bridging Large Multimodal Models to Agents MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.386551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.386551Z digest=sha256:28c81cdf4a1a7d17886771749ec925550ade3e13e432600626df416aedc93974

Observation 73af8712-c1c0-4b49-a29b-129dbda5c1ac · outbound

This paper cites A Survey on Multimodal Large Language Models.

MageBench: Bridging Large Multimodal Models to Agents A Survey on Multimodal Large Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.391392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.391392Z digest=sha256:71a8393e0494835bf14fae15b45167cb697215e65c66532ec87481885dad3f32

Observation 6d169184-9103-4bd1-97cc-a89ed4e17aca · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MageBench: Bridging Large Multimodal Models to Agents Yi: Open Foundation Models by 01.AI

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.396491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.396491Z digest=sha256:32aa327fa4abb09d9066d83144a88670d5d0700d34d414b166094633c19d4736

Observation ea5e4b54-3e0b-4f4d-9ee5-47072a66dec5 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

MageBench: Bridging Large Multimodal Models to Agents MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.401490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.401490Z digest=sha256:e236135967c2c5ca6c92bac5ddb70493e0e43762953ebc6b9f01d0154dff8f06

Observation ee2878c6-602b-4667-b1be-62b80f17641a · outbound

This paper cites Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi.

MageBench: Bridging Large Multimodal Models to Agents Mmmu: A massive multi- discipline multimodal understanding and reasoning benchmark for expert agi

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.406415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.406415Z digest=sha256:2819033e508810520faf267f51f38cd7deabc83894e84b713403397a57d87a01

Observation a0c21e55-b939-4b44-b525-3193f2df6b98 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

MageBench: Bridging Large Multimodal Models to Agents Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.411029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.411029Z digest=sha256:0f3039c6634346246930d644f8880616f353ed12a77e2acfe562c859e85028f3

Observation 021b4051-b969-4c37-b467-b6a6d6de3096 · outbound

This paper cites Transporter networks: Rearranging the visual world for robotic manipula- tion.

MageBench: Bridging Large Multimodal Models to Agents Transporter networks: Rearranging the visual world for robotic manipula- tion

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.415588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.415588Z digest=sha256:95e63aa65576237f6cd5543672487adf12ff729cce7e8e42cd50b136e21510a4

Observation 72cce690-f700-413f-9491-5e59939442e6 · outbound

This paper cites Large language models for robotics: A survey.

MageBench: Bridging Large Multimodal Models to Agents Large language models for robotics: A survey

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.419922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.419922Z digest=sha256:1e9e7f043f86f96997d683965773e3383968fa38e2660e4cc1a999cca9496962

Observation 67cc5245-6342-409f-b239-b3ede1801e52 · outbound

This paper cites 12 Prompting large language model for machine trans- lation: A case study.

MageBench: Bridging Large Multimodal Models to Agents 12 Prompting large language model for machine trans- lation: A case study

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T21:36:09.424513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:36:09.424513Z digest=sha256:fec54a945de0b8f730393801c84e930d794f903cef7a4a63ca84aa7a389aa608

Pith citing papers

Observation 9c172641-d4e1-47f8-9e53-b03f9e409b3f · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL MageBench: Bridging Large Multimodal Models to Agents

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.355848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:e3f3a802472a507d58df72942872f0e7fd7fe5e96569121eb3fa62d3fb6fdcc7

Observation ba845c56-06b6-4390-8b63-95c6ee186084 · inbound

LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator cites this paper.

LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator MageBench: Bridging Large Multimodal Models to Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:21:21.330732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T23:20:01.094730Z digest=sha256:6c4599bba842b84bfcab1c048c6d95dfff879deb072d5254b0c595e834254b20