Pith. sign in

Paper Citation Record · LEDGER

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2605.12703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12703 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-14T20:55:50.873271Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact13
  • verified fuzzy8
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e00dabe7-4512-4bb7-9983-54f7d60efa9c · outbound

This paper cites Cl-bench: A benchmark for context learning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Cl-bench: A benchmark for context learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.477216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:4ceddc49d0a889157c447e7dcf8b5e7c16594d46e78a73022a187b85c299283e

Observation 940a004b-6b36-4823-9c16-eaa00539fe9c · outbound

This paper cites CL-bench leaderboard.https://www.clbench.com/.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence CL-bench leaderboard.https://www.clbench.com/

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.574390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:40331b07767860d01dfb6554f360c60b70bffece3676bc9ed34304beb6c74f51

Observation 115147b6-6a04-43d2-aa6e-e7cef5472f11 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:59:27.465430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:2c01f2680a290031473e3afd1b3c5d1139d3e57f4b36c030a9eb36a35e60976d

Observation 9d2b9f35-9e88-48ed-bcb7-abeb8bdc09fa · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Long-context LLMs Struggle with Long In-context Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.471087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:7055c93b22430aa05832136932fca9834ac92d55615a52d8574797859d0abf4e

Observation a953cadd-7c65-487a-ba33-653c989f3728 · outbound

This paper cites Mmlongbench: Benchmarking long-context vision-language models effectively and thoroughly.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Mmlongbench: Benchmarking long-context vision-language models effectively and thoroughly

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.460138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:2c0d29fa53bd5b864259dfe704982be72a72a807ad82e8989d21d404e25621f7

Observation 5a435841-c8b9-4dfc-ad51-0309e22c2e8b · outbound

This paper cites MileBench: Benchmarking MLLMs in Long Context.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence MileBench: Benchmarking MLLMs in Long Context

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.482431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:fb8e4eeb7bbba463b7d4cb82f47524c6cab2f8743c7c1f76e3ec4e8a9b37e78c

Observation e8012ecd-0b3e-48e4-a85b-587e8896339c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Flamingo: a visual language model for few-shot learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.591397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:ff15bdb78cd5ceb1ecdcf3837fb2ac2dc29ce4c00cf3320046c9d21eec9c3063

Observation 352b51bc-8eb7-447d-a147-9b4e415cfae9 · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.481171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:d59b8c305e71ab46f6b5bf5afcb97df30e4c939b12c47e7aa7f0a81c31dc4694

Observation 590339c4-5ad7-4e7c-a458-0c99529196e7 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.580399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:f56be443579dc81c6a74a228cee92644229d6055b9614cb637daf381736f68ac

Observation 04e23971-e7e0-45ed-9158-43abe108f9de · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:59:27.463094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:adaeebc1b7bc16154d7c971aa9ab5a9260b141fca403705c7533d5fc0f58c7cc

Observation c1bae3e1-9449-4d3c-b57c-b9090ac68e72 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.584386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:1a28623e2ae97b680bc2a87d4258d5965e0b2f70008fefa079c89086eee5c2a0

Observation 07ead220-755a-47e8-aaf1-476cffe1c0ce · outbound

This paper cites Towards vqa models that can read.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Towards vqa models that can read

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.578420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:3d1101a211461058e7ac71fcbde435eac7ff4424355cc7705482c3a126f193f2

Observation 45c7f5f1-275a-4eef-b1fe-bbc3fcc6dc77 · outbound

This paper cites an unresolved cited work.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-15T13:10:50.595893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:1df0fd5abad94ade06ce744fd3d49fc0750624d013b006a768f10b65d9f35f9a

Observation bf2118fd-ef09-4835-90c9-4a3241c6649f · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Ocr-vqa: Visual question answering by reading text in images

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.582374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:eb6a5c1d0bf87d849a34e0ffb007238f480f9decc615626803c25d533fd9bd39

Observation 577c9b37-88e4-4c83-8c6d-81925e24e634 · outbound

This paper cites an unresolved cited work.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-15T13:10:50.586507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:291ceac6808193a895973504483f7300601b7545c7c4c4e97403c8e8281acef5

Observation 41c0385d-8aa9-4cb3-acf9-8acb821913a6 · outbound

This paper cites Hierarchical multimodal transformers for multi-page docvqa.Pattern Recognition.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Hierarchical multimodal transformers for multi-page docvqa.Pattern Recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.589116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:407d944c2ef464541a1161119f1383a51d0558b60e54c552bffd2cccb937b044

Observation f4c58843-b0c7-4c18-8b48-f342b0e582cb · outbound

This paper cites an unresolved cited work.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-15T13:10:50.593740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:9cc4b9db2bcd68a567bf98a222a840dc95a30061321ef6f57bb118f2894647c5

Observation 204b37e4-2144-475f-bbda-eff516f96a6c · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T13:10:50.576322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:9f77f1fe4cf96eca98b8df5ff5f18be1357e97c1c79422b48562487e9693f940

Observation 44f91405-ceb0-4645-87cc-a4fa844a4657 · outbound

This paper cites What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.435741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:0862c74851366846a313d63872d019239fa5728c2e905157997714c09b797466

Observation 9a18e8bc-75b3-4e66-a435-f3097e4faf73 · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.443071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:bcae2e09a886669843f2fdc8a00823c669f9eca0605082a280368a52cd8d3aca

Observation e0494347-3837-40a3-95e8-2e0c2420a272 · outbound

This paper cites Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.451890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:061fc24ef37ffc0153c6b179e99b9ab32863a969bd9b6b473ca55f4b1a1d4879

Observation 817684a6-0b8d-479f-87ec-5c2447b7bdfe · outbound

This paper cites Mt-video-bench: A holistic video understanding benchmark for evaluating multimodal llms in multi-turn dialogues.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence Mt-video-bench: A holistic video understanding benchmark for evaluating multimodal llms in multi-turn dialogues

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.488076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:111713cd60ae6ed3375e98a32c9658109f4e637f38a231f830c8a79101bc6088

Observation e609a570-1da9-4af5-ba00-a91498e7fc37 · outbound

This paper cites MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.475069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:1ad84dd9f78c21d8cce5ece1eaa5286df73405b5b086924ff4064687d0f53549

Observation 9b796f75-f6d0-4afc-9c6d-3828fb8f5f0c · outbound

This paper cites VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:59:27.468715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:8d58992ed14cb54d057155f8ca8539d1ffa35abe2909d9e4c1fcde760fd9a7b6

Observation 35dda440-56d5-47d0-a120-ec6a11098604 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:59:27.457543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:55:50.873271Z digest=sha256:a56f3973ad1f1171b72f1f061e7a99f7fda69fe1708901ec8e0d1ab311c03422

Pith citing papers

No inbound Pith citation observations are available.