Pith. sign in

Paper Citation Record · LEDGER

MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2410.09733.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.09733 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:23:44.140834Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e61be1d3-bed9-4b26-a8b1-bdf0a84e9d41 · inbound

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity cites this paper.

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:23:44.140834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:23:44.140834Z digest=sha256:2e7804a906866fb7fd17c2654e83b6d1e0cc434c481142d42b23b784d29bcefe

Observation 72696188-7aeb-4108-aa46-b636a46f7385 · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.105066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:3c427b4bde1614454189ef839a1e762a8a044ef13444757218178e07b716eecb

Observation a27e9937-2230-4ba7-bf44-1d7987b2a080 · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.709860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.709860Z digest=sha256:c8e7160ca780b2c1582a2a297aeec9c6bfcab71bab320fdebca4982afbe8b5ef

Observation 9e34ec5b-9fbe-4501-974f-fa54928b9f1e · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:04.089523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:04.089523Z digest=sha256:cc318396c9b4136df6f30cac849263feca7d0543dd46754fadd05344a95cffe2

Observation 08c3f604-a780-46fe-b6b7-cb5b067c7cd4 · inbound

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models cites this paper.

Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:48.390311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:48.390311Z digest=sha256:47a62fa97c99ac511d509a5fdb91588f8a6300dbe54ebf9f7cf131bcdef71078

Observation 1cc8abd9-bf34-4c5a-b4da-50f6edddfa9c · inbound

Evaluating Compositional Generalisation in VLMs and Diffusion Models cites this paper.

Evaluating Compositional Generalisation in VLMs and Diffusion Models MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:53:59.285702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:53:59.285702Z digest=sha256:2ea41f80536ff9af154269c8dedd3632714def745f99b1d32c0bcdf8370804db

Observation eca0d95a-d76d-416e-bf40-aa99013712fe · inbound

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding cites this paper.

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:32.954917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T22:39:06.113655Z digest=sha256:c1bed66878dd35af0519de82137bdb26be44454684298167637aeb80c809933a

Observation 27372a24-07a9-4c29-9881-909e16265e62 · inbound

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents cites this paper.

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:58:14.947470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:55:36.734758Z digest=sha256:454751bc9af63a13059f4dc8271f05cde5eb3ab3a833a25e7ce26464791de000

Observation 94c3e4ce-b0ce-468f-a1b9-62f70045d21c · inbound

Agent Skills Should Go Beyond Text: The Case for Visual Skills cites this paper.

Agent Skills Should Go Beyond Text: The Case for Visual Skills MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:23.895716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:11:37.457810Z digest=sha256:155c8786ff76d67e46d3322fbf8b916f864797df035d09bd32350c6c8beddbcf

Observation 026b55e2-a04a-44d3-be46-718b85d65819 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:57.789709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T02:03:45.564122Z digest=sha256:2b87c3b478ecc02865701f4b8d520d5db6725478fc435d21ba8d110ce668ba8b

Observation 30318801-6c61-45bb-a607-f13110a5bcf6 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:40.670430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T06:25:58.872140Z digest=sha256:569cbc7c93b574d21979b34e893915821468a3069ba1dfc6c8c9aca0971621eb

Observation 7104efbb-ce55-4f2d-ab61-b9bca81830da · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:22.793660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T20:52:28.444524Z digest=sha256:e4fb5fa29f434f25f9b7169a9bd695ef453588fa6ba78cea3e090c067d4006e9

Observation 4aafa573-09b4-42ca-982f-898331afc419 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:49:00.915059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-03T22:44:16.272541Z digest=sha256:f989f4281844b51787dd9d904e334f02bd72a7b064fe06b76e3ab4e1f04f445d

Observation 00aa0689-4103-4c6e-891b-8480f87b4465 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T17:14:19.770867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:14:19.770867Z digest=sha256:e11af28bcf04708bd404bd3c2282e5819e0aa54d07dbbc1703b4f57e1f426f53

Observation 84daaa27-8626-4168-b9ee-cc3c1c2e9e5e · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:10.824477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:10.824477Z digest=sha256:25a28fd23b9bb4af082a7dd89eb596da0c2fd03fb588f9729968e8e0a6fe817e

Observation 42c656a9-cb30-441c-9236-974d2d20ed31 · inbound

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval cites this paper.

Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.310166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T15:11:57.228949Z digest=sha256:ffc994214a36430044cff5f81073d551e8a4a55f19eeed1c57228ae60f36133f

Observation 7b4dbd12-f7ba-49db-8bed-d4a396b798c9 · inbound

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing cites this paper.

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:16:44.676377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:16:44.676377Z digest=sha256:86ef4d44f604e6bdedf371e02e7027f9635ae2b8aa29abae986102101c6930a4