Pith. sign in

Paper Citation Record · LEDGER

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.21741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21741 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:31:26.112925Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b290fad-7789-49b4-81c4-471ea9f4f717 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.450246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.013359Z digest=sha256:dbc907a7b4d54f4647eb2f3ba68474ee81164f3caa7802823d22c777929ac517

Observation d6158cf1-a03d-4fa7-82d0-0b2f80fb5002 · outbound

This paper cites Uniter: Universal image-text representa- tion learning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Uniter: Universal image-text representa- tion learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.414687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.029256Z digest=sha256:687b32129c12b44132e99f9f4d1d1749d32e60bafa4ec868f82c0adc82c38786

Observation 19b30b09-1918-4e97-9c92-8b68e719fc18 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.032238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.032238Z digest=sha256:5bb649973c019333cee0de8ad8ae06eadd367a69d43ecaca649159ac8aaecaab

Observation f5cbb349-b375-4e07-979b-b13e43915fe4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.038126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.038126Z digest=sha256:3aa1fba25f963cc7ec6ffcd98ee12eec3476e0262bf3798f326efa3d5a92a804

Observation c1c8e38f-9a7a-46f5-88c9-4a4c482968e3 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with in- struction tuning,.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Instructblip: Towards general-purpose vision-language models with in- struction tuning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.399980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.040867Z digest=sha256:a3d518c995e7545089a4bb114a8f9b4ace131fa75f7b9d87ba2173ffa31e14c0

Observation 20f951cd-4963-4261-a512-14d0829a679b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces NVLM: Open Frontier-Class Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.043505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.043505Z digest=sha256:5a1f3c59a1faf889553f5db120f75f3635a511fd043259ba5b3f2decdbb9a912

Observation 68acd8d9-cf1d-4eb4-b295-f68536314c2a · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multi- modal large language models,.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Mme: A comprehensive evaluation benchmark for multi- modal large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.391906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.046621Z digest=sha256:35a4533db67ec1ef60253b62f5fc2b5dcb6ef40fce8743f361fe7bcdceb8b674

Observation a444735a-947f-4d70-b860-00624c37a323 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.049030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.049030Z digest=sha256:eaedfce5667fbb20512ea48b540cf522d69ce0fd8eacb0b1fb89c84b74af6134

Observation 8b4015ee-c20f-46d4-ab7a-afa9f0be0a91 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LoRA: Low-Rank Adaptation of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.051888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.051888Z digest=sha256:59445023d0275e7efc47fd26f6f737fe2dc160165f7780262f3db020399c4551

Observation b2704b91-d00a-4a38-96a5-37ccec7f7ff1 · outbound

This paper cites Dvqa: Understanding data visu- alizations via question answering.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Dvqa: Understanding data visu- alizations via question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.383131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.054721Z digest=sha256:83cdd6b26559404c4f3f566578a078640df96ef6cbaa82eac6d247637a53f029

Observation 7ba075e9-3b99-40b1-8010-230270d6c175 · outbound

This paper cites VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.059865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.059865Z digest=sha256:5e294e24427da805c5e0f4c7a36c8cbe8a409e110958f0a7ef476f1d284a75a0

Observation 1c167ae1-5c2a-43d5-825b-1f490f60785a · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.366203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.062802Z digest=sha256:81c0876ce98d76ae1556536c30903ec1d9aa595a5e78f2f0017365cc87136cc4

Observation 1ba7ca47-eae9-479e-a316-ec72aeb044c8 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.065302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.065302Z digest=sha256:f8565f5705c5f5eb71792ecc8da4140d2f54d3a1997fe1d3155e09f4f6a8e0cf

Observation 76aa291a-bc34-46bf-a010-0976aa03728c · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Evaluating Object Hallucination in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.068491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.068491Z digest=sha256:dbd3c382811f630cf24c5448a52e4f766ad09ad5ac03a5327f8156cc44515e6a

Observation ede1ea7f-a126-498b-8996-33326c899b63 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.071331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.071331Z digest=sha256:ea424fbc3cbb57c3b711fc330d949532139744c4e42c78c78c626c6acac8c040

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:d511fbd0b4d920dfdac6b08ff9c63beb937fc99773e0bf682ec950a294de22f8

Observation 1dd6caea-b43d-4aef-9ce0-1164a14fa382 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Docvqa: A dataset for vqa on document images

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.356567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.080676Z digest=sha256:e15dca42cc86fed7d5225d4165534720ac0e8a31a906ea2715b2f871e3f059d1

Observation 0249a39e-5b92-43f2-8dc0-9867770aaaa5 · outbound

This paper cites Hello gpt-4o.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.347968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.083312Z digest=sha256:837eb80efe3b6144b0db6f083765a6bcabcc418cd1734fb2387f0ea7907111bf

Observation d1136b45-0eb7-4423-b9cf-45ab13f36844 · outbound

This paper cites [Radford et al., 2021a] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces [Radford et al., 2021a] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.338397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.085780Z digest=sha256:ca46ab8c06591510590966c5115b7148a520473889ab4644be2f2dcef5314739

Observation 36f9e785-7302-4ec7-a97b-9e954a01647e · outbound

This paper cites Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.329223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.088243Z digest=sha256:f53d0f7ded5729ec35f6186ef761b1c5617487da4baca2b0fccbd8d4f28ff3a8

Observation 3fab29e7-4af2-4062-99ec-f3d863f35600 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.091282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.091282Z digest=sha256:fc6c4029df02177873f2cca4c4336a8f9f91f951b8c7fd6459684e437ddfb0b2

Observation 3a37acac-9ece-45ac-b4fd-db999bb3fd8b · outbound

This paper cites Attention is all you need.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Attention is all you need

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.320574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.094018Z digest=sha256:1f1044dbe6f7cb12ff6c5a6cda30558edc2db9a4ac4a96ffbc48ec4cf60c4421

Observation af5e91bc-0260-4800-a078-d8b76e320927 · outbound

This paper cites Introduction to convolutional neural networks.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Introduction to convolutional neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.311833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.099172Z digest=sha256:03fd171204e767722096742815db6d11fbb8a03633d4d81fea2f9727e8b90c02

Observation 755eb3ef-c83d-46cc-b854-83e06da943dc · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.101813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.101813Z digest=sha256:3e922a57c577822598e66d460f790139fa6fd3ff3c3d2a89a1a7b30f5a8262f9

Observation 601549d3-cebe-4ac9-940d-289cc83354ca · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.104757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.104757Z digest=sha256:e8b1f8d225212fc1d6c363b597a8523ed8945c5b9598637094a4772e1fa17529

Observation 45b1656b-993a-4acb-a062-d5813e053d15 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Florence: A New Foundation Model for Computer Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.107503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.107503Z digest=sha256:d01bbda2f66f05ffd9296e75720431f2fdf261c8d3f764a98a4d67dee5029dd0

Observation 41f0bfe4-120a-48ef-aa2e-b2bde586493c · outbound

This paper cites Easygen: Easing multimodal generation with bidiffuser and llms.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Easygen: Easing multimodal generation with bidiffuser and llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.303148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.110181Z digest=sha256:0c2aef2687948658e32a1b3b44ac1254a192799ff0a4b30dfb2530a67c93359d

Observation 832fbd14-48b9-4ce9-a888-c694d51f440d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems , 36:46595– 46623, 2023.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems , 36:46595– 46623, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.293335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.112925Z digest=sha256:606da384b51e5ba2fd283ddd5b1a713d753238ab6295768d3d634cc15a5fd652

Observation 67244016-c8d3-4d9f-9475-48c2196fe6d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.096481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.096481Z digest=sha256:2b1570a1ccb6c08a7b488e747177b22ac51c9038523ec901115f15ccc8206a91

Observation f39d8168-5047-4033-9f8d-a1275fdd40f3 · outbound

This paper cites Ocr-free document understanding trans- former.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Ocr-free document understanding trans- former

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.374882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.057257Z digest=sha256:663e0ac14dee4ee0d0d7d52fed4c34adde0f9af6fb85231838774c9039713ea4

Observation e640e00d-ee1a-49ad-9e67-2782314a0c7c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Honeybee: Locality-enhanced projector for multimodal llm

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.423521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.026185Z digest=sha256:c722f99b2e853b29ce63a103e41d238f31a110f63771eb571e2555d8c4b1f246

Observation bc5c1853-1c5a-4793-94e1-21bd4be5d566 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.035221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.035221Z digest=sha256:edd0c32a307f3474e104e44db2127f73826cea20291e42837b05c6b3bfdcce50

Observation 2a606dd0-01a3-457e-987e-de605d458991 · outbound

This paper cites Claude 3.5 sonnet.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Claude 3.5 sonnet

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.441144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.016801Z digest=sha256:29c0f6bd514a08c65ed172084ba8d4a10c528b85a3e98d9d740d8c72d80d92bb

Observation 9781c94c-d3a5-4783-b2cd-f10bc6e367d7 · outbound

This paper cites Language models are few-shot learners.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Language models are few-shot learners

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.432442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T12:31:26.023150Z digest=sha256:dd9502b84467e2ef7b3e3a685aa16598b9865294463f12a1599dc5266465d49c

Observation fabbac7f-b0b7-49b4-a84c-845b2a2754cc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.019701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.019701Z digest=sha256:140de23e88d4f2099f53907f772fd8b864277297ec5ce75b6dbd0948e3057360

Observation a858a163-883c-4807-a054-80b7481666a0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.077842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.077842Z digest=sha256:44b18b60409df3f201aaf4c208122bf4f4f48e14d03bdc5eab775d00eef33372

Pith citing papers

No inbound Pith citation observations are available.