Pith. sign in

Paper Citation Record · LEDGER

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.08497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08497 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T06:30:28.462917Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact26
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 118692d1-fc3c-44b5-b6d3-f4442849222f · outbound

This paper cites GPT-4 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.563355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:fd608bd6371de577c5f1c90f94d4b765572e183fb3f8e566eec8f9af8f959db3

Observation 11e1e380-6d78-435e-ab1d-989592fae595 · outbound

This paper cites Qwen Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.572324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d870a33167d8882df8d6c4e163dbb4810edb1648176664d6d7ee5d7507679166

Observation ee0602c2-0354-4ea2-bb65-7a2d4e0d4002 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.975236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2363e8957b3f4ed2e7cf457439c112f3a28da348508ead58d126671e5413c353

Observation e567e0b5-1cc4-4e0d-915f-6656802e901d · outbound

This paper cites Qwen3-VL Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.558779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c49a6cdeb8fd538f1beb05c5e86cf3650affba8f9122b61679e3f436a74f9601

Observation acc2e0eb-e4c1-4c16-a46a-6db0cfd7c988 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.973493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:a1c0bc0f4521b83f056b1b3ec717920f7d9dff104502dd91cbff96b6382dcdba

Observation 7a6a227c-33ce-4bf4-9ceb-f9880e313684 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.570188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2f869eca6fe16a5370b8a5556be724203bfde3af5511be4cac36a12844e75bc3

Observation 238ab7e4-30dd-4ca3-a8c5-82e808772937 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.567988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d33f04833f9053c8716905e56b3c56c74ca87e8760c03821f890a7a2173041c6

Observation 2a2f523f-2b61-476e-8eee-b7856980268b · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: A memory-augmented multi- modal agent for video understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.969766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1a1edd20469f078c56a6d925b7cb3610b2a6486ac03f5f9046f28b6febe0f69e

Observation 614e779c-392d-4a80-a56f-5936bfe66faf · outbound

This paper cites Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.553796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b2eb38ef9122d3c1b22d6b1afd78c638e19503c83c0259a20e77652098e7aff9

Observation 3d65c253-8658-4c0e-a091-d89bfa607dbe · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Metagpt: Meta programming for a multi-agent collaborative framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.968036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1bd8cd2b8b4b1a59a37c62ce3fae108363cc11ef8f96846526c8900956485b08

Observation 4c0f5849-0085-4fe7-a67c-4d89f32e16da · outbound

This paper cites Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.526893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:dd02dab32c1b7b13322204b9695dca95e3d060f4e7de4f277b7238c6271d8989

Observation 451c7b3c-4205-45bc-bd1f-1b7927b3efd1 · outbound

This paper cites Wegen: A unified model for interac- tive multimodal generation as we chat.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Wegen: A unified model for interac- tive multimodal generation as we chat

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.971543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c337586575ab65f95f0ff96541100d75dfa48035e9d0af97b5c162132b396a36

Observation 4250f6db-719a-48d3-938d-93dcd311fe5c · outbound

This paper cites SYNAPSE: Synergistic associative processing & semantic encoding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing SYNAPSE: Synergistic associative processing & semantic encoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T06:36:52.549002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ab106a42b2599a3788afb4fd4e05f7548e38c504ed8db8089624f7f6ba6d75ce

Observation 868cf897-d97a-45b8-b55e-48b8ee1be68b · outbound

This paper cites Videomem: En- hancing ultra-long video understanding via adaptive memory management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videomem: En- hancing ultra-long video understanding via adaptive memory management

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.535630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2d17132ad627b61cc5a81229048afece96885ff0aa0df349b771fdeda58078f7

Observation 364476b8-984f-488a-ba55-9972166d8930 · outbound

This paper cites MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.551375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:9d6259c6f442e974f1a1f047fcdcd4fed463a7f6394172c6d270f7cceaa76d57

Observation afa2fd87-cf78-499d-a3e0-80358efa4eb8 · outbound

This paper cites Camel: Com- municative agents for” mind” exploration of large language model society.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Camel: Com- municative agents for” mind” exploration of large language model society

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.962455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:11f5333142c4cf9d7e81f683a76251b48a9fbfde1aefce27cec6773f34ad1cf5

Observation cb6cc774-ac5c-40ea-bb86-b702ebf7bb22 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.966212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:564a9560e7225d93ae23f2eca61e2ea978760bbf444242cb2b8d386f123869a9

Observation bfa69c8c-ae61-41f7-8130-cf3306aba638 · outbound

This paper cites Iterative trajectory exploration for multi- modal agents.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Iterative trajectory exploration for multi- modal agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.958907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c46b2d410059572c60c81cb22bfb681671c39d4ef37b5cad02411dfe5d657081

Observation 1a600f12-802a-4133-9427-85247b16fb6a · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.960735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:8a3eb7afaa921805483c6f4d8bdfe37d6978eab31f5521f57c16aa48a2117b6e

Observation 86aecfdd-2dbb-4b59-9ad8-8604735bb7df · outbound

This paper cites Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.541423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2ce967898c6a365fb5ad1f599080b843acdfb5790e839d38601d077181a96d94

Observation 86d93bfe-f407-41d6-a0f6-a399cf026589 · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2401369611c04bf5fdff5217f3c4927def90b754c6c1a65e386e11be2866f93d

Observation d28de4e5-65ed-4b91-9306-d7bc82277630 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.956862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:11091bcbc97e52692b83f88b2a471b026dea2209c1c2a4f65872827088631b5b

Observation 4044ca7d-07cc-42da-bdcc-eed259536201 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.546322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c2f50ac7da2a9c7cbd7b9e5fc21256b4f417b99d209aa4c8b350b986778537fa

Observation 9bd478fc-be42-4874-84ef-3c93137fb3e5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.543951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d1b2a543e73307dd3cc24fc2a549eea13256d9c4dced02ec1ec7a686a3ad5f31

Observation f010496c-7246-4375-b471-caec90bc69a6 · outbound

This paper cites Chatdev: Communicative agents for software devel- opment.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chatdev: Communicative agents for software devel- opment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.955062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:343e191b87ee3374376a8d35c6ffd3514e5113ee7d09c9f2ef6a0b06a04d12fe

Observation d8ca0203-eac4-4df3-8090-c8ae668ff738 · outbound

This paper cites quiet" and.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing quiet" and

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.953298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d1f0f3ff562605c14039eca798144b13dbc451f6bbceebf7c1f6dd75bfb38179

Observation 1632f313-2d01-48c0-a98e-7ad050f361fa · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.529778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:6df3996dce1bf912f7647ec2c7d1787e00ce9863247327af9105b338657e1e2c

Observation 3ac3107a-04e5-4f82-8ccd-8cb50a706ec4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.576992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0ed18ad0c36d7a1c43a97cda2732a08caa5af33a4e0ee75bc6d3d0c72c09c62c

Observation d770cabd-c3fc-4b06-a29b-5cbfc367fbcf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.516126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:790b70599af0ce45557cbb4325f491c6fcd0da58aa32137f6ccdc34b69684c87

Observation 659ce5ac-57e5-4001-82c4-708bdb4ff38d · outbound

This paper cites Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.948389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:863acf1b86c0a7ee15206efd5ca9c6af84ffdc3642e3a3274f0e030bf2c56a21

Observation 141f7f3d-8d8e-4a1b-9082-f83562ca4b1d · outbound

This paper cites Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.532440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:af165b5233456dce31b66b98b5acf04454fac5fd29ebda3302fe1ba97fd3ef10

Observation a69fc2c7-550d-4561-9c66-dac08667474c · outbound

This paper cites Multimodal needle in a haystack.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multimodal needle in a haystack

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.950001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:6c0736dd54fbfb51f90833b4e69d949bcebbeebb37da2e28ef7748994371079a

Observation c5b8d679-e97e-4e89-8220-6f2fb9b67d5f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.556292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:f9e5f10feac14fc42bb5bda6f697081091da4724182875f0598ed3bbdb5e81f1

Observation 1fab9c4f-c2c3-49c5-a1d7-b5ad56bbebd2 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emu3: Next-Token Prediction is All You Need

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.574682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2f500addd86e807db73af2aec6440e5dc05e705e43ea53c01e3eddbb9a8f4a08

Observation 39bdbb5b-3444-48b1-a8a7-5554bbce8171 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: Long-form video understanding with large language model as agent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.951712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:550538a944f1da9ad4221d27da2fa566c9fa0bd364d50364afbae42fe7e96a7c

Observation 7ba7a256-fdbf-4092-b012-feb8aa1ee298 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.964348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:71e24ea3422e1eab6e24ac70d4603a9e45c6176842a7c4fe4c2614a9281a4259

Observation 11c699b5-43dc-49f3-b604-801e2d9687d7 · outbound

This paper cites Qwen-image technical report, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-image technical report, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.946383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:30970a7ffcbd163ba2b0f35a2bfd93e10cc5884168c7709b9e26d7a8cca141c7

Observation 7fdc33c5-da4a-458a-b748-ea7865372227 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.561097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:409bf1aa0fd39245fdf5797bd6fd8da9272588e5a91f7953072ca6fb66402bd8

Observation 981b8bcc-a252-40a0-996c-1ef5eb69dac3 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Show-o2: Improved Native Unified Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.521661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1ae858c4a8b545adbc763442410bbe7d644e9f754b1ca3acddd8d0d74f957a46

Observation ebc799a5-f80d-4189-8158-bb72c5c21345 · outbound

This paper cites Qwen3 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.510592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ca24c104351c3ead369ee9dacbad1ec861979598514b01ada5bda8cc405ce3af

Observation 62c9e693-29be-4823-999c-b826b20f61d8 · outbound

This paper cites Agentfold: Long-horizon web agents with proactive context management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agentfold: Long-horizon web agents with proactive context management

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.518800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:24d7a37ef5829c8116219411cd399b1e4c2f3d4526cfb18e7190d22efa485b7d

Observation 7b42fe35-f250-400d-a039-20e52b447e34 · outbound

This paper cites Worldmm: Dynamic multimodal memory agent for long video reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Worldmm: Dynamic multimodal memory agent for long video reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.513341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d50d9085619e2d6bed25cb00e540013fedb8e1e9794cb60105d958187dea85dd

Observation ed3a82d3-2bbf-40c7-928b-a0d057c187df · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.565720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c8585ef14d9bfa3e346cc27faba8ae7aeaceb1e1468ed383aaa0e20c67eba1a9

Observation e9555c9c-9e65-4e8e-beb9-46a671eaf0d0 · outbound

This paper cites Multi-turn consistent image editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-turn consistent image editing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.944652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1f4fee0460311fb71af21e803259092c16f40d81301a88b036ab628ad89f5c71

Observation 8e0fb181-2552-434d-9b5e-cafdf3777fa3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.524292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ccea62f78b0e5f881518e29186be3bb90ac9d8818f06e8f8fbaa2e8dca20d96c

Pith citing papers

No inbound Pith citation observations are available.