Pith. sign in

Paper Citation Record · LEDGER

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

As of 19 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.08497.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08497 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-10T06:30:28.462917Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact26
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 118692d1-fc3c-44b5-b6d3-f4442849222f · outbound

This paper cites GPT-4 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.563355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:23aef29a882c951311b79eaf9097b4c9d5ac7b2bda1b133447928b8fc2b0b655

Observation 11e1e380-6d78-435e-ab1d-989592fae595 · outbound

This paper cites Qwen Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.572324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c81d69687c46f7813d5decbd74575959ac31cd780a837552d889038200cbc423

Observation ee0602c2-0354-4ea2-bb65-7a2d4e0d4002 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.975236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:fba0367326510ec35bcff63aa08b16e1968a87b5d43fb673a899d440134adcb3

Observation e567e0b5-1cc4-4e0d-915f-6656802e901d · outbound

This paper cites Qwen3-VL Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.558779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c81d71682ad352b7da21db233b0d7a51766ac7970f723ed18630250c53be6c28

Observation acc2e0eb-e4c1-4c16-a46a-6db0cfd7c988 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.973493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:fa760f31c9b35d5e2ba716de81ceb40176489e34475418eca430cb3131ff8108

Observation 7a6a227c-33ce-4bf4-9ceb-f9880e313684 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.570188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:62f6ef9f7616dcc7aca5d9c61a2bc7b1f6143f72c30df7c0b056a7732cc2d8c6

Observation 238ab7e4-30dd-4ca3-a8c5-82e808772937 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.567988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:96ca73c11b535a07ce5f79b2b098e9123b50e2b034afa79e3e46325e75d4b4c5

Observation 2a2f523f-2b61-476e-8eee-b7856980268b · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: A memory-augmented multi- modal agent for video understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.969766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ecbce7987499b7f0e0f7da5ba29eb0e145cc31fd2a40cb4fc67e38ec336e922e

Observation 614e779c-392d-4a80-a56f-5936bfe66faf · outbound

This paper cites Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.553796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:1a2972f2a6324b14210894b806e9068c779ce9e416b0fbfadcb3905def843bc2

Observation 3d65c253-8658-4c0e-a091-d89bfa607dbe · outbound

This paper cites Metagpt: Meta programming for a multi-agent collaborative framework.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Metagpt: Meta programming for a multi-agent collaborative framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.968036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:8df071c7afbaa705b0ac4052fad997a599097e451770b1749a784206d02c70b2

Observation 4c0f5849-0085-4fe7-a67c-4d89f32e16da · outbound

This paper cites Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.526893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:5009ce2add23c9622d2f1d4407c3fe23563ca1cf4dc4184bb919cdb8e23fed48

Observation 451c7b3c-4205-45bc-bd1f-1b7927b3efd1 · outbound

This paper cites Wegen: A unified model for interac- tive multimodal generation as we chat.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Wegen: A unified model for interac- tive multimodal generation as we chat

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.971543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:4bdab910a1e69f11a64adf66222a76a6a1510936b9c5fa0d3ba5bdbf7677f1d0

Observation 4250f6db-719a-48d3-938d-93dcd311fe5c · outbound

This paper cites SYNAPSE: Synergistic associative processing & semantic encoding.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing SYNAPSE: Synergistic associative processing & semantic encoding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-10T06:36:52.549002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:3dc4707b6a7fb5f2760b90d45b7fb83fe4490cbd57d85f8595f707746ee16bd4

Observation 868cf897-d97a-45b8-b55e-48b8ee1be68b · outbound

This paper cites Videomem: En- hancing ultra-long video understanding via adaptive memory management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videomem: En- hancing ultra-long video understanding via adaptive memory management

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.535630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:297ade6093c9fe4bfa4335aef3e5e9792664c8d42c9cb669af1d855599205519

Observation 364476b8-984f-488a-ba55-9972166d8930 · outbound

This paper cites MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.551375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:e28af846a260ad5ac8df0298effd1fd3b90d0dff546ecd6216ecb5093308ac94

Observation afa2fd87-cf78-499d-a3e0-80358efa4eb8 · outbound

This paper cites Camel: Com- municative agents for” mind” exploration of large language model society.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Camel: Com- municative agents for” mind” exploration of large language model society

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.962455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:929fd9fc96bb33b7fde806427d537e31312ab7cde8f203be6411376871e5fcdb

Observation cb6cc774-ac5c-40ea-bb86-b702ebf7bb22 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.966212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:01b68bcd2e7d487f06d8c27032e95e5ed07b4f83db25c1fb6c49c4795165a6b9

Observation bfa69c8c-ae61-41f7-8130-cf3306aba638 · outbound

This paper cites Iterative trajectory exploration for multi- modal agents.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Iterative trajectory exploration for multi- modal agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.958907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b6a47c53f2fab495075b00d874bc5360dee6fd4359de5288f3fbbbc976dc55de

Observation 1a600f12-802a-4133-9427-85247b16fb6a · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.960735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:06e326f308af5d41b79b9bd853ada8cf184946447a375e31770ce2918a5c6b6e

Observation 86aecfdd-2dbb-4b59-9ad8-8604735bb7df · outbound

This paper cites Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.541423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:7dc6ed9a51d15ad5b4681b8abedd039b13b1e9e9c3f5d02a066870d181b6965b

Observation 86d93bfe-f407-41d6-a0f6-a399cf026589 · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.538636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:8214d98aad9574219bbcab28b3044e15aa3e6a6bd951b5a56bd262cf08210f30

Observation d28de4e5-65ed-4b91-9306-d7bc82277630 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.956862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:dcd03812df9bcdc008dd1caf73938dc3f3a348e5c8f50165fd57127054af4616

Observation 4044ca7d-07cc-42da-bdcc-eed259536201 · outbound

This paper cites ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.546322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d980023e7e953b5a4d41f0939548f4e7e6f069e5d2affa9e7e9e338fcef117b4

Observation 9bd478fc-be42-4874-84ef-3c93137fb3e5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.543951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:e6383bbf9a7a24dbbe71c8b60dd6ede16b13fd1c56bbc74bc1234d4bce1b23a0

Observation f010496c-7246-4375-b471-caec90bc69a6 · outbound

This paper cites Chatdev: Communicative agents for software devel- opment.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chatdev: Communicative agents for software devel- opment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.955062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0b316f7b74cf84f4959dbac721ea5e1b3590b30e1c372b6bf856de1a314a0517

Observation d8ca0203-eac4-4df3-8090-c8ae668ff738 · outbound

This paper cites quiet" and.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing quiet" and

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.953298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:b93ff560f7c7d67346f6b931672f698d46d0d3ac8fcfea0e4550cce44ab6f33e

Observation 1632f313-2d01-48c0-a98e-7ad050f361fa · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.529778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:05fb51f60c352f29219a3c965e700280e89368917953697e103132176058e262

Observation 3ac3107a-04e5-4f82-8ccd-8cb50a706ec4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.576992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:a26e552aaa80cda00d2d9aa4b35376fa89a2eb2da3b96e7b4f39103acf09adcf

Observation d770cabd-c3fc-4b06-a29b-5cbfc367fbcf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.516126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:f0e2faa0374783290a925a84ed1c8c9f08a0aa938771239b1610e46520f95279

Observation 659ce5ac-57e5-4001-82c4-708bdb4ff38d · outbound

This paper cites Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.948389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:c8a8b50e04a38e6e4889dc6c9b8a89ee75019b46cad71f5482679918a64af178

Observation 141f7f3d-8d8e-4a1b-9082-f83562ca4b1d · outbound

This paper cites Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.532440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:ddf0bc0cc502dd715a61d3d068821fbee5c818a0bfa812ad594f15f6885b711a

Observation a69fc2c7-550d-4561-9c66-dac08667474c · outbound

This paper cites Multimodal needle in a haystack.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multimodal needle in a haystack

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.950001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:bb11e7995a0fb95f2ceac1dc78f21e285fdb45095158743d5235c4e6cab00001

Observation c5b8d679-e97e-4e89-8220-6f2fb9b67d5f · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.556292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:954c79b3a5a6d8265c525627eb298662edfe4d93cf96b8ba35a7a3c0ce365f7b

Observation 1fab9c4f-c2c3-49c5-a1d7-b5ad56bbebd2 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emu3: Next-Token Prediction is All You Need

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.574682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:198c0ea9cb52d1d9b83b9fdd72e441412b1170637d00e087c3563c14b69f521a

Observation 39bdbb5b-3444-48b1-a8a7-5554bbce8171 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: Long-form video understanding with large language model as agent

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.951712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:df0817711358dbc8670105ea949f3d2d90ad69d21c86a35e17609ced9d2ac0f2

Observation 7ba7a256-fdbf-4092-b012-feb8aa1ee298 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.964348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:35e77fa28d32ce03f2cedb69fd57d3d7e8e1ad1c169edd4d91dbbaecc0c28aeb

Observation 11c699b5-43dc-49f3-b604-801e2d9687d7 · outbound

This paper cites Qwen-image technical report, 2025.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-image technical report, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.946383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:30105dad55899be939272451ace9ec3a570c5d1e23bbfd26e5a4af8457d582b7

Observation 7fdc33c5-da4a-458a-b748-ea7865372227 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.561097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:d803f40bf7190ef1195d61967815cb2f7cd6a3691c48c502cd9281bd590f0d18

Observation 981b8bcc-a252-40a0-996c-1ef5eb69dac3 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Show-o2: Improved Native Unified Multimodal Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.521661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:84c14d43fcf80087efbed3fdf04acc1631ef70a41f87fd6d23a4c33b4da5594e

Observation ebc799a5-f80d-4189-8158-bb72c5c21345 · outbound

This paper cites Qwen3 Technical Report.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.510592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:cd8a01165836b2369832a592b3fd970258fe1213ca9a7fb3314499c28133fbf4

Observation 62c9e693-29be-4823-999c-b826b20f61d8 · outbound

This paper cites Agentfold: Long-horizon web agents with proactive context management.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agentfold: Long-horizon web agents with proactive context management

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.518800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:003e2874e8870b6648aa3e79f77602ddaffa4933d1e836c18b07c2a1030c0c21

Observation 7b42fe35-f250-400d-a039-20e52b447e34 · outbound

This paper cites Worldmm: Dynamic multimodal memory agent for long video reasoning.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Worldmm: Dynamic multimodal memory agent for long video reasoning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-10T06:36:52.513341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:7ab65f55c0cfafdf04eb15904136c2097a27debf262b8106b161cecb0466a2c1

Observation ed3a82d3-2bbf-40c7-928b-a0d057c187df · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.565720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:da271a9da673e59414195cbe8ea0d7c7dba3cd7e58c1700d6c825c62679430ba

Observation e9555c9c-9e65-4e8e-beb9-46a671eaf0d0 · outbound

This paper cites Multi-turn consistent image editing.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-turn consistent image editing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-10T06:36:52.944652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:0cc327cbc7aadaaeeff98a21a8537536555909fbb7dbc229fcbfdeed749a2c6f

Observation 8e0fb181-2552-434d-9b5e-cafdf3777fa3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-10T06:36:52.524292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-10T06:30:28.462917Z digest=sha256:2c21581a2773ca0d5213ab97abb0da88cabefa557c5c49381758d85576720933

Pith citing papers

No inbound Pith citation observations are available.