Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T06:30:28.462917Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2607.08497.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-10T06:30:28.462917Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 118692d1-fc3c-44b5-b6d3-f4442849222f · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11e1e380-6d78-435e-ab1d-989592fae595 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ee0602c2-0354-4ea2-bb65-7a2d4e0d4002 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-vl: A versatile vision-language model for understand- ing, localization, text reading, and beyond
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e567e0b5-1cc4-4e0d-915f-6656802e901d · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation acc2e0eb-e4c1-4c16-a46a-6db0cfd7c988 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing In- structpix2pix: Learning to follow image editing instructions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a6a227c-33ce-4bf4-9ceb-f9880e313684 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 238ab7e4-30dd-4ca3-a8c5-82e808772937 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emerging Properties in Unified Multimodal Pretraining
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2a2f523f-2b61-476e-8eee-b7856980268b · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: A memory-augmented multi- modal agent for video understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 614e779c-392d-4a80-a56f-5936bfe66faf · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d65c253-8658-4c0e-a091-d89bfa607dbe · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Metagpt: Meta programming for a multi-agent collaborative framework
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c0f5849-0085-4fe7-a67c-4d89f32e16da · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Talkphoto: A versatile training-free conversa- tional assistant for intelligent image editing.arXiv preprint arXiv:2601.01915, 2026
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 451c7b3c-4205-45bc-bd1f-1b7927b3efd1 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Wegen: A unified model for interac- tive multimodal generation as we chat
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4250f6db-719a-48d3-938d-93dcd311fe5c · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing SYNAPSE: Synergistic associative processing & semantic encoding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 868cf897-d97a-45b8-b55e-48b8ee1be68b · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videomem: En- hancing ultra-long video understanding via adaptive memory management
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 364476b8-984f-488a-ba55-9972166d8930 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afa2fd87-cf78-499d-a3e0-80358efa4eb8 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Camel: Com- municative agents for” mind” exploration of large language model society
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cb6cc774-ac5c-40ea-bb86-b702ebf7bb22 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bfa69c8c-ae61-41f7-8130-cf3306aba638 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Iterative trajectory exploration for multi- modal agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a600f12-802a-4133-9427-85247b16fb6a · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86aecfdd-2dbb-4b59-9ad8-8604735bb7df · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agent0 -vl: Exploring self -evolving agent for tool -integrated vision -language reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86d93bfe-f407-41d6-a0f6-a399cf026589 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d28de4e5-65ed-4b91-9306-d7bc82277630 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4044ca7d-07cc-42da-bdcc-eed259536201 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bd478fc-be42-4874-84ef-3c93137fb3e5 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f010496c-7246-4375-b471-caec90bc69a6 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Chatdev: Communicative agents for software devel- opment
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d8ca0203-eac4-4df3-8090-c8ae668ff738 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing quiet" and
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1632f313-2d01-48c0-a98e-7ad050f361fa · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Unilip: Adapting clip for unified multimodal understanding, generation and editing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ac3107a-04e5-4f82-8ccd-8cb50a706ec4 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Gemini: A Family of Highly Capable Multimodal Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d770cabd-c3fc-4b06-a29b-5cbfc367fbcf · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing LLaMA: Open and Efficient Foundation Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 659ce5ac-57e5-4001-82c4-708bdb4ff38d · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Crea: A collaborative multi-agent framework for creative content generation with diffusion mod- els
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 141f7f3d-8d8e-4a1b-9082-f83562ca4b1d · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Yanyun-3: Enabling cross-platform strategy game operation with vision-language models.arXiv preprint arXiv:2511.12937, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a69fc2c7-550d-4561-9c66-dac08667474c · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multimodal needle in a haystack
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5b8d679-e97e-4e89-8220-6f2fb9b67d5f · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fab9c4f-c2c3-49c5-a1d7-b5ad56bbebd2 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Emu3: Next-Token Prediction is All You Need
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39bdbb5b-3444-48b1-a8a7-5554bbce8171 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Videoagent: Long-form video understanding with large language model as agent
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7ba7a256-fdbf-4092-b012-feb8aa1ee298 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Genartist: Multimodal llm as an agent for unified image gener- ation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 11c699b5-43dc-49f3-b604-801e2d9687d7 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen-image technical report, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7fdc33c5-da4a-458a-b748-ea7865372227 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 981b8bcc-a252-40a0-996c-1ef5eb69dac3 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Show-o2: Improved Native Unified Multimodal Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ebc799a5-f80d-4189-8158-bb72c5c21345 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Qwen3 Technical Report
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62c9e693-29be-4823-999c-b826b20f61d8 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Agentfold: Long-horizon web agents with proactive context management
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7b42fe35-f250-400d-a039-20e52b447e34 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Worldmm: Dynamic multimodal memory agent for long video reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ed3a82d3-2bbf-40c7-928b-a0d057c187df · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9555c9c-9e65-4e8e-beb9-46a671eaf0d0 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing Multi-turn consistent image editing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e0fb181-2552-434d-9b5e-cafdf3777fa3 · outbound
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.