Pith. sign in

Paper Citation Record · LEDGER

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

As of 7 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2507.18009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18009 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:43:00.032297Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02b908cb-b9f0-41d1-8dfb-5d4185b77753 · outbound

This paper cites GPT-4 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.986783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.986783Z digest=sha256:886a3c3ed54ffcdfb6dcd490b086493e44940782bf3a55371dbff3f85010cca8

Observation 4959d51a-6a5c-41ad-af12-928e4361ad27 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.993290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.993290Z digest=sha256:0936a1d9c4e8793bd23dc164622b6c6f80f8e73db3a07625c51660b58b0f8690

Observation e4af88cd-e715-43e2-8ae2-81172465c4b6 · outbound

This paper cites GLU Variants Improve Transformer.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures GLU Variants Improve Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.000440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.000440Z digest=sha256:fd31a77ac76b0618e410e333c11ad14dfe1852cdb26469e2c6784c4fe29cc8ef

Observation 366b64a3-709b-4993-abe5-1beac0f64fe9 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.005418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.005418Z digest=sha256:29cb5715bca0a21e48d2dd54091032f28a59d42f42a2e621f5f06c763dae4fac

Observation 2bf0d6a8-50f7-4d32-bdd5-0a79c85994fc · outbound

This paper cites Nemotron-4 340B Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Nemotron-4 340B Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.011505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.011505Z digest=sha256:4fb5c8a59c05e5aed7a4f9683aa576363a424a59a3fb6ea54d1b76149cde33a1

Observation 105069bb-a8f4-4e98-8efc-ce049bef17fa · outbound

This paper cites Qwen2 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Qwen2 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.015406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.015406Z digest=sha256:3b311cc759012eaaeb61df844e0258ae4e9538dcaa7d0c82ecfb402f2fa90eff

Observation e6dfe577-ba5c-448c-aeeb-3d3faf8fc4f0 · outbound

This paper cites an unresolved cited work.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:00.211792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:43:00.018854Z digest=sha256:6d93ed6059c5376c8e9daebc759ec3381de7592e387fdc1d7f99b3f6d51bddd5

Observation dc74ace1-043e-4c9d-bc44-68ec55b530a9 · outbound

This paper cites PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:00.200477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T14:43:00.022201Z digest=sha256:2d3f98456126255c18804d4f4e75f66f5f917f5e72b1782ea4e41ddd5885a736

Observation dc60dfb6-d565-4060-9dee-5f4b50b356c5 · outbound

This paper cites Decoupled Weight Decay Regularization.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.028975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.028975Z digest=sha256:ef6776b267148e5201f20bd6131ebcd28adbb68f0882f156017de0b1c841ded1

Observation 780af361-7dfb-4aec-aae2-d6d3920c797b · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.032297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.032297Z digest=sha256:98aae4b44715d315c62ff89f6cdeb65674046eaae0a02aaaa5bf0c1682e24f3f

Observation 38bce55e-d6ab-4720-8219-927013cdbeb7 · outbound

This paper cites The Llama 3 Herd of Models.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures The Llama 3 Herd of Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.978684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.978684Z digest=sha256:ea532871340b4b4844d5bc8095a5692f6d33b57f67ba198da8122e7eb59ed984

Observation b2d2b7a8-e42a-41f3-9700-33d02ff64959 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Imagenet: A large-scale hierarchical image database

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.025506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.025506Z digest=sha256:8b43a163b8c2308dd1388dae61328502b83fd5cdbea64a2d4f920b8dbd5a234c

Observation 7439bea8-daca-4d88-8499-61210ebfd298 · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.996885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.996885Z digest=sha256:6d338e79c750325e6dc8abf0c20e41e0dec82b4a5abec338fc5fa67681093c79

Observation b405ada9-b939-422c-b40a-30ae82eda205 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.989946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.989946Z digest=sha256:4dc78822361125a4e99739f55101882b8e9e2009683d1550071fc456d5821577

Observation 385e81e0-0776-48e7-a04a-dae1ff7b274f · outbound

This paper cites DeepSeek-V3 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.982607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.982607Z digest=sha256:56b59d008c4331472b3e36c77df3cfd18f39c1e73fd2f777f9929baae69c0c5f

Pith citing papers

No inbound Pith citation observations are available.