Pith. sign in

Paper Citation Record · LEDGER

DataComp: In search of the next generation of multimodal datasets

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2304.14108.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.14108 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:53.468171Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

74
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 53e4f70b-9858-4a6d-889c-4aa69b2da6f1 · inbound

Objaverse-XL: A Universe of 10M+ 3D Objects cites this paper.

Objaverse-XL: A Universe of 10M+ 3D Objects DataComp: In search of the next generation of multimodal datasets

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:02:11.604349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T13:02:11.512409Z digest=sha256:07119875987657d3f4ce70b2128c95f7c6efc40aa27bc1270a6edca1065b4196

Observation 5b0763c5-fe42-473d-ae4b-031e297c053a · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation DataComp: In search of the next generation of multimodal datasets

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.712246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f5c6ebe78f7dbabbf00cb755cdcea328bc1e7deeee28e4a5ff1c330b8522bc1d

Observation b7853a04-19fe-49b5-bcb6-5fa3b35955f9 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models DataComp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.415728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:572a7658e144b8d8fcdeecaf55a63487e31109321f74749b123934770f628b4b

Observation 50537838-c88e-44dd-add6-32810a4947d5 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration DataComp: In search of the next generation of multimodal datasets

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.844203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:75fb50a22f1badb15049e715743324c7e9802428aeb75c237d379abf8c366298

Observation dbd967c6-4515-481d-8b9a-4f0905432019 · inbound

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions cites this paper.

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions DataComp: In search of the next generation of multimodal datasets

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:08:12.842273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T17:08:12.727773Z digest=sha256:329ebdf8bf612ed2f4380ca9e3b32123dccebf9aef8bf44cf85d58b402a04aeb

Observation 16a27061-3fdc-4a50-b658-94fc2332560b · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models DataComp: In search of the next generation of multimodal datasets

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:16.961906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:5e62d072102db646ded0542d1098c0078987d523312856b1966f75277f466c6f

Observation 26e6ffb9-7845-4953-af59-dcce8109799e · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models DataComp: In search of the next generation of multimodal datasets

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.377313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:60ccc6b01d436d5b1c9f72b74c785ed8fb762d4208e11994ea711b5c5bc226d1

Observation 46aae501-ac6a-4601-a235-b9969bd1f334 · inbound

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images cites this paper.

jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images DataComp: In search of the next generation of multimodal datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T17:37:15.258312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:37:15.258312Z digest=sha256:39c8fc4b292e03d3562d18313edaa865ac2cb4d9b4afdcd086d0f6cc73a5dcad

Observation 4e3bda14-adab-4da1-99bc-5aef939eb0d6 · inbound

FastVLM: Efficient Vision Encoding for Vision Language Models cites this paper.

FastVLM: Efficient Vision Encoding for Vision Language Models DataComp: In search of the next generation of multimodal datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:23.176565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:23.176565Z digest=sha256:e6984b35c59b2e3ea3d9ea0672c3f6cf814d873e4e75671cd663ccb2dea7bc02

Observation fa64068f-00f8-4ec1-990d-c6e4c2d63e7e · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey DataComp: In search of the next generation of multimodal datasets

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.673140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.673140Z digest=sha256:9419a6ecf576c147b8d1b042f905d77c962ff805e90afee3db4626ba96722080

Observation 7dc61a48-3420-4b62-82db-fccc3d0d19ae · inbound

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models cites this paper.

LR0.FM: Low-Res Benchmark and Improving Robustness for Zero-Shot Classification in Foundation Models DataComp: In search of the next generation of multimodal datasets

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T00:12:36.209585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:12:36.209585Z digest=sha256:0b92f5b1b18c1e5d1da3f1cc1beb77c1c2f49c0d4983af5ba2e759036491b0d9

Observation 30ac765f-d39d-43fa-9d45-b57bc71a3f6f · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding DataComp: In search of the next generation of multimodal datasets

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.977529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:558b1de4f7991f9f2181d96bbc531afefc7b5e87f6ef87dd3fbc143bbcf6d9b3

Observation c957d152-6d94-4247-b56b-21a7ecc89e37 · inbound

Position: The Most Expensive Part of an LLM should be its Training Data cites this paper.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.468171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.468171Z digest=sha256:3292f4f7f9c3487639a2ff6ab57e3fbc0662855aa6b9815538e228beb240caa7

Observation 01858a9a-0c1a-4c10-9b5b-b5316d65ceee · inbound

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models cites this paper.

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models DataComp: In search of the next generation of multimodal datasets

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:27.611469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:27.611469Z digest=sha256:dbf2ffb4a193329c056d4faa7ca1391640f3e1b1ec84739c2d70f50d15f8caaa

Observation a37584e8-29ff-4584-abae-b57acb129cfd · inbound

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning cites this paper.

OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning DataComp: In search of the next generation of multimodal datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:29:57.290807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:29:57.290807Z digest=sha256:e3ca254f3420f885f4951f6d80b36f6ed92ee74f594091d3ce11c88b630775a2

Observation 13386499-bd3b-45c4-a6cd-e137f993a7f3 · inbound

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning cites this paper.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DataComp: In search of the next generation of multimodal datasets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.317518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.317518Z digest=sha256:29c69b8becf5983d81a998ab2affcf233c4bdaa98363f89713b7ee53535c07e3

Observation aadfe9c0-a28f-4092-b195-fc2eae3e4cd4 · inbound

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment cites this paper.

Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment DataComp: In search of the next generation of multimodal datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.345850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:32.345850Z digest=sha256:d58d1d18cbf2cf5b6454e6c5f4cda695096961fcd76412b0d28db58b3be1e8cd

Observation e6df2df8-30ed-41f0-b903-f650583f6d3b · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence DataComp: In search of the next generation of multimodal datasets

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.875949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:6fbfffae6809b0345e07f061bd1099b0917ae6f6ca53041b6439c604b18583e3

Observation 38170498-c1f5-4e46-8004-1602482eff3c · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence DataComp: In search of the next generation of multimodal datasets

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.274211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:205f5df50e917523791d988630891b96acb6ef7dd352680f3a5ceffe6e747249

Observation fceae8d0-9b28-4509-a9a7-255461e1da22 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences DataComp: In search of the next generation of multimodal datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:56.445045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:56.445045Z digest=sha256:28dd366e4c56a8760f52eafdec5ad5cd033482e26796de9e8ec52102d797bb4f

Observation 1a413808-640f-485c-a0d7-4112960d33e7 · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems DataComp: In search of the next generation of multimodal datasets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.025685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.025685Z digest=sha256:ef1002d989a334ee616ab2792a708d5c52b0720f6055eb6fc15879ec00681ef2

Observation 28da12cc-c86a-4b77-a6b4-f5ab6e92c7cb · inbound

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models cites this paper.

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models DataComp: In search of the next generation of multimodal datasets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:40.265102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:40.265102Z digest=sha256:abb721948d47768fc0b0120110903653aac3f18d562535ff8b1501eb9b94c989

Observation 12352945-800f-405c-90d6-e72bb1f3bac9 · inbound

Ambient Diffusion Omni: Training Good Models with Bad Data cites this paper.

Ambient Diffusion Omni: Training Good Models with Bad Data DataComp: In search of the next generation of multimodal datasets

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:13.736890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:13.736890Z digest=sha256:84afca940c8368d4bfdf74762b6e658666ccb9cef983a2e3d39a4d4d7a1ae840

Observation 8ca22854-6d9e-43e6-a931-eeb949699ba6 · inbound

CLIP-like Model as a Foundational Density Ratio Estimator cites this paper.

CLIP-like Model as a Foundational Density Ratio Estimator DataComp: In search of the next generation of multimodal datasets

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:46.175499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:46.175499Z digest=sha256:b06cf607bf03bd63ab355f0ac2229a8879d846900fc0f1069f190ddcd71e0761

Observation 6823ab5d-aa99-4459-8da0-55367a9087ae · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training DataComp: In search of the next generation of multimodal datasets

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.528028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.528028Z digest=sha256:b935559d00a1e09f169f891f89e3d2f6114da143dfd65891933a8202ecf6a1a0

Observation 98faa756-30e0-4b35-a4e4-807594854e0b · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning DataComp: In search of the next generation of multimodal datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:33.659440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:33.659440Z digest=sha256:193a37b8f7149406553c5ada5cc2e1cdceb1d3052af6b4914d59e09de283730d

Observation ba5b2e99-95be-4cbb-8b2c-e65ec2e8cb55 · inbound

QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild cites this paper.

QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild DataComp: In search of the next generation of multimodal datasets

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:01:17.173725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T10:56:28.973065Z digest=sha256:12d345179b2a99943007877224f78342b1f4a68fce628d13574cca95de6d0e37

Observation 36a23a66-588a-4566-86fd-137999944328 · inbound

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining cites this paper.

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining DataComp: In search of the next generation of multimodal datasets

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.623943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T20:07:29.544613Z digest=sha256:9b015a15ecd3cadefae4928c9e4a95b994fc5aadf77e5a7c652eae76c51aa0dd

Observation 354b2574-26db-433b-b021-4ab0150da4c1 · inbound

Prior-Aligned Data Cleaning for Tabular Foundation Models cites this paper.

Prior-Aligned Data Cleaning for Tabular Foundation Models DataComp: In search of the next generation of multimodal datasets

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:17.041197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T16:47:31.905149Z digest=sha256:c281cabcce4712cc795fd67d138f61c2c6471676f33b465ecd2bd9758cb2d6fa

Observation 44239207-3e69-40cd-b826-4e2a54fcd34f · inbound

From Cradle to Cloud: A Life Cycle Review of AI's Environmental Footprint cites this paper.

From Cradle to Cloud: A Life Cycle Review of AI's Environmental Footprint DataComp: In search of the next generation of multimodal datasets

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:31:12.927636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T15:43:50.422887Z digest=sha256:216ff5196f9bc789980491646966a94caf38e42271b31d63536f2c68df5b51d2

Observation dd4dc2db-6d56-4ac3-be94-13ed77e75d64 · inbound

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs cites this paper.

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs DataComp: In search of the next generation of multimodal datasets

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.568450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:10:27.595446Z digest=sha256:65215b700a3a3ea997380ecc0346a4a652c935b867e9d1f883c001e13d533b1b

Observation fb96e050-6f1b-4a26-8252-11400aca9b15 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison DataComp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.732465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:b7fd843ceb655e5e60e7d7fd234525bf26586d5c9611b547cc1177746cc51a37

Observation cc80062b-c6f6-46c5-be8d-15b48b4c6989 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison DataComp: In search of the next generation of multimodal datasets

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.323685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:f384e8025f1225699e94e20b8361d2b9b77789ece31b093f7a18b5b8a429a5a1

Observation bffc5f5e-6bfb-4ee6-a4e5-d39898cd29a4 · inbound

GPIC: A Giant Permissive Image Corpus for Visual Generation cites this paper.

GPIC: A Giant Permissive Image Corpus for Visual Generation DataComp: In search of the next generation of multimodal datasets

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:14.139724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T07:36:21.262064Z digest=sha256:dbe1ad201157f0f256bfffe22ecccabcc13533a016fea487279190bf119e842b

Observation 7fab5c0d-92cc-4308-b2b2-1ece77cb957d · inbound

Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory cites this paper.

Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory DataComp: In search of the next generation of multimodal datasets

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:46:55.266557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T03:07:52.730713Z digest=sha256:fad181d989f909fb19685322ac6fae164c23322c024593c0fef37b5493c3b5a9

Observation 8becee99-1461-4517-9165-db60f18cc0b9 · inbound

Instrumented data for causal scientific machine learning cites this paper.

Instrumented data for causal scientific machine learning DataComp: In search of the next generation of multimodal datasets

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.232799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T22:24:28.643956Z digest=sha256:2e84a7d632d5e5a32f563dd9d56d9fd3e297b8a69125532cbbbf3b028889a9b6

Observation 02ffbe7a-5687-4cdd-a96d-6257e73f5924 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation DataComp: In search of the next generation of multimodal datasets

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.855348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:9f0166363923b00934e1adab4c48c3798ec297b9b911129c03e37a599b8effe9

Observation 1ed58a5c-454a-49fe-9e34-bd224feed00c · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All DataComp: In search of the next generation of multimodal datasets

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:29:57.904493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:567adc957d623db9688c739361efe05563376f3eb9b345485054fb53eae64558

Observation 60f42309-0369-402a-a0f7-a289360b21fd · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All DataComp: In search of the next generation of multimodal datasets

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:03:32.239700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:57749084d2bfc3b75735e97843273ad8ca11f6689110d7ac44d628aeb904e3a8

Observation 06c80a0a-6a21-4901-9a71-0ac0cf997d2b · inbound

MIRAGE: Protecting against Malicious Image Editing via False Moderation cites this paper.

MIRAGE: Protecting against Malicious Image Editing via False Moderation DataComp: In search of the next generation of multimodal datasets

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.596740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T01:52:12.291420Z digest=sha256:0263e85f6fc66c77b0887c6c47b8e6b5ce6e2a04475a8011145610dae21fc67c

Observation ad1a47ec-70b8-44bb-9cfa-7ce7b22c1616 · inbound

MIRAGE: Protecting against Malicious Image Editing via False Moderation cites this paper.

MIRAGE: Protecting against Malicious Image Editing via False Moderation DataComp: In search of the next generation of multimodal datasets

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.521084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T04:46:56.552601Z digest=sha256:f32bcd554920b3966a83cc8cb71bb773226076958f8a6a9faebcf6350766d56e

Observation 64f9987d-9b2e-495e-97c9-e375d3f97e63 · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling DataComp: In search of the next generation of multimodal datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6bbac0373efbad664a03ee8ca5919d728813d04340400b478009d92a1e486999

Observation 139999be-369b-4d71-8716-c21c3f3d0e95 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report DataComp: In search of the next generation of multimodal datasets

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:435247cf898199de8c7aa5a02ac6bc3adf7d4e5070a8ecf558455313458e8676

Observation cf286f86-7e59-4e71-9d37-597deed39d0f · inbound

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement cites this paper.

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement DataComp: In search of the next generation of multimodal datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T01:15:07.149546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:15:07.149546Z digest=sha256:66e7da31faae4d5cb1fe9dd9979e35e09e0f85130d1bc9a737384f9df7c08da0

Observation 96d3de4c-1eb1-4273-ba32-125730c3fefc · inbound

Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders cites this paper.

Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders DataComp: In search of the next generation of multimodal datasets

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T11:09:10.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T11:09:10.532529Z digest=sha256:be01698f3a19ac26934229218115e9f3d00a859943aa4e8bd51b94f563a8632f