Pith. sign in

Paper Citation Record · LEDGER

Vision as Unified Multimodal Generation

As of 5 August 2026, this Paper Citation Record lists 100 of 231 outbound references and 1 inbound Pith citation observation for arXiv:2607.06560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06560 v1

Coverage vector

measured 100 of 231 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:54:30.649092Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:39:37.738395Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 231 outbound references displayed

  • verified exact29
  • verified fuzzy59
  • unresolved5
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14329a86-1436-4cd2-b60e-bc86ee11fd5f · outbound

This paper cites Dataone-synthetic-v1.0-sample, 2025.

Vision as Unified Multimodal Generation Dataone-synthetic-v1.0-sample, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.142232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:05392c625084214753c1a61e566351cca98bdb1020a10e7c28990e3649bdc16e

Observation 0f379d01-c37c-406f-84a1-f41c0a1e8773 · outbound

This paper cites Ttpla: An aerial-image dataset for detection and segmentation of transmission towers and power lines.

Vision as Unified Multimodal Generation Ttpla: An aerial-image dataset for detection and segmentation of transmission towers and power lines

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.120057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a13f3430bdd0ea1f3f8ed118d5f88124585343564b19e63d13b8eb9c988d60c2

Observation 149c2d2c-e268-42df-9bbf-d7d30721f8c2 · outbound

This paper cites Matting human datasets.

Vision as Unified Multimodal Generation Matting human datasets

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.091082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:f8677e70970207155358f51ab0900d764961d1ac88477266c0c21282d2a58f20

Observation 5c910035-4b5a-43ac-a8b7-a9e65fcf7bdd · outbound

This paper cites Flamingo: A visual language model for few-shot learning.

Vision as Unified Multimodal Generation Flamingo: A visual language model for few-shot learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.113959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:93a2bfcf353ba255df9b8bb64d3558aec5c9a67580fbbf98565c913c35403035

Observation 6b4dee0e-baea-4063-ac71-e8c5f0766b80 · outbound

This paper cites IDDA: A large-scale multi-domain dataset for autonomous driving.

Vision as Unified Multimodal Generation IDDA: A large-scale multi-domain dataset for autonomous driving

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.041635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:8e232225ffbec9aa72a6acb9659caf2717a4d23d85b914dbe33edffc2baba54e

Observation 1fb8eb27-9844-4836-88d9-7c4cc7c374f1 · outbound

This paper cites Open-world text-specified object counting,.

Vision as Unified Multimodal Generation Open-world text-specified object counting,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.132260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:feec746a5ec02c4e0a913be811a215562a6d61274b282d8c7d3854a26fa69709

Observation fae6a27b-5208-4aaa-8fd6-16ceb9f082c9 · outbound

This paper cites Open-world Text-specified Object Counting.

Vision as Unified Multimodal Generation Open-world Text-specified Object Counting

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.373912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:1f1e2aa1a957670b9f40d5baeff03c5f5a887ad180710a459a403417bf174fe7

Observation 28634544-7d04-4fb9-9d07-98d9605dd16d · outbound

This paper cites 2d human pose estimation: New benchmark and state of the art analysis.

Vision as Unified Multimodal Generation 2d human pose estimation: New benchmark and state of the art analysis

Reference 8

Resolution
verified exact
doi, observed 2026-07-08T02:04:26.040736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:68d08955c10c4719c00c4a79b94fc755471a5d8582478269c96d61ce544f6a07

Observation fec34e38-0e46-405d-a881-ff6c68e85ed5 · outbound

This paper cites SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model.

Vision as Unified Multimodal Generation SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.375455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:37f96905af843fcd1980caad363f36f5ca33c41e61c466cab44ace9df268900e

Observation 63184c61-cfac-4a3c-93b2-8836e5c74007 · outbound

This paper cites an unresolved cited work.

Vision as Unified Multimodal Generation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-07-08T02:04:27.099221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:f2e8c26176711ab3ccbe07798b5fbb7a865e37ae2aaae22bb3d8a89b31caaa40

Observation be3d0ff4-98df-47ad-82fc-fe3cf664aab1 · outbound

This paper cites Qwen3-VL Technical Report.

Vision as Unified Multimodal Generation Qwen3-VL Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.433787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:7d082cd3a9ea035bba813343544daf571018ec4b7d9e7d6af3b0b54f7d93b9ca

Observation c43e6536-ceac-4a6a-b059-ee513b2c506c · outbound

This paper cites Zerowaste dataset: Towards deformable object segmentation in cluttered scenes.

Vision as Unified Multimodal Generation Zerowaste dataset: Towards deformable object segmentation in cluttered scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.702708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:1aaf00a36bcea163bfebe9aa702b544125f496d51e0aa3940cdf1fa11e145332

Observation 390795ff-a94f-4104-8af0-5f39ffec0b03 · outbound

This paper cites an unresolved cited work.

Vision as Unified Multimodal Generation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-07-08T02:14:26.702881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:1ca94832c44c48faad9941158c35f734258e23fbac2006ef937e1ddbdd0a164e

Observation 578c729c-c99f-4981-ad3b-a2a0a9d5b220 · outbound

This paper cites CDLA: A chinese document layout analysis dataset.

Vision as Unified Multimodal Generation CDLA: A chinese document layout analysis dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.689995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:099b9d71bc81ba5d8237b90ef3fd9fb08af199fe09a2f28f8498e525b6718f64

Observation 872e9d5b-71a4-418d-b45d-4c1999c2eb8b · outbound

This paper cites The 2019 davis challenge on vos: Unsupervised multi-object segmentation.

Vision as Unified Multimodal Generation The 2019 davis challenge on vos: Unsupervised multi-object segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.697432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:20684ce14cd617a12ea8b08a496415b0be75e818801b061fca4f55ef8c396c03

Observation f5531b8d-d21b-43cc-a91d-ecbcd87ae490 · outbound

This paper cites Rethinking object detection in retail stores.

Vision as Unified Multimodal Generation Rethinking object detection in retail stores

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.700974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:b0250515e0df3ac79665d10f3f59a6e64dd0a5ec9af4d6063eb9506778315ede

Observation 33583841-cdc0-4c88-8d40-79e876d06be2 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

Vision as Unified Multimodal Generation Scaling spatial intelligence with multimodal foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.693711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:d568aa63b8a784a7c2c2c003f36e2752ef79b63b4fa14d91b1b94a473bfe7757

Observation f9b3e33c-b0d9-4ed0-a6f7-78f1ff5a49c0 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Vision as Unified Multimodal Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:25.987284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e0e4e903f9f5ee00bc0d449c3c83e01306446653a4bffb2b5adbc0d5e290d0cc

Observation 2e0019e9-6e56-4974-8fdb-0d3aa5c3cda9 · outbound

This paper cites Industrial-site-safety-detection-v1-dataset.

Vision as Unified Multimodal Generation Industrial-site-safety-detection-v1-dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.692172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:4420c81283b239825c02445390bedfa0ceead2b03b79c9a86ff8e62f9e69dfc7

Observation fe29178b-76b0-4ea7-95d9-189c0c4fe96d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Vision as Unified Multimodal Generation Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.428515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:9187c02c794b3858aeb7644b639b797012eab6ab759d12ecfb4ee2d81b0674cd

Observation b93e2bd0-317e-4f1e-a7be-712c75f321f8 · outbound

This paper cites Fleet, and Geoffrey Hinton.

Vision as Unified Multimodal Generation Fleet, and Geoffrey Hinton

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.686607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:c51bd3d0f862211a5525bfa72c444494d1807627238de8b52774053f5460b9e0

Observation 49bbeccd-a0fc-4500-b1cb-9273644bb7a3 · outbound

This paper cites an unresolved cited work.

Vision as Unified Multimodal Generation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-07-08T02:14:26.701135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:000798bb4e78efa1517c430507b9fdd8cea6a54465af2c48cadb15fb4dd683c3

Observation dbdab1f4-633c-461f-a69a-01dc2ad692d0 · outbound

This paper cites Large-scale structure from motion with semantic constraints of aerial images.

Vision as Unified Multimodal Generation Large-scale structure from motion with semantic constraints of aerial images

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.699233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:bda51e752474c74295665973bcf07b792337adc4d482f6b7ad866e1fc5896796

Observation 18cbea02-7b04-4548-8d03-99a2b8ee1966 · outbound

This paper cites Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vision as Unified Multimodal Generation Intern vl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.706562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:157c85e2b18377fd41ad8b3931ebba8bbf116e07cdb0d7be5094c1f6ea894885

Observation e7548657-eb4f-494b-87ea-105884b49fba · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Vision as Unified Multimodal Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.708408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e910d172bee29c47b6bd759db03e7deb55c3798fe6ce9a48333103ba06b14fbd

Observation e7ec7402-0896-463f-b38b-6a71f328799e · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

Vision as Unified Multimodal Generation Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.686271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:9f623e6c785b41c8092fc1eab255345c29c1280abbff5692d47a965f29dd44e8

Observation f7601215-6e08-4693-a61e-f4effe6e0e0b · outbound

This paper cites Domain adaptation for traffic density estimation.

Vision as Unified Multimodal Generation Domain adaptation for traffic density estimation

Reference 28

Resolution
verified exact
doi, observed 2026-07-08T02:04:26.007256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:c47e9c4c40acb8b67fe6b1e894fc3108ba14371bed52bebe3dcc1e528a5ced6d

Observation 45098d0c-0eac-4f31-892c-97322c6af01d · outbound

This paper cites The cityscapes dataset.

Vision as Unified Multimodal Generation The cityscapes dataset

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.704342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a84f131ea0a1e2838c0d4e546c9e7ed1d4378bc676f785031c7fb91a273e58e8

Observation c18d2d5b-d843-4915-891f-0d2f98754520 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Vision as Unified Multimodal Generation The cityscapes dataset for semantic urban scene understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.694108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:c539ef8311813a1f11c5972d5ff06d1981517679fab7cb00855e32b6c77591e3

Observation 07ec1fb6-7fd6-4f81-8812-a23afa29a50e · outbound

This paper cites ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes.

Vision as Unified Multimodal Generation ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:04:26.436129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:872771401f5b1a41cca00088a02190bc07c0fd938b880db7332eed85a67a5a3a

Observation 5c73ca71-cb9e-474a-8c8a-73e9975931d8 · outbound

This paper cites Objaverse: A Universe of Annotated 3D Objects.

Vision as Unified Multimodal Generation Objaverse: A Universe of Annotated 3D Objects

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.303666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:8651f106c4929f5b4cbba41ec8d47ccc5126820916a138979d51dfe02f759372

Observation eb15dd6d-103b-46fb-8167-286a773cda28 · outbound

This paper cites Smith, Hanna Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Vision as Unified Multimodal Generation Smith, Hanna Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.082212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:8a18aca0d7d945b95f78bca6ac1768c05c7c87fa86a90a6c56e1bcc3f4ab5013

Observation 206f9ac7-7a21-48b2-8137-3e761ceba446 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Vision as Unified Multimodal Generation Emerging Properties in Unified Multimodal Pretraining

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.301186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:b89a8d277f06798dd5529b2b348c843f93cd2b25ae0cd4cbbd53560fcd94da35

Observation 2af2eccf-9a14-4dec-bb15-fec01c0c8e58 · outbound

This paper cites Coconut: Modernizing coco segmentation.

Vision as Unified Multimodal Generation Coconut: Modernizing coco segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.044055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:6eb97470d93bbc83b5d1d18dacdd10e442d3a855c2c633596637107bc3837ed6

Observation fc9d60c5-a1f8-4468-a9f1-a5ba17bf8b9e · outbound

This paper cites SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture.

Vision as Unified Multimodal Generation SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.364613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a51418ee99be0ee16b1fd8479fcf80621404978735a9ad2dbc05c5a09dc08b73

Observation 8645e850-cdc1-401c-9730-56547636fde9 · outbound

This paper cites Object detection in aerial images: A large-scale benchmark and challenges.

Vision as Unified Multimodal Generation Object detection in aerial images: A large-scale benchmark and challenges

Reference 37

Resolution
malformed identifier
arxiv_id, observed 2026-07-08T02:04:26.385721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:35cf4c7a5735c99493325bb39b936a178548783031b8aa3533d4edc436246c78

Observation a5453fca-0f35-48a7-8af3-59c426c60ed2 · outbound

This paper cites Visdrone-det2019: The vision meets drone object detection in image challenge results.

Vision as Unified Multimodal Generation Visdrone-det2019: The vision meets drone object detection in image challenge results

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.138374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:6539e71ab1a76bef275b7e56f2dd6f3d0383b2e645b21ab23f208c7b8fb12eec

Observation 8759c6f8-ae54-46de-93d4-1fe9137305b3 · outbound

This paper cites MoRe: Motion-aware feed-forward 4d reconstruction transformer.

Vision as Unified Multimodal Generation MoRe: Motion-aware feed-forward 4d reconstruction transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.624196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:383279df742f483f13978251875bdf90cfaba1926574b21b1da1f967f13aadad

Observation f17162fc-0190-4ba3-8c69-501b72c73dc0 · outbound

This paper cites Mvtec d2s: Densely segmented supermarket dataset.

Vision as Unified Multimodal Generation Mvtec d2s: Densely segmented supermarket dataset

Reference 40

Resolution
verified exact
doi, observed 2026-07-08T02:04:26.031869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:d36efd1cdeefb9b3f92e19e8e33639a8b36fe398660c91033de31a93e2dbf8bf

Observation 9a92b372-58d7-412a-8ca1-9dd2099a1132 · outbound

This paper cites Panoptic nuScenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking.

Vision as Unified Multimodal Generation Panoptic nuScenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.353443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:6d884b678bed8bb00d158715936fb5c49cd1a1d8caeba893bf9eb8cca54d2db1

Observation 77621a5c-c86c-4207-b700-6f4febb8225b · outbound

This paper cites nuscenes revisited: Progress and challenges in autonomous driving.

Vision as Unified Multimodal Generation nuscenes revisited: Progress and challenges in autonomous driving

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.340292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:0f05d22fc60120019bc0305dd9aaa16a3470e9541b962cdeaab1a237ffa8971b

Observation 91402eb4-a29e-4bab-a710-22842666e0c6 · outbound

This paper cites Image Generators are Generalist Vision Learners.

Vision as Unified Multimodal Generation Image Generators are Generalist Vision Learners

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.419598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:c142de5c2a4714735bebd466b02641a113982596d412127a62cc1e053ace2984

Observation e96c58c5-935f-4380-8b57-cc192b68fee3 · outbound

This paper cites Virtual worlds as proxy for multi-object tracking analysis,.

Vision as Unified Multimodal Generation Virtual worlds as proxy for multi-object tracking analysis,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.708685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:18846dbc29bc5cc55f411a19a624af2eb43c91e2b011f08c3ca3a8d18e78e9c2

Observation e92083f0-082c-4159-809e-b61c4e7bbe7c · outbound

This paper cites Virtual Worlds as Proxy for Multi-Object Tracking Analysis.

Vision as Unified Multimodal Generation Virtual Worlds as Proxy for Multi-Object Tracking Analysis

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.425342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a8a90f832a6d12a2b21eb017b1d44e452e62c453cf8dc97475458e941291b719

Observation 97f9b6d1-81cb-45f5-ad05-90c4e27d1504 · outbound

This paper cites Visual bridge: Universal visual perception representations generating.

Vision as Unified Multimodal Generation Visual bridge: Universal visual perception representations generating

Reference 46

Resolution
verified exact
doi, observed 2026-07-08T02:04:26.075870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a0dca89e3cb4069d75dde66ee93c2b9154d0be565dcb9c3dee055525a311981f

Observation cefa197d-09ec-401d-91a3-d31a3f813bb7 · outbound

This paper cites Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations.

Vision as Unified Multimodal Generation Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.306295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e91199e67b92f32306d2af175e2a2737e60658080e3b6657e1d000a63a1b771b

Observation e095b4e9-36b8-49ba-866c-e1bf4ed5077b · outbound

This paper cites A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images.

Vision as Unified Multimodal Generation A versatile benchmark for detection, pose estimation, segmentation and re-identification of clothing images

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.057515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:dde4a9646612cd9993db8d4cb3c815ffa4066bfdade81f5e32ed6caecfcf8cb3

Observation 88c0661d-2770-4749-8c92-b550a01dfa7d · outbound

This paper cites Vision meets robotics: The KITTI dataset.

Vision as Unified Multimodal Generation Vision meets robotics: The KITTI dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.089280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:8ac5501c7a1747a0e6f22c0db5dc2fa66078e2d844a6da0f4ade2be6aa0b2092

Observation 5a97b4c7-56b1-41d9-8765-2dc420a50575 · outbound

This paper cites GenEval: An object-focused framework for evaluating text-to- image alignment.

Vision as Unified Multimodal Generation GenEval: An object-focused framework for evaluating text-to- image alignment

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.067418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:31051c23868dab81279e972b77d88f932285c324409d67190a02b3773273ae50

Observation f08b6a3a-6663-450e-a6fb-a3bfbd8f7340 · outbound

This paper cites Precise detection in densely packed scenes.

Vision as Unified Multimodal Generation Precise detection in densely packed scenes

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.085769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:d1dcd202c038ad831ee5e43054fe07a7499091633e387e26f616c2b00226a42b

Observation d1ad69f4-5c69-4349-8258-95dad91400a9 · outbound

This paper cites Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing.

Vision as Unified Multimodal Generation Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.084015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:78f0fcb03ab442a8d29fbfe8cb396e5f8c20755bcc968826c344780cdd943030

Observation 72d53621-0981-4ff4-a260-eb87b6e2b67e · outbound

This paper cites Instance-level human parsing via part grouping network.

Vision as Unified Multimodal Generation Instance-level human parsing via part grouping network

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.060503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:35909489252b0d4732d472bd15140f71ac4249c8d45a3f54919f28d1fd8dacdc

Observation eb240ce4-dd34-4035-bf62-b07c17e33431 · outbound

This paper cites LVIS: A dataset for large vocabulary instance segmentation.

Vision as Unified Multimodal Generation LVIS: A dataset for large vocabulary instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.028249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:14aac64c965641d915e5e3456b2f68da3d995d39105d40efbfb4dd2c75689f3d

Observation 40c041f1-e2f4-4f02-93a0-6372f678b722 · outbound

This paper cites Synthetic data for text localisation in natural images.

Vision as Unified Multimodal Generation Synthetic data for text localisation in natural images

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.005728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:4d9f71a56eb2bb1f45711a2b828cf908a430b1f48613f6e3ac4ae37573f93a24

Observation fa3296e9-e6d1-4c61-a58e-3b39c3c6c154 · outbound

This paper cites MinneApple: A Benchmark Dataset for Apple Detection and Segmentation.IEEE Robotics and Automation Letters, 5(2):852–858.

Vision as Unified Multimodal Generation MinneApple: A Benchmark Dataset for Apple Detection and Segmentation.IEEE Robotics and Automation Letters, 5(2):852–858

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.075842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:d3f1d3dbca58f5c5bb6b16624fc804cd199cf7cae7f0a94dad4648333bf534f9

Observation 4bc79d60-2b56-4ac4-be46-df1fa376dc8b · outbound

This paper cites Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model.

Vision as Unified Multimodal Generation Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:04:26.377903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:ba821bd5fcfb50694f60af7d04a532d694c46e3d4e768285a222fd32fb16af7c

Observation 3128bc55-143c-45ad-9479-1ec153480ef9 · outbound

This paper cites Lotus: Diffusion-based visual foundation model for high-quality dense prediction.

Vision as Unified Multimodal Generation Lotus: Diffusion-based visual foundation model for high-quality dense prediction

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.643819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:91a39308db0b594717c7a49b06399b40c64d7bb55008b7db81c5c034bae417f2

Observation 99a0fb86-acbd-4a19-975e-0520e22fbe92 · outbound

This paper cites PartImageNet: A Large, High-Quality Dataset of Parts.

Vision as Unified Multimodal Generation PartImageNet: A Large, High-Quality Dataset of Parts

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.388942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e40635284559bfe66cf0a78de8bd490e542f425e763ef3da47018abbeebc9892

Observation 862e2637-35dc-4807-b733-54f0949e4576 · outbound

This paper cites Partimagenet: A large, high-quality dataset of parts.

Vision as Unified Multimodal Generation Partimagenet: A large, high-quality dataset of parts

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.112008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:70a39d7e869036ff6c9a823f7b6d4befd0a841e050a020845a892b317e5c1eaf

Observation e881f11f-8639-4fe1-98cf-36236f41122f · outbound

This paper cites Mask R-CNN.

Vision as Unified Multimodal Generation Mask R-CNN

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.668338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:aa84982d32e15747d278a4b9ea61f603c5615e25a2550ad9d5f38c27aa93c555

Observation 16d2dfac-bdbf-4f01-93bf-26727d7ba2de · outbound

This paper cites Masked autoencoders are scalable vision learners.

Vision as Unified Multimodal Generation Masked autoencoders are scalable vision learners

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.117270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:0e350e45d0e5ed9bd25b05bbfb0b9f2c356b5425acf07fb44bcc27b50b120bc2

Observation d04fc53f-6234-4d9f-a055-b34ad9fec609 · outbound

This paper cites Icpr2018 contest on robust reading for multi-type web images.

Vision as Unified Multimodal Generation Icpr2018 contest on robust reading for multi-type web images

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.044821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:18f6f55b16d9ad9064f1dffe513ab7b748c4140e9bf3fa5c1ae039bc7183b638

Observation 358a424d-c711-4c7a-b80c-e7bd93b062d7 · outbound

This paper cites Scaling out-of-distribution detection for real-world settings.

Vision as Unified Multimodal Generation Scaling out-of-distribution detection for real-world settings

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.639572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:baf4877804fa56901d4a5478e2605ecb17f9d8aaace881269be7337219fbe58b

Observation 3d017355-2ba4-49c2-898f-78bc3c125515 · outbound

This paper cites Lvis fruits and vegetables.

Vision as Unified Multimodal Generation Lvis fruits and vegetables

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.628069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:4ff991bc3535b486a3b8631328846c5be9052c0ecb48a8674c57cca115ed33ff

Observation e5438663-a401-4db8-9c7b-8b5398122582 · outbound

This paper cites Denoising diffusion probabilistic models.

Vision as Unified Multimodal Generation Denoising diffusion probabilistic models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.114760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:60b45c51d5ef395fe0ff5fe2eefae00c8b25f0891cf0da3327d8ae7fed2a1997

Observation 75391a31-7bdd-4226-8276-38f7377ad0eb · outbound

This paper cites TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris.

Vision as Unified Multimodal Generation TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.405271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:efd6458a8cc0d0dfe56f32756b74c5912886dbf60293c89c0bf43b07a1f7ca1d

Observation 1e6f2d2a-9387-4d69-b88f-bf03758dce82 · outbound

This paper cites an unresolved cited work.

Vision as Unified Multimodal Generation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-07-08T02:04:27.124011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:f148c9b5e3b8af1b358ad54ab8e852eb7c4322fa2a6828b9f481967434961c82

Observation 7e52e22a-660c-4de5-a90e-a80278efab66 · outbound

This paper cites G 2VLM: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.

Vision as Unified Multimodal Generation G 2VLM: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.093203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:24c3deeadc91dfc2515c7737a48897aaaf1418ea32aa7c0f3c3f1aef49f0db5e

Observation 77dee460-54a9-484d-a851-f4805a40c75c · outbound

This paper cites Deepmvs: Learning multi-view stereopsis.

Vision as Unified Multimodal Generation Deepmvs: Learning multi-view stereopsis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.016009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:77fdf58c7c9f06266c44ee40952aff44d2adb787ec9e4aa6b8e54484e31150c0

Observation b5549b05-b2b1-4b9b-bc88-9a3f8a640b45 · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

Vision as Unified Multimodal Generation Semantic image synthesis with spatially-adaptive normalization

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:04:26.054930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:243d02584635ac6f80ef4696e8c9f056a3715967b97e3064e64fc0886b8d9984

Observation a8c9f7b0-0306-4871-be61-c30c980cc2ca · outbound

This paper cites Semantic Segmentation of Underwater Imagery: Dataset and Benchmark.

Vision as Unified Multimodal Generation Semantic Segmentation of Underwater Imagery: Dataset and Benchmark

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.118162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:34780e1339d3a7eae961957660c3a87f3761c553cfc4538a18eef69e40e87127

Observation c8d3d9f0-0888-4759-8f98-c67a3d536e03 · outbound

This paper cites 2016.280.

Vision as Unified Multimodal Generation 2016.280

Reference 73

Resolution
malformed identifier
doi_truncated, observed 2026-07-08T02:04:26.036466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:7fe2c8b024cdbb1f7f639b5356323aa8c97ec1d3c80dfbfdddc2766fe91cc405

Observation c5f07d0f-4837-4124-8952-06876dbc2227 · outbound

This paper cites Communications of the ACM 65(1), 99–106 (2021) https://doi.org/ 10.1007/978-3-030-58452-8 24.

Vision as Unified Multimodal Generation Communications of the ACM 65(1), 99–106 (2021) https://doi.org/ 10.1007/978-3-030-58452-8 24

Reference 74

Resolution
metadata mismatch
doi, observed 2026-07-08T02:04:26.034109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:dc593a1e7b9aeaab0b72c3fdec13708b47ffa12d925889777b15bf3a52c05827

Observation b5777c8a-499d-4155-b36f-cddd5cfe3f2b · outbound

This paper cites Megasynth: Scaling up 3d scene reconstruction with synthesized data,.

Vision as Unified Multimodal Generation Megasynth: Scaling up 3d scene reconstruction with synthesized data,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.101611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:7f88a8926654ec7993cfc883ae2c3b138aeefeb48e34b0f0f263985f529c4cbe

Observation 421d11e2-10f3-40a8-8811-24f9ccc84ee4 · outbound

This paper cites MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data.

Vision as Unified Multimodal Generation MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.350257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:542006446630d81016c5f16c209df11b79e8a44ccd2f55f784ab3389d91fa440

Observation bd0898ff-794f-494b-91c6-da4a1dec1311 · outbound

This paper cites ChatRex: Taming Multimodal LLM for Joint Perception and Understanding.

Vision as Unified Multimodal Generation ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.361272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e32759ffe0e5506cf0ac8adae0d94340a37de3770bbaf93d5f964102e1cb8f3a

Observation 2d97a606-4e5b-4792-8af5-7aceaa92d4a2 · outbound

This paper cites Detect anything via next point prediction.

Vision as Unified Multimodal Generation Detect anything via next point prediction

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.704588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:174815a359d3fdbe617d75e98a9c6e3658ea811807fdf33bcdad4913fd08e2a3

Observation 093db996-1565-4ea3-be6e-bffac95dcd92 · outbound

This paper cites RaidaR: A Rich Annotated Image Dataset of Rainy Street Scenes.

Vision as Unified Multimodal Generation RaidaR: A Rich Annotated Image Dataset of Rainy Street Scenes

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.449387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a129f4beaab2ffd32387f3652670574db544e95029d2e76c6b90d1c0b725765a

Observation 245d09a2-bcab-49cc-821c-849248a7bce2 · outbound

This paper cites Ultralytics datasets: Medical-pills detection dataset, Dec 2024.

Vision as Unified Multimodal Generation Ultralytics datasets: Medical-pills detection dataset, Dec 2024

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.684508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a5f68ab208fe5a46c2da3d68a71290786cce54852a064ae58c518c82b3384d53

Observation 7d610b34-3d97-4f9d-9fba-96d63b3d567e · outbound

This paper cites Ultralytics datasets: Homeobjects-3k detection dataset, May 2025.

Vision as Unified Multimodal Generation Ultralytics datasets: Homeobjects-3k detection dataset, May 2025

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.048239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:8884b51aac8392599acfaacdf0c303fff1de1d9f393c2ef5c7d4bd44ec25a97d

Observation b74be24a-f932-4d7e-8d37-5d11d7082ffc · outbound

This paper cites Human-art: A versatile human-centric dataset bridging natural and artificial scenes.

Vision as Unified Multimodal Generation Human-art: A versatile human-centric dataset bridging natural and artificial scenes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:14:26.695974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:02af74c1704dd76f79aa3b3b9d7e8962e1292f508a5cb7b3d1bb3dfb42c6a10d

Observation a079404a-0c3b-4bd7-a37e-0503128c7408 · outbound

This paper cites Cloppet, V.

Vision as Unified Multimodal Generation Cloppet, V

Reference 83

Resolution
malformed identifier
doi_truncated, observed 2026-07-08T02:04:26.020922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:bdf33170c737f0bda5646cd5b3228aa927d2dce6efd118392785093c71f9f0a9

Observation 376882d1-4e31-48a6-a018-e8a123872785 · outbound

This paper cites Icdar 2015 competition on robust reading.

Vision as Unified Multimodal Generation Icdar 2015 competition on robust reading

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.072891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:894d99479d6d681d9404608e347bd981842d9b0f4d37b6ed6da4a863d1d748fe

Observation 99111827-049d-493e-a7d7-40bf5e3e6797 · outbound

This paper cites ReferItGame: Referring to Objects in Photographs of Natural Scenes.

Vision as Unified Multimodal Generation ReferItGame: Referring to Objects in Photographs of Natural Scenes

Reference 85

Resolution
verified exact
doi, observed 2026-07-08T02:04:26.069720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:093046720f3845e2ba46d46cf25739e826addf3e52b8a22691e38363a4a94b52

Observation e367475d-0a3b-4948-b037-44a32bedf1e8 · outbound

This paper cites an unresolved cited work.

Vision as Unified Multimodal Generation Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-07-08T02:04:27.053696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e44b25e66504b0e7c2d3cba6002dac5c02a8bbf8c33636f2375eeb45582c6867

Observation 85b8f798-50d4-43b2-80a4-a8aa9b26ee94 · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

Vision as Unified Multimodal Generation Repurposing diffusion-based image generators for monocular depth estimation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:26.996543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:9ef53ec0af390e3d8e183d54e5fcad23513c68b18fefd08243f1d8e1d4842407

Observation f6f83ed6-b653-4342-b220-448780083cc6 · outbound

This paper cites MapAnything: Universal feed-forward metric 3D reconstruction.

Vision as Unified Multimodal Generation MapAnything: Universal feed-forward metric 3D reconstruction

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.030764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:4a2ccfec060c0e448780154463fe61b463aaef2d7aacb7ef6ed719433c092b1c

Observation affa4c61-cce0-4945-9e16-831b4ba968cb · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick.

Vision as Unified Multimodal Generation Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.089415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:2ff8d778c16f5bb096b39619bc1bd14bea1e6c734c7844a0370fbd85e457470e

Observation 7bf58d39-baed-4408-9a4d-763284091221 · outbound

This paper cites Waterovs.

Vision as Unified Multimodal Generation Waterovs

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.071524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:d06201702d59547d33dd4ccf027e65fcab3a982b9c6f797d299ff79aff6224cd

Observation a8366567-8341-4891-929b-6f443c4c8bf9 · outbound

This paper cites A hierarchical grocery store image dataset with visual and semantic labels.

Vision as Unified Multimodal Generation A hierarchical grocery store image dataset with visual and semantic labels

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.073748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:4384c27ceff37abb0b36d6f5deec4ab8eba6b5f10de8d461fcb92b7665555d3c

Observation c451a99c-23e0-4dd8-b5e8-316733be0140 · outbound

This paper cites Evaluation of CNN-based single-image depth estimation methods.

Vision as Unified Multimodal Generation Evaluation of CNN-based single-image depth estimation methods

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.078159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:1a6f4e4f170c7a92a2cc7b00ac9c61ea9d2673d543542c54fd4b2e688c58bd70

Observation 33ed5c08-7a36-4cf8-a92c-0c729bc634a1 · outbound

This paper cites shoe-dataset.

Vision as Unified Multimodal Generation shoe-dataset

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.085588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:96fcdb2fac2e720ed8464efbc98fce2982f1e46dc6bed304607422d802a4430a

Observation 59876151-82b1-47ad-8fa5-31720e281e41 · outbound

This paper cites Alet (automated labeling of equipment and tools): A dataset for tool detection and human worker safety detection.

Vision as Unified Multimodal Generation Alet (automated labeling of equipment and tools): A dataset for tool detection and human worker safety detection

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.087628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:74435bace07cff7a174b68a24effcc3fd1d65d6643689d6a5c4cde447be8bc00

Observation d22093b2-c380-4690-96b8-646fcd3b29b8 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Vision as Unified Multimodal Generation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.067549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:acc9ace309866932dd102a10982a2a373185de7641faa46b1ee9c01e986276ff

Observation 53d178ea-1d2e-4b0c-9228-f410689b5951 · outbound

This paper cites in the wild.

Vision as Unified Multimodal Generation in the wild

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:04:26.079301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:3e2921f14a43bd2095cd8d8b1b4f6bcbc738a047d63899a21abac966edbad541

Observation 4ced6689-14ef-408d-991b-e00a87385d3f · outbound

This paper cites LISA: Reasoning segmentation via large language model.

Vision as Unified Multimodal Generation LISA: Reasoning segmentation via large language model

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.082044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:29c18d527da2f25492c3cfa68ca830ff3e11ca636a0ecb1a89bd459698f41043

Observation 98e17c95-466f-4ffa-91d9-471890444613 · outbound

This paper cites Text4Seg: Reimagining image segmentation as text generation.

Vision as Unified Multimodal Generation Text4Seg: Reimagining image segmentation as text generation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.062902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:711e7aaaabf22f423ed4eeaf363fb206cab4151ebe12a18aa61121af6a36d2b1

Observation 621538ba-4de9-4106-aa03-8be5f884392c · outbound

This paper cites Uni-Perceiver v2: A generalist model for large-scale vision and vision-language tasks.

Vision as Unified Multimodal Generation Uni-Perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.071702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e646a7ed1c842abc878297eb99b3757f0478604ff7a8d874babd1fb00c4f60ff

Observation 64076dd6-dcaa-441d-8aa4-f3a38dc63264 · outbound

This paper cites CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark.

Vision as Unified Multimodal Generation CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.312887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:a6be32d0bc29fd9eacb08bc2c9ebb56c79622d225eeaac140be2fd05f90c0154

Observation 05b4fd5e-cd60-444e-a7ca-02df13f12593 · outbound

This paper cites Tablebank: A benchmark dataset for table detection and recognition, 2019.

Vision as Unified Multimodal Generation Tablebank: A benchmark dataset for table detection and recognition, 2019

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:04:27.080328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:9c44619c9563c417cb7e9ff18ecd78aa59609b382e82e177c43a917bbca400fe

Pith citing papers

Observation 23f1e4ba-6321-4b1e-a023-24f3724d5ea7 · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Vision as Unified Multimodal Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.738395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.738395Z digest=sha256:8a6fa9e334d079e67b551ca9c0e808b63edb42a078ec2f039d1b78423c82b360