Pith. sign in

Paper Citation Record · LEDGER

MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2104.12763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.12763 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:37.398957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T11:09:22.426941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2219cda-1b5c-4460-976a-a4b5fa618fa1 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.430430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:1600c6d7f70035f388e8640a02cd902fd6069c2b26400af288475eb1d44b9201

Observation e5177977-f620-44b3-af16-da781b15a257 · inbound

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning cites this paper.

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.398957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.398957Z digest=sha256:d490e201f7768e14b9d512711cd184620645ebef7e754a8408d5a917dacc9ceb

Observation d0c7ffae-0651-4c6d-88ca-aa9489c7923b · inbound

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset cites this paper.

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:25:20.637769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:25:20.637769Z digest=sha256:f30d0c728d56cb9a740fb4068871009251780fe7f4f75186c81ca3791833074d

Observation dd15ed5d-d967-417b-8ed0-c9edfe0d9bd8 · inbound

STORM: End-to-End Referring Multi-Object Tracking in Videos cites this paper.

STORM: End-to-End Referring Multi-Object Tracking in Videos MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.501349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:25:31.777907Z digest=sha256:47cd7a6a37f1ca276a4a9a689b92b7b7df6c76d50ceb1a3627e37bb4fa3b3eab

Observation ec4bf94c-242f-46d6-8637-a030cbb128e8 · inbound

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation cites this paper.

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.455044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:35:44.016054Z digest=sha256:88d7d912bbe4968220034f62c499711e4108c8719bbe71de91d2452fe90ed40f

Observation e4e15ac1-93fa-4776-b7e4-8f2b71b184a1 · inbound

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval cites this paper.

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:57:53.190292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T23:54:32.646978Z digest=sha256:c46e6a57cf41b6f4d766bb6f521f5fe58352139496a1c0a2a996e54e808ff470

Observation 59a5792e-d8c7-485b-a0e8-c3f2bce2d99d · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.442806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:5d79f88befe41527eb51f66558a4352deb2e00d2a5bdfaa8fb4d855e0093409c