Pith. sign in

Paper Citation Record · LEDGER

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:1505.04870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1505.04870 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:13:25.708993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:29:57.874633Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 634c8446-581f-4fc9-a4e7-62490b733990 · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T09:38:09.538355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:a51593b5ececbafdd26e433b645d99e6779067aae109dc49603ce646b512790e

Observation 5841e810-0639-45f3-bec8-1f34f573c01e · inbound

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning cites this paper.

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:52.472320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:47:52.472320Z digest=sha256:9b264d3043945772f2b8166224a077a7f98096b1bde065f6c238373a4b10d1a9

Observation 5e06a9ef-f940-4230-9e47-9a1a2fc2b20c · inbound

POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation cites this paper.

POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:25.708993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:25.708993Z digest=sha256:893464099697c2eea53bd4480457678d5226cb4dd8e59668708f754fc3eaa208

Observation 41bfaf0e-3822-42db-9510-5eb1cf42d5f1 · inbound

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision cites this paper.

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-16T05:10:32.976270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:10:32.976270Z digest=sha256:963e39c880085d329f065bfe48ec825b5e302c2ab83722a1e7e7bc0bfe0131ff

Observation 878890d2-951f-4555-9dd9-6a3e351a436f · inbound

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing cites this paper.

PointArena: Probing Multimodal Grounding Through Language-Guided Pointing Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:24:01.101202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:24:01.101202Z digest=sha256:4b1950f9427b90266d02b9f2484f92aef874dc0b7997e644073862d5c1acc26e

Observation be44f958-c33a-4fd3-8214-91d47c34f5a9 · inbound

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP cites this paper.

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:00.396259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:35:00.396259Z digest=sha256:b8e5e3c238840244570f591a695a32a46b497a1ce8b04563d6eafaa79a973395

Observation 8b6af184-38bb-4335-a2ee-a826093c41e8 · inbound

Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing cites this paper.

Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:04.406711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:04.406711Z digest=sha256:7194f3f9f86698d98a0cd4f08bfd9060dc100cae66b129f8aefadedd9ffc01f0

Observation 2f61d586-ef08-44ce-8b03-c26ca978f2b0 · inbound

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models cites this paper.

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:59:41.110128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:59:41.110128Z digest=sha256:34879be393b05c03a7f83de4ca85e332faa828a3d2bdcd1bd91947f3e88bc8ad

Observation d1e6a688-4d00-4f77-913b-2b58ee3d0394 · inbound

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models cites this paper.

GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:53.686184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:53.686184Z digest=sha256:76eb0d2a6aa316a7e086f6ecacf65c1cadccc7f5a73d1bbe50855dd04e3d6122

Observation 37fe6998-8d8c-4dbe-b484-a1a27a21b5f7 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.349883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.349883Z digest=sha256:28435fa759633081af4cff838122069b64cf8f0c845d9a7e0d01a69269a6ad97

Observation 7cd2decd-4774-4b03-8060-2878214813bf · inbound

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini cites this paper.

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T18:23:50.301057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T18:23:30.681253Z digest=sha256:397c77274799b88531711c7acb2c8f4d402654358bda2ba3fb680474f66f9b9d

Observation 4a6d6ff8-6333-4b91-b2cd-441785d0ea0c · inbound

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding cites this paper.

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T17:53:46.733539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T17:52:59.346637Z digest=sha256:390ec65ded5a97cfbf15ba3031ab98431d33e0700e02339eb01d55dee8af9e2e

Observation bbe0e001-b902-4287-b621-f74b2714a4ba · inbound

Grounded Scaling: Why Agentic AI Needs Deterministic Environments cites this paper.

Grounded Scaling: Why Agentic AI Needs Deterministic Environments Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:42.668100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T10:51:34.720931Z digest=sha256:caa1cf5c012be81018cfd3713c47da276c44996dfec030d316274a5e8f5e9198

Observation db2ec889-8a69-4add-ab49-22d2143fd77c · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:29:57.875807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T00:29:41.291832Z digest=sha256:6af57d1f5830a42638dcb6c0a975a8d36d289df689842f03dfb29111c733ffac

Observation 87159781-7554-4b8f-af18-b7686068359d · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T15:03:32.190184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:6448b0e9d23e278b1c07cc9bf24ae68283b81f23be9519e9fa510e90e80b207f

Observation 27ef41ee-da4e-41f5-aafd-2ca362cdedc0 · inbound

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval cites this paper.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T16:32:55.757864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T16:32:55.757864Z digest=sha256:304aee5285a9facba63c3e5d1346457cb6501d7ee4a7d2bcb7ce811b63629655

Observation 70a783c4-b83b-49c6-b3a1-bf9c97b06cb4 · inbound

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval cites this paper.

Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T16:20:19.874172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T16:20:19.874172Z digest=sha256:ec018dd9d1e6c01695da385477105e840f398721d4f35d952a895217fae9d1ea

Observation f0ec3cfe-b7b7-4b23-afae-c7f11f703ec1 · inbound

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval cites this paper.

MASCOT: Model-Aware Submodular Coverage for Composite-Attribute Text-to-Image Retrieval Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:12:08.692697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:12:08.692697Z digest=sha256:e0f91e5b8d1c6b3e8b588350c321c235e8f83a6a113813a5a9813bd1fefddceb