Pith. sign in

Paper Citation Record · LEDGER

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2303.06594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.06594 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:23:32.836660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T14:22:18.666015Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f85ff31d-71c3-43bd-99a0-5650941854ba · inbound

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models cites this paper.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:37:01.800800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:37:01.617345Z digest=sha256:7d1f96827ff41a544931282a5e71332b1e61380e5bbb05e13df21a30a83a0a9b

Observation 4fac5e0e-1e55-4b6f-8834-4bcd53a90ae5 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.603376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:57abca17cc09775aa649c821d31603015898ff6e2c194fff79f1bb048b426ba1

Observation b0407eb2-6147-421f-ae70-a19b3137e8b1 · inbound

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning cites this paper.

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:13:09.013573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:13:08.867745Z digest=sha256:a9b4e163d194dce3e5c886fca3db9a0cd4f2b07f19d38043be596c9c6bb0f00c

Observation 658192d9-7847-432d-8a7e-57eada58229f · inbound

An Embodied Generalist Agent in 3D World cites this paper.

An Embodied Generalist Agent in 3D World ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:22:18.668885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T14:22:18.606817Z digest=sha256:6ab63ebe0a1350c8d8e5dacf18fe3fb9ab62f0c2bd804902397a11478e94d04b

Observation 7d012bea-a758-4e5b-b7e6-288bc1cba0b5 · inbound

Adapting Lightweight Vision Language Models for Radiological Visual Question Answering cites this paper.

Adapting Lightweight Vision Language Models for Radiological Visual Question Answering ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:32.836660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:32.836660Z digest=sha256:216fb01b310bc9d521cf8b6cc9e1214afac38ffa6afd9dc4935f83fc78b9b754

Observation 5afc66ba-1db5-4322-8cac-109c8ec26260 · inbound

ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation cites this paper.

ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 88

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:36:50.080336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:50.080336Z digest=sha256:f8e1600bda57b35f0ae194e95d43afddfdcb4197959e5b9eb46c0a29362269e4

Observation 693d6196-8f77-4f8b-8a6c-a0382fd50ecd · inbound

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges cites this paper.

Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:03.863141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:56:03.863141Z digest=sha256:fb8ddb45cfb3ff6d33d6628494cddab8e376101acbdd7bff3ce43ab790fe109d

Observation ac2d11c2-308f-4e49-badc-dc4acc1c35df · inbound

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning cites this paper.

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:55.687390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:37:55.687390Z digest=sha256:06a6eb3546337253a984c617041d2e45f3011b693d25ada59d33a6a629d98324