Pith. sign in

Paper Citation Record · LEDGER

Performance Analysis of Traditional VQA Models Under Limited Computational Resources

As of 8 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2502.05738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05738 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:10:38.360640Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93978d07-bc26-455b-8dd6-472c274992bb · outbound

This paper cites Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.619877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.244971Z digest=sha256:59598583563aaa20b54cc2f57974551eb057ee1ea9d4c4cd0b13d721c6c56954

Observation 54be280f-3706-4b34-9961-c3a2669e43b8 · outbound

This paper cites Learning Convolutional Text Representations for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Learning Convolutional Text Representations for Visual Question Answering

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.594649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.260605Z digest=sha256:f8e3ced70bec882ba9b4d0b0d8cdb4378c1419e369664934e15de0a0ef25604b

Observation 4ef1d1db-b1c5-4d74-9e93-0b166b93b7e1 · outbound

This paper cites Structured Attentions for Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Structured Attentions for Visual Question Answering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.579059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.265022Z digest=sha256:dea4db5330fda7bb1be15422af3ebf9f6f71e7c88e5b1b50b30ede1f2b72db3c

Observation 36ec7fbb-05a3-481e-8938-e2da7d1e3f1c · outbound

This paper cites iVQA: Inverse Visual Question Answering.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources iVQA: Inverse Visual Question Answering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.563465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.269298Z digest=sha256:5ee2c9c238e88b8ae8201e614123295db160aa44da6755e1dd26df87f10bee60

Observation 0cbe02dc-f50d-47f4-b404-8ea3b1e126d1 · outbound

This paper cites Self-critical Sequence Training for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Self-critical Sequence Training for Image Captioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.281868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.281868Z digest=sha256:b2ab715ee580a9bb958f99cacea17c162193ab422dc41c2f54780c06ec30d8ad

Observation d9b0ba65-3614-475f-a3e6-74181f19a943 · outbound

This paper cites X-Linear Attention Networks for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources X-Linear Attention Networks for Image Captioning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.520079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.295039Z digest=sha256:3421903157a996ccfb6658546089bac488bc042e602b9b5cf5601852439f6463

Observation dee30146-85d3-47fd-849a-dbc2a2bdfa95 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.299321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.299321Z digest=sha256:29c843408ffec8691349c38c68b4816f743205e31afcd97405d44d07f4165ef0

Observation 99ef28dd-9510-4c9b-9edb-cf56bd770c80 · outbound

This paper cites Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning,.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Beyond a pre-trained object detector: Cross-modal textual and visual context for image captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:10:38.632483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.303290Z digest=sha256:378989f1979257f071770b5dabc7383ef1478fdef995b3418354a802bd224d1c

Observation 11fe739e-246c-4a01-afcd-f6a528289fe1 · outbound

This paper cites Transformer-Based Multi-modal Proposal and Re-Rank for Wikipedia Image-Caption Matching.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Transformer-Based Multi-modal Proposal and Re-Rank for Wikipedia Image-Caption Matching

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:10:38.494657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T18:10:38.306986Z digest=sha256:9fd469d2dd218da61cb2bc437431509869c3ac177e4f216840441f186324c285

Observation f6c2cb35-f96e-4aea-9037-75b6b9e78df8 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.311056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.311056Z digest=sha256:12b53d22f5c5caae8d30b428f3eeb0992394a5dcac1bf88aa413010a9eb22c23

Observation 0f24e703-7482-4897-8b24-d8eca8e55813 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.315617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.315617Z digest=sha256:ac883e19081c8e4889f86fdd799c2ec8f7fb585d6de2070d55609106a70b96c5

Observation ed3f86ac-51f9-49e3-a784-ef6b44ceaaf6 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Training data-efficient image transformers & distillation through attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.319474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.319474Z digest=sha256:fbb4970484c1e908157bf822fc2255e9f212d80965867cf09e4668cb56997328

Observation 5b875199-5513-4653-99ad-c25625dfa6a5 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.323443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.323443Z digest=sha256:f4436675650d43780dfb879ecde06d57a15d3af996aea9565e9bdd1be3785aa3

Observation e23c349d-d4e3-4c1b-b770-a7998624d357 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Distilling the Knowledge in a Neural Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.327827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.327827Z digest=sha256:d6b858ac7850c4cb81e5bb230112232f533f7995dad14c0450c9b10e35281b5b

Observation d5d66f28-4c3d-4039-9f6d-cc525750cd45 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources TinyBERT: Distilling BERT for Natural Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.356476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.356476Z digest=sha256:84dde81f6507301f081ad676cd202ce3f4ce6bbf80da1fd192a5c7c2553634ea

Observation 3851e900-dd65-4cdd-8426-fa4fe8fdbd91 · outbound

This paper cites Learning both Weights and Connections for Efficient Neural Networks.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Learning both Weights and Connections for Efficient Neural Networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.335940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.335940Z digest=sha256:cf24b50b959e4acd97dfb9e6d7cf68dab1308dadbf0edccfb6655fbaa21ca83d

Observation 50dcd689-b2e8-488a-af13-aaa88a2cbecd · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.353136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.353136Z digest=sha256:9f54f0914fe371a384368277db51e4b1e14b6ecc0108ad9748d2cd5d0c473346

Observation d01917d4-6eed-4a4c-b8c1-de86dd67f13e · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.360640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.360640Z digest=sha256:9c67887fe0a05a7bf9998455880644a9d6ea8abb8b1c7b611ae04a9ea148d022

Observation f4d9f3a9-d543-4878-8903-1c708158c3f5 · outbound

This paper cites Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.278124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.278124Z digest=sha256:a04359be713f6e83a48eaf163aa373512f6157b4e08e71602105ff982c84dd6a

Observation 9d0941e8-7218-4156-b52d-027e66c44e7a · outbound

This paper cites Bidirectional Attention Flow for Machine Comprehension.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Bidirectional Attention Flow for Machine Comprehension

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.256348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.256348Z digest=sha256:37c206d72d61863a4171f35e4ad02575cfb9834c2d23c351abe6fa33c1c4b3c7

Observation aa2525b9-c4fe-4f2a-90b8-d18bc7ab16ff · outbound

This paper cites Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.344144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.344144Z digest=sha256:85bca82c1e5b4e0193e8fcb796d3f7c21c927eafdd9e48fb2957fbfdaee7d537

Observation 2385e11e-2888-4327-9185-b82d5e498178 · outbound

This paper cites Meshed-Memory Transformer for Image Captioning.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources Meshed-Memory Transformer for Image Captioning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.290999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.290999Z digest=sha256:dab34e395664e4edd6862803628108f3fc8433806f32bc6e6bb09ccefebd726e

Pith citing papers

No inbound Pith citation observations are available.