Pith. sign in

Paper Citation Record · LEDGER

KAT: A Knowledge Augmented Transformer for Vision-and-Language

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2112.08614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.08614 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:20:01.008761Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:29:06.156996Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 314f70ec-9202-4977-92e7-e119a1d396f1 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.515015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:7445453b552e45d7ff2ae4cd9c67eb0489b04e60eb4b22f2732dd70ac4b45d5b

Observation e6066f6a-638b-4807-a821-03f2b87b3c70 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:29:06.160017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:f633fc21db748c9ca6fd05b41cfde534d7cfa8710f0330143b61a33b65340cb8

Observation f80fd8c1-cbf8-4f06-9b3e-ffd6659b9abf · inbound

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation cites this paper.

MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T23:20:01.008761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:20:01.008761Z digest=sha256:7a290121fa427d1f276371d683ddda158985f9979f9ddcb199afb150e5238735

Observation 976b4ea1-7398-408c-a3d0-a5c54e7311f6 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:00.032601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:00.032601Z digest=sha256:1cc506f88a6aedcfb11d8bb1b8412f1a831ddf14d32b39a509360541487eab62

Observation 9e38a3e0-8a2c-4e85-833e-23fce8378416 · inbound

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos cites this paper.

Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:40:42.680705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:40:42.680705Z digest=sha256:39c944fb2fd58fc12df1cae4e29dce2e96bf0916c8d8d2fb7f7537ce83c9847a

Observation c52725fd-73e1-406e-99ad-cbe082c6f92f · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.886456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.886456Z digest=sha256:e60de7f4896b663f7bbc917cc6f0c5c0b38f38e9b7ab261fe69e1d3501af3115

Observation daa8d02c-ec27-49cd-b82c-d503d87aedbf · inbound

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering cites this paper.

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:08.543771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:44:08.543771Z digest=sha256:321df4973591b91d2ac23fe3913ac9f4217f69b6d9f60db499dad6855c90a9d3

Observation 30f15c67-efb6-4c83-8562-bc0caaa8c014 · inbound

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum cites this paper.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:baac5713b397bc2781e74a538efd93827c39f05d64e3156a149675d9fa1633e0

Observation 74f849e7-7ed7-4da1-8285-2917f20e487c · inbound

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data cites this paper.

SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T09:59:08.317646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:59:08.317646Z digest=sha256:5b7bc3ac57f439a8ce67e11cf9f7156fafa95a06c4b4892a804c523b347f2844