Pith. sign in

Paper Citation Record · LEDGER

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2411.10252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10252 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:52:34.445566Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6bab4c03-8d17-4f37-9210-c0c87e688353 · outbound

This paper cites Vqa: Visual question answering.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.677761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.350474Z digest=sha256:1a46db72564ccea25c281db41d7df5781c8f7fc689b53a25626cb39e5e9ef1cd

Observation 8164ff06-df70-402e-a88e-394ba00bd617 · outbound

This paper cites End-to- end object detection with transformers.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.668781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.354065Z digest=sha256:bdaceafd1c2a86ec48d759d70a88862b7703e4e39b558c405363432109b05b75

Observation 497d0734-49e0-43e4-98b1-ea480eb2aa38 · outbound

This paper cites Spatial memory for context reasoning in object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Spatial memory for context reasoning in object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.660024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.357385Z digest=sha256:c0d7a766f132a1d0d40c16a193c90a8e80cb9feb1a2019dc5e767c288fb9e791

Observation 4c7d22df-37b1-4823-9f10-fb6af8f61c46 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.360271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.360271Z digest=sha256:4681eb5f441c524a0b1c1f9a9d579718d5a0acbd9e8fea607910688bee83b746

Observation 57f6ee1e-1a6e-45f4-8e5b-ed80f5d3c9e8 · outbound

This paper cites Relation networks for object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Relation networks for object detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.651356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.364262Z digest=sha256:ebca99af00260e227a5e3f320e5036fd2166aafdeef54cc8e9eef6aef0a3bcbc

Observation 7d996949-d5e4-4ee2-96a7-865a816aec46 · outbound

This paper cites Dac-detr: Divide the attention layers and conquer.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Dac-detr: Divide the attention layers and conquer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.643131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.367837Z digest=sha256:fcde777a51cdc7fe7e3703b6d9b715a6bffe08a21e2fb014b8508648719fb9b5

Observation 2a9169a6-ae50-4225-b18f-51e92addc00a · outbound

This paper cites Hugging face.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Hugging face

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.371002Z digest=sha256:5c20456c86acd03f7bff7faf905a44166ffc99cb32bcf296616426baf57a1b27

Observation 748f5a05-3d9d-431a-81fc-ac3f7c2dff51 · outbound

This paper cites Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.374839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.374839Z digest=sha256:af023014f7df15d36bd60ec2186dabc68c510d6424b31949e93e9e6119f685e0

Observation 16e860df-6a81-4d28-aa0d-6d050ea154ce · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning YOLOv11: An Overview of the Key Architectural Enhancements

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.378225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.378225Z digest=sha256:82911209fdeac8c718a5c89eb4ca96855a7fb0baa9ac7e0e8616521e55ba5a99

Observation 1c198296-dd7d-45cd-88c7-d7975601b408 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Seed-bench: Bench- marking multimodal large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.381052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.381052Z digest=sha256:7966f06fdad2b0760f871050ba2cb536686957b1535dd398cfcd17550d2c50e6

Observation 22fb73db-7ecb-4055-82cf-01c456cb688d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.384365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.384365Z digest=sha256:e89d74febc95a452fa414c535db867bf6cffea4b96f9792562bf08e920726772

Observation 5456ba81-9a04-4517-8dda-7cf406361a37 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Microsoft coco: Common objects in context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.387936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.387936Z digest=sha256:2b77675e8b5225f4a214afe3bc0df98c703ecea83b588b58345177be4e03a840

Observation fd188497-08a5-47c4-92bc-71550d0c102d · outbound

This paper cites Visual Instruction Tuning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Visual Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.390932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.390932Z digest=sha256:0dcdc258679602050e049620ff78d41f2f5bab153dc966459d02a6ff7402476b

Observation eba186b5-0727-42b6-9274-7cccfc43bb54 · outbound

This paper cites Cigar: Cross-modality graph reasoning for domain adaptive object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Cigar: Cross-modality graph reasoning for domain adaptive object detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.613730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.393995Z digest=sha256:669a41ce0a493c88d1e24c41bb647f564658ef1600a2f89a6b2645da552f5a79

Observation bc87118b-775a-4597-aa73-379630865c19 · outbound

This paper cites Rt-gcn: Gaussian-based spatiotemporal graph convolutional network for robust traffic prediction.In- formation Fusion, 102:102078, 2024.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Rt-gcn: Gaussian-based spatiotemporal graph convolutional network for robust traffic prediction.In- formation Fusion, 102:102078, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.605679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.396516Z digest=sha256:4830dadc6686376b4f4c6bfed8807dee0f04364a8dc59bdd175899318dbb8d25

Observation 0e8b8201-ca0b-4297-a91e-492235df3795 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Compositional chain-of-thought prompting for large multimodal models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.399174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.399174Z digest=sha256:89a314409609d6f670b88ab282ade7a9b91e562b8e3e537a62cdea330a51017e

Observation df02a91d-858c-4a10-898f-251054fe041b · outbound

This paper cites Improving multimodal datasets with image captioning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Improving multimodal datasets with image captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.402350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.402350Z digest=sha256:b077f4906f27a9fe79d2558c4bdee4a87ea483f7bd47e731aac8ae82339e047a

Observation 1c40f5f1-9591-4320-aa40-0ad5103418ea · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.589887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.405498Z digest=sha256:99cf7ef915b6b1bc6944346b280a71da0d549e106be6f21f16f9551aeb45243a

Observation 9a08821b-b50d-4e3e-9799-65fc3ddab5ab · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.408174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.408174Z digest=sha256:ad4602a7d25b3638f120304a3dd2019fb0d26132113ec29c8b41866fac88260d

Observation c65c301f-da18-4661-95fa-6e4a9b7e33b4 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.411067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.411067Z digest=sha256:20e6af10df4053e3c083cd40f2dc162a4505eb20608834854e611d40ac1ab0a7

Observation 8d59b5e0-b21d-4934-aa4d-f180653602f8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.413960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.413960Z digest=sha256:a3574cf5faeb2a6f9c8dd5d87ed1dc0a636a7d3ae1266cf72a79ff264365a53f

Observation beb84e66-a76a-482a-ab85-60f7be330d96 · outbound

This paper cites Sw-yolox: A yolox-based real-time pedestrian detector with shift window- mixed attention mechanism.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Sw-yolox: A yolox-based real-time pedestrian detector with shift window- mixed attention mechanism

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.573232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.416975Z digest=sha256:6fa44d32086381c9bacad38d9dac0123b4e1a796d1f60c26fb36d9746b266fa5

Observation f56eda0a-6fef-4042-ae26-7cbf19d71c6b · outbound

This paper cites Robust motor- cycle helmet detection in real-world scenarios: Using co- detr and minority class enhancement.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Robust motor- cycle helmet detection in real-world scenarios: Using co- detr and minority class enhancement

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.565167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.419779Z digest=sha256:8b6916f680bce3f252bd3c48894df3cc062b00648c9795980ca9d0e857df1b4e

Observation c256903f-50eb-447b-8689-ebdac6f56ea5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.422536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.422536Z digest=sha256:74593f64e9f7272e152e363ed8d7338504ed4c79139b4e2b900c7c4c8961e76d

Observation 1f29ecd7-33d0-4a94-81a3-511b922490d9 · outbound

This paper cites Spatial-aware graph relation network for large-scale object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Spatial-aware graph relation network for large-scale object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.553104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.425201Z digest=sha256:f5673435fc8fb07571c2ead6f764ba785f1ff4fbeb36e555ba4261cd984379b0

Observation ba35c89b-9d60-4587-ac8f-b00bf13a94d0 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.545221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.427800Z digest=sha256:100e1d9efd265911a6e65aa98bf6b5c02aff40e2876d8dcb991dce340e452cbe

Observation 862dfc8f-5428-4ee9-9e01-39481688e8ba · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.430716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.430716Z digest=sha256:20ed6d1228fb73cac397ea76a1af88aa8f7110e093702b3c9c1d3a4bdb67aeed

Observation 13e55193-fec1-457e-ab7d-3578acd0cec5 · outbound

This paper cites Ms-detr: Efficient detr training with mixed supervision.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Ms-detr: Efficient detr training with mixed supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.537183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.433880Z digest=sha256:b2d7381a7746114cfedd70bbb383745cc68a07e9185dbc2f5111ac57107d988d

Observation a7856eac-afc3-44f8-9bc2-2cb943cffae9 · outbound

This paper cites Rgrn: Relation-aware graph reasoning network for object de- tection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Rgrn: Relation-aware graph reasoning network for object de- tection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.529019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.436818Z digest=sha256:3257e32adcd025aaa7992715615f2e8d479dd5afa39e11a7ac15a31f34cc135c

Observation 4fdae871-b264-4c8c-b8bb-3cd709c50eb5 · outbound

This paper cites Semantic relation reasoning for shot- stable few-shot object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Semantic relation reasoning for shot- stable few-shot object detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.439458Z digest=sha256:06d8e5a037ebe25cec5c40b47ffe985b395b04ed2d7e790aa97810ae6aa7d2d6

Observation 11a22183-2bb6-4e6e-a4bb-d76a42d65ae9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.442780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.442780Z digest=sha256:af0b4ba1348872d1b612ff1e03b81ce96ce9bfb762eeb00682ec1182d90c915d

Observation 6f9c0d52-e485-4146-8a64-2404eadc8adc · outbound

This paper cites An efficient two-state gru based on feature attention mechanism for sen- timent analysis.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning An efficient two-state gru based on feature attention mechanism for sen- timent analysis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.510345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T19:52:34.445566Z digest=sha256:32c5ef273933d0fd789ae46b4cca2e637ea13213f64e6b834f18e38d9ca89b3b

Pith citing papers

No inbound Pith citation observations are available.