Pith. sign in

Paper Citation Record · LEDGER

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning

As of 12 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2411.10252.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10252 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:52:34.445566Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6bab4c03-8d17-4f37-9210-c0c87e688353 · outbound

This paper cites Vqa: Visual question answering.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.677761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.350474Z digest=sha256:aaa73441e25d549289eaabec743860dc8717700e256d823697ab570352d4262a

Observation 8164ff06-df70-402e-a88e-394ba00bd617 · outbound

This paper cites End-to- end object detection with transformers.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.668781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.354065Z digest=sha256:491d7e95d4d523362fb8cfabd8580f48e90aa2b38f12493a8c26b7109068f79f

Observation 497d0734-49e0-43e4-98b1-ea480eb2aa38 · outbound

This paper cites Spatial memory for context reasoning in object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Spatial memory for context reasoning in object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.660024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.357385Z digest=sha256:746aea498db2954cd525319b32540839bfaa474d9d3b4db863eaecd425f6c53f

Observation 4c7d22df-37b1-4823-9f10-fb6af8f61c46 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning PaLM-E: An Embodied Multimodal Language Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.360271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.360271Z digest=sha256:d2e5a15d8917839d03866f6a1ccb5e45b22b7a316915d1c7b653e8b25397d420

Observation 57f6ee1e-1a6e-45f4-8e5b-ed80f5d3c9e8 · outbound

This paper cites Relation networks for object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Relation networks for object detection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.651356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.364262Z digest=sha256:a7c8a90bee085ba4ce6397e41d614605d74215444e58a30ce5f6ebffd5b9f0e0

Observation 7d996949-d5e4-4ee2-96a7-865a816aec46 · outbound

This paper cites Dac-detr: Divide the attention layers and conquer.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Dac-detr: Divide the attention layers and conquer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.643131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.367837Z digest=sha256:9e7c05e6bf331e59dbe209689b670b1074308ee223118275df4a48b00faefbca

Observation 2a9169a6-ae50-4225-b18f-51e92addc00a · outbound

This paper cites Hugging face.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Hugging face

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.634427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.371002Z digest=sha256:b57faa13c415801c424bda73e26b3ea9a1306ecff5b7e1d4202378cf55078a87

Observation 748f5a05-3d9d-431a-81fc-ac3f7c2dff51 · outbound

This paper cites Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.374839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.374839Z digest=sha256:e34b17ae11ccc838e948dd9b850ffa0f08a0bc2fbbcf9402a01cd0d261f3ff28

Observation 16e860df-6a81-4d28-aa0d-6d050ea154ce · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning YOLOv11: An Overview of the Key Architectural Enhancements

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.378225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.378225Z digest=sha256:8264cded4ccd0d377a11167c93c0c91c25561aea0e366e8eacec2468112ba8d8

Observation 1c198296-dd7d-45cd-88c7-d7975601b408 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Seed-bench: Bench- marking multimodal large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.381052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.381052Z digest=sha256:c0c37fd1545f7b1035f4c680ed1178a976f4370c35d1aff39074d0f854442d9f

Observation 22fb73db-7ecb-4055-82cf-01c456cb688d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.384365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.384365Z digest=sha256:43c01add618007efa47954b3d37f4b5a4dbc528eaec7ce67463e72105b4ec905

Observation 5456ba81-9a04-4517-8dda-7cf406361a37 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Microsoft coco: Common objects in context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.387936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.387936Z digest=sha256:5e8c5d2e8c531b9f0c40d3074ebe665da29f37a741ad45f7579ede74ac9bfade

Observation fd188497-08a5-47c4-92bc-71550d0c102d · outbound

This paper cites Visual Instruction Tuning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Visual Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.390932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.390932Z digest=sha256:0fab71a0472139b3e3901eae6589e1797e600bf7fcfee112eebc561e11f08d60

Observation eba186b5-0727-42b6-9274-7cccfc43bb54 · outbound

This paper cites Cigar: Cross-modality graph reasoning for domain adaptive object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Cigar: Cross-modality graph reasoning for domain adaptive object detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.613730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.393995Z digest=sha256:2705b59e50f82741c47a9ae08399e3d48e1fd5a39ffdb5b0017cb2edc1f4ee23

Observation bc87118b-775a-4597-aa73-379630865c19 · outbound

This paper cites Rt-gcn: Gaussian-based spatiotemporal graph convolutional network for robust traffic prediction.In- formation Fusion, 102:102078, 2024.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Rt-gcn: Gaussian-based spatiotemporal graph convolutional network for robust traffic prediction.In- formation Fusion, 102:102078, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.605679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.396516Z digest=sha256:a56195dd6313a8bca16d7252492576822980276696acd2d4701feb8cb965205b

Observation 0e8b8201-ca0b-4297-a91e-492235df3795 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Compositional chain-of-thought prompting for large multimodal models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.399174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.399174Z digest=sha256:9acc78c63cc7afd7b9287fe0e3156c91d77ccbb2d4881696c8084cca28f84050

Observation df02a91d-858c-4a10-898f-251054fe041b · outbound

This paper cites Improving multimodal datasets with image captioning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Improving multimodal datasets with image captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.402350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.402350Z digest=sha256:f65b6d94e1a867d26cdc503964d9dffd53ea6bc17e33c122199c3b603dfe76ff

Observation 1c40f5f1-9591-4320-aa40-0ad5103418ea · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.589887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.405498Z digest=sha256:c9bacec8109d2ef7b61a90631df773a5aa2a7d5b4ca80f028045d1b40b1ecaf3

Observation 9a08821b-b50d-4e3e-9799-65fc3ddab5ab · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.408174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.408174Z digest=sha256:0fe70924c213724403a87a849e96ad388f69bf40fc64d54119fbe90edb79787c

Observation c65c301f-da18-4661-95fa-6e4a9b7e33b4 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Vipergpt: Visual inference via python execution for reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.411067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.411067Z digest=sha256:ff784325559e03412f47b65cfa7104fe960bdfa324123555e5be54e13693c3e8

Observation 8d59b5e0-b21d-4934-aa4d-f180653602f8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.413960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.413960Z digest=sha256:3f570d6509ba28064dafd1e8a3606c3bd4b096b74b47988ce8727106ae4e9f5c

Observation beb84e66-a76a-482a-ab85-60f7be330d96 · outbound

This paper cites Sw-yolox: A yolox-based real-time pedestrian detector with shift window- mixed attention mechanism.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Sw-yolox: A yolox-based real-time pedestrian detector with shift window- mixed attention mechanism

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.573232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.416975Z digest=sha256:3c421df53b2c6897f0a09b63095062007e818f9636fc525ef2139e91874251e7

Observation f56eda0a-6fef-4042-ae26-7cbf19d71c6b · outbound

This paper cites Robust motor- cycle helmet detection in real-world scenarios: Using co- detr and minority class enhancement.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Robust motor- cycle helmet detection in real-world scenarios: Using co- detr and minority class enhancement

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.565167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.419779Z digest=sha256:8d9dce24cf5d91a234952be5de21e97b9d793fd591295de72f20cf6ea45426fd

Observation c256903f-50eb-447b-8689-ebdac6f56ea5 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.422536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.422536Z digest=sha256:0a163e6c701c7745700b0cafac8d751cb4809887bb192dab0b19695dac3dac5c

Observation 1f29ecd7-33d0-4a94-81a3-511b922490d9 · outbound

This paper cites Spatial-aware graph relation network for large-scale object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Spatial-aware graph relation network for large-scale object detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.553104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.425201Z digest=sha256:0db6a5e71203db8b7ebc044e63208985d1b6f6e347d3dc7da7bae1530c7fbaec

Observation ba35c89b-9d60-4587-ac8f-b00bf13a94d0 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.545221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.427800Z digest=sha256:5498a2b9ce4e0060ccb9f3eec35039319f866d3baf3898593db7814e90133ec0

Observation 862dfc8f-5428-4ee9-9e01-39481688e8ba · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.430716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.430716Z digest=sha256:912d1d4acf23829c5edfbda7e39861d05846a831af1dd43633168758e40fa8bb

Observation 13e55193-fec1-457e-ab7d-3578acd0cec5 · outbound

This paper cites Ms-detr: Efficient detr training with mixed supervision.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Ms-detr: Efficient detr training with mixed supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.537183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.433880Z digest=sha256:e5bc73269b39a346f8268044b6d8a7bf21dfc6bf29de97579fbc0d1006b1623b

Observation a7856eac-afc3-44f8-9bc2-2cb943cffae9 · outbound

This paper cites Rgrn: Relation-aware graph reasoning network for object de- tection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Rgrn: Relation-aware graph reasoning network for object de- tection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.529019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.436818Z digest=sha256:eaa2678c9962c62e932b75595d57eeffe18f090aadab50dfd09ca2384cb1ca83

Observation 4fdae871-b264-4c8c-b8bb-3cd709c50eb5 · outbound

This paper cites Semantic relation reasoning for shot- stable few-shot object detection.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning Semantic relation reasoning for shot- stable few-shot object detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.519833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.439458Z digest=sha256:663df7e8d88ffb7c5e0762779d0781f6d221809777a646a4c3419d9dd41d159b

Observation 11a22183-2bb6-4e6e-a4bb-d76a42d65ae9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:52:34.442780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:52:34.442780Z digest=sha256:fafe22709093cccf6bd1b9ad92cc4afc58f0f8730daa3db81871b38121f8631c

Observation 6f9c0d52-e485-4146-8a64-2404eadc8adc · outbound

This paper cites An efficient two-state gru based on feature attention mechanism for sen- timent analysis.

Visual-Linguistic Agent: Towards Collaborative Contextual Object Reasoning An efficient two-state gru based on feature attention mechanism for sen- timent analysis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:52:34.510345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T19:52:34.445566Z digest=sha256:ffd12d0d103d61c336bcdd2752dab0937b0855c07871813757412b6186c668ea

Pith citing papers

No inbound Pith citation observations are available.