Pith. sign in

Paper Citation Record · LEDGER

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.05859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05859 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T22:27:52.861022Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact23
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4c887a8-553f-4055-92b9-b475a4163984 · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.316541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:c4b799cf45d4639019a4590d65d10b69d21bfb825d9d9a2490840c6d83b1e274

Observation a86a0f61-0c8d-4e7c-8256-b81a4a0d4b11 · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.277778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:069431a084cca5630bd0239dd202f1e603777f5ceae6c6d17f56fdfbe9d7888c

Observation a8204efc-3275-425c-b072-f346397254a0 · outbound

This paper cites Spice: Semantic propositional image caption evaluation, in: European conference on computer vision, Springer.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Spice: Semantic propositional image caption evaluation, in: European conference on computer vision, Springer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.286591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:2b8bc1191186187892354400bf51e53359bc605c00e51eff206fa56a54802658

Observation 7d62fa39-4409-48df-8b4a-4e9159c111c0 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.299263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:4c69ad3f313abcf98608744e9f55f709d61ad9d981ccf1bcac960aae258b7d4c

Observation 5609b157-c88b-48ac-a63b-e511b5f6377f · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.272360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b12aa64abe7b749836a0cb406bc92e9a07dae85965982e7709c29fd036f8d445

Observation c04100e3-cd8e-4ba3-a284-21d5c58c1a73 · outbound

This paper cites Context-awarevision-languagemodelagentenriched with domain-specific ontology for construction site safety monitoring.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Context-awarevision-languagemodelagentenriched with domain-specific ontology for construction site safety monitoring

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.267250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:06096d3ddd40888d8cd02f48ec1619463687c79e86c058d9a0f23734a5803c02

Observation db022c64-081e-45f2-8731-59aa3cee9036 · outbound

This paper cites Enhancing vision-language model for construction safety inspection via visually grounded reasoning and reinforcement learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Enhancing vision-language model for construction safety inspection via visually grounded reasoning and reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.274283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:019c3601e71e0bf17c224380c8d8cd25b5ea05d05c0783edf5f01f30dd5c95ab

Observation 15e6ae3a-8d24-4bda-a6f4-97f1cf8b41fa · outbound

This paper cites Augmented reality, deep learning and vision-language query system for construction worker safety.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Augmented reality, deep learning and vision-language query system for construction worker safety

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.449701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:6395b850422aac80577e905e5f08a49086c472608d650d0adefcacc07d21030b

Observation b97a474a-15c1-478c-a0d9-7f4144f25c6a · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.588253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:3a6e6708d5d3e37739c34b5b2c6076dacd0cb0408354596d807767efd570b8c0

Observation a8b5b3c8-cc53-4915-913d-0bbd6b662b3b · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.573081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:139036186ec716422a8871065190033344604474c68f4df60375241a45dd7fb0

Observation 50a78136-ca05-4dfb-b2c1-b1ec9f796d1c · outbound

This paper cites Tailored vision-language framework for automated hazard identification and report generation in construc- tion sites.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Tailored vision-language framework for automated hazard identification and report generation in construc- tion sites

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.452922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:6e004d6df44a0a4438e7cce3a136e48839b5318bebbb5cdec2c772debca65199

Observation 79dd781b-948c-4f84-b7e2-c20924fd489e · outbound

This paper cites Canmultimodallargelanguagemodelstrulyperform multimodalin-contextlearning?,in:2025IEEE/CVFWinterConferenceonApplicationsofComputerVision(WACV),IEEE.pp.6000–6010.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Canmultimodallargelanguagemodelstrulyperform multimodalin-contextlearning?,in:2025IEEE/CVFWinterConferenceonApplicationsofComputerVision(WACV),IEEE.pp.6000–6010

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.311086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ac2c7cea4312c81133d1822e6bf1483ca19009b20b5cfa0017780e85cd3c7418

Observation 1db83eaa-394d-4813-a22f-8ed71a09fd3d · outbound

This paper cites Are large pre-trained vision language models effective construction safety inspectors.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Are large pre-trained vision language models effective construction safety inspectors

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.439667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ff435b600e02a657ad643ea139b75ad1849f92524f003d662fed384b8493b5a6

Observation e8539243-9771-4b00-9826-177b464188bd · outbound

This paper cites Expandingperformanceboundariesof open-source multimodal models with model, data, and test-time scaling.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Expandingperformanceboundariesof open-source multimodal models with model, data, and test-time scaling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.308515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:67a62e3a5701742d099a5448711ac52ac001f2702e3bccbccab8a25ec0318a2a

Observation 41ae7e5f-ee17-45c6-a353-c484f16f95ec · outbound

This paper cites Chain of Thought Prompt Tuning in Vision Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Chain of Thought Prompt Tuning in Vision Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.569827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:734dae5dbf1096d6cd417eefa4edba2a9eb421abc633f7f616f09a2a0b455ec2

Observation 1d220ae5-72a7-43e9-8af5-4a2abd499063 · outbound

This paper cites Intelligent virtual assistants with llm-based process automation.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Intelligent virtual assistants with llm-based process automation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.313135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:1c66f748c3edc244664f73d3ab2ea9c6f5f81836d019ae578375bc7ab6ffa0fe

Observation 64dbea47-ccd8-4b3e-a774-abadb1139a23 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.595517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b9524c8c70d813b4ba279c74e6bf0f2c70b72e07f234e591e67841a006804bac

Observation 34e022b9-5582-4941-837f-beb8b64786bd · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.306558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:9488e1980bcf45281790f47b4b003c4b87067e4946320e91539235e507a57ce1

Observation 8cd8e49b-6f76-4ef8-9070-523d55ac4409 · outbound

This paper cites Vru-accident: A vision-language benchmark for video question answering and dense captioning for accident scene understanding.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vru-accident: A vision-language benchmark for video question answering and dense captioning for accident scene understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.302910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:aa2e3340f4d952b8084515967207ea42faa7bbcae63630b1a713cb50a775e925

Observation 39d9e36e-2391-4d2c-8153-146d5dd508ea · outbound

This paper cites Safe-llava:Aprivacy-preservingvision-languagedatasetandbenchmarkforbiometricsafety.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Safe-llava:Aprivacy-preservingvision-languagedatasetandbenchmarkforbiometricsafety

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.601075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:499e7e3b7b6df4934c9c010715f1ffbde3ecb7a51458d5095745231134ef7ffe

Observation 33641ca8-8c68-45d1-b53b-fddf2cd063c1 · outbound

This paper cites Res-bench: Benchmarking the robustness of multimodal large language models to dynamic resolution input, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Res-bench: Benchmarking the robustness of multimodal large language models to dynamic resolution input, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.304705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:64779f82b2cce91c72960d48f07cc25d2a31b884557ae593bffb9670b5e7a2a0

Observation ea9338c7-0f61-4527-99cb-c07ee5d7f74b · outbound

This paper cites Construction site fall hazard identification and automated captioning using adapted vision-language models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Construction site fall hazard identification and automated captioning using adapted vision-language models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.467360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:114db873ef725a5e567845072fd9def363f539def797f4730d1b226c909af54e

Observation f3830f6e-72a3-44e8-b6b8-88a0da821de7 · outbound

This paper cites AdaptVision: Efficient vision-language models via adaptive visual acquisition.arXiv preprint arXiv:2512.03794, 2025.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring AdaptVision: Efficient vision-language models via adaptive visual acquisition.arXiv preprint arXiv:2512.03794, 2025

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.551322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:8ac1dcd9c45028341748b5108d880db1683742a66d9af442e16ff67c0f7e6ad5

Observation 3761a0a6-e163-4383-8877-fd02b169a6b3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Llava-next: Improved reasoning, ocr, and world knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.314843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:c8c12c87d0fcdb929faf51490ca326066af62ce2f6952256f7e1f0e9a59b7c9b

Observation 40e2f463-7018-460c-b20b-fa0004ecc42c · outbound

This paper cites Visual instruction tuning, in: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visual instruction tuning, in: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.301085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ba06e45c2f3d2d0f3b13b303d4020ffbd4b792108ba8687e5b72086c2745f22a

Observation 0782e0f5-ef7b-4b50-bb0d-c21cd184fb6a · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.580445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:cd9ff8c2f28fbc6553581098809524979f201f7edf33b5ec425c27a898c5c5fc

Observation 675779f7-775f-42f0-8e63-74351b6ab11e · outbound

This paper cites The Llama 3 Herd of Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring The Llama 3 Herd of Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.575764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:908a70c2cb4eb0a868425b860eb41e793261e4d6b63d88bd348fc78fe7a9087b

Observation 7b142371-113d-4e6a-bd64-fcc9f6097183 · outbound

This paper cites GPT-4 Technical Report.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring GPT-4 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.578047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:4a5d23ca17491145cc74a7b7d0517dcdc1ad1ffad178f43f336765998f80200d

Observation b911e575-3131-4eb4-a6d0-3ddc7845e535 · outbound

This paper cites Privacy-Aware Visual Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Privacy-Aware Visual Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.590687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:854d097e13c150a31e629d804d330c180b35ac336dc9f4e219784de38098ac4b

Observation 81a41807-1877-427e-a3fc-02ad0209459e · outbound

This paper cites Vlm-robustbench: A comprehensive benchmark for robustness of vision-language models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vlm-robustbench: A comprehensive benchmark for robustness of vision-language models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.603788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:c99c8f42d3f27f341c8a8759184eaf7acb0b20a05297466b2f1cf79efb04e03b

Observation fa13d331-0a3d-4d9b-a0a6-27624befe93f · outbound

This paper cites Vision-language models for edge networks: A comprehensive survey.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vision-language models for edge networks: A comprehensive survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.292231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:27b7bdf89e03ca23693789b7a615024244c81b8c752f273327c5acbbf2d28373

Observation 4ef34baf-f4b8-43b6-b59e-71e9e1cfc553 · outbound

This paper cites Upop: Unified and progressive pruning for compressing vision-language transformers, in: International Conference on Machine Learning, PMLR.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Upop: Unified and progressive pruning for compressing vision-language transformers, in: International Conference on Machine Learning, PMLR

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.297412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e7e6ca0f9254a46581055198dc0ffd298df638297f88a1eb3cff21084c5048b7

Observation 01db1329-854a-4e32-a16d-f8f63d9592ac · outbound

This paper cites Real-time safety detection on construction sites using a vision-language and nlp- based model.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Real-time safety detection on construction sites using a vision-language and nlp- based model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.446394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b7a6a4c2e4300db4f803ec673c5cd33c8061da1b0df196566a068a0e510eb76a

Observation ef1d064f-93a1-4846-9930-8abc511a1541 · outbound

This paper cites A double thinking enabled visual language model for open-set construction site safety inspections.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring A double thinking enabled visual language model for open-set construction site safety inspections

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.290434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:1dd6512166ec08fc1c6df3e766032c913bf5e9b271714588e087e949d54e4420

Observation 5a347992-f026-482e-9bb5-576ef879123e · outbound

This paper cites Region-level vision-language model for detecting distraction behavior and mobility attributes of vulnerable road users.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Region-level vision-language model for detecting distraction behavior and mobility attributes of vulnerable road users

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.461810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:d40d4dd706780f4bd782bdc71c2c8e7aaf93550d34d8c098a18522b3628f30fc

Observation a9ca898f-8578-4960-ac15-8a82a00adb27 · outbound

This paper cites Visual question answering-based referring expression segmentation for construction safety analysis.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visual question answering-based referring expression segmentation for construction safety analysis

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.458666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ad4c779c871709be1baea98e23e7b2c53c1e038f2ab9e0e8cad0a226044c7a28

Observation 03238cdb-052e-49f1-8880-de628907165f · outbound

This paper cites Selma: A speech-enabled language model for virtual assistant interactions.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Selma: A speech-enabled language model for virtual assistant interactions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.295114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:65ffc5c6553d1dc5e7f800bf494d88927240fa860dba8d9599be096c55dcf40e

Observation 34c9128a-0134-4b53-9856-11c20b6f6627 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.288555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e13536a749604bebceacd46e1ce4b76d2ccdf57575111a5bccaeea1e3112621e

Observation b89f6176-1db0-4b14-99ce-d82dc42d1d68 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.582870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e7a0a04293883985f609b83b32b2429e4f8d622a9621cc8c08135bd52f9f3e39

Observation 845780f1-5d67-4024-8896-568df97ed835 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.609039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:c961e56c9eeb6f0b5d15cc63160d01a243cb998f07b7374eed56aa0be2822364

Observation 55cfce7d-25f9-42f7-ad8f-44484ae995b2 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.606423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:8d91ac29cc2ce6ca2702762c251a581f25c299379c12747b98ea4cc65acad2d0

Observation 309f9b4b-22ac-4542-9287-bb29ee9c9b1e · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, in: Proceedings of the IEEE/CVF International Conference on Computer Vision.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Llava-cot: Let vision language models reason step-by-step, in: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.281322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:cfebb4e8bdf40c50b885cfced4f7e281671bc0a373e007b9d06a84ec0f3dd7c9

Observation ba0ffa3c-a517-4530-aa69-323d59955cd3 · outbound

This paper cites Edgevideoanalytics:Asurveyonapplications,systemsandenablingtechniques.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Edgevideoanalytics:Asurveyonapplications,systemsandenablingtechniques

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.282854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:265f15ac466467b9fc11c35b872552aca676e5c51cecf90ce607d6b5e9850345

Observation b0423c76-7678-438b-bf77-54c87ea12c95 · outbound

This paper cites Qwen2.5 Technical Report.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen2.5 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.585730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:83cb77b154990dc4bd13de99f0dfe6f653bedd50db51b5adf4849c847210c66d

Observation 1c8752b1-5b71-4006-b0fb-0b9598995377 · outbound

This paper cites Vision transformer-based visual language understanding of the con- struction process.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vision transformer-based visual language understanding of the con- struction process

Reference 45

Resolution
verified exact
doi, observed 2026-07-08T22:35:40.464393Z

Source-reported events for the cited work

correction dated 2024-05-21. Source: crossref record 10.1016/j.aej.2024.05.064->10.1016/j.aej.2024.05.015:correction, observed 2026-07-11T03:14:22.632143+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ced7ce734e7a2ebd0bf979a529a400432c864bae885f4c33703226917ec028dd

Observation d220c57d-6314-422a-94b9-67342edbee60 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visionzip: Longer is better but not necessary in vision language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.284722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:fd9b38f9fe8a85640542053438712c7a5928109b6cef3358839d348c6c793622

Observation 3f04773a-9683-4cc4-9277-5bd91c9cf822 · outbound

This paper cites Visionthink: Smart and efficient vision language model via reinforcement learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visionthink: Smart and efficient vision language model via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.279512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:cacea4a2ca508e75d88304c3aabbe2f78e9e8a9015a67ba133db60f2b75b5274

Observation aec977bd-d0e7-416b-90e9-f9dca407730b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring BERTScore: Evaluating Text Generation with BERT

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.611532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:494b80a9d38b0109566f70afb8616c627937524406227e22827b01e53082ae91

Observation f908bd59-8cb1-41b8-b779-8b2e0a937fef · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.598163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:6e0aea5761394e3afa12177085b207f7ca72947d8015fe870eda6d926b3819f9

Observation 01c5da41-714e-464d-9706-0f192a0a4c77 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Multimodal Chain-of-Thought Reasoning in Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.593114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:52a9ce6551be216f8d9c4e74b3e47cfccaffb591f141f9346cf099112f742394

Observation 370f17ac-fcac-4490-a6de-21e2d736c31f · outbound

This paper cites Mmicl: Empowering vision-language model with multi-modal in-context learning, in: International Conference on Learning Representations, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Mmicl: Empowering vision-language model with multi-modal in-context learning, in: International Conference on Learning Representations, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.276107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:4630793f3de95a8e2ff5816b277d6d4f15ff907443e8fdd9c089bab40dc6e79a

Observation 31197953-97f3-4862-a023-8cb7add18a2a · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.269087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:c0273c8cb00854f55daf39d5575e5ebe22dfe8239657c216898d3477eb4502cc

Pith citing papers

No inbound Pith citation observations are available.