Pith. sign in

Paper Citation Record · LEDGER

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2607.05859.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05859 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T22:27:52.861022Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact23
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4c887a8-553f-4055-92b9-b475a4163984 · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.316541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:34c78a27a2b3ed842e54bed4b16ee781bde47bbdea3e71a802b2783f192ad9dc

Observation a86a0f61-0c8d-4e7c-8256-b81a4a0d4b11 · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.277778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:48e16d5232df1fb57a34777ad71a919ed86eb3d4587063d21fd4534d67983211

Observation a8204efc-3275-425c-b072-f346397254a0 · outbound

This paper cites Spice: Semantic propositional image caption evaluation, in: European conference on computer vision, Springer.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Spice: Semantic propositional image caption evaluation, in: European conference on computer vision, Springer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.286591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:d2a4e101b24bff1fdd3040066b249fa3ae6bb9b2ed7ea57dbd27102673a4d526

Observation 7d62fa39-4409-48df-8b4a-4e9159c111c0 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.299263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e31f08f0861b444b4ab83e92d3aca1ad214b29ba9af5af05d36fa75fbfc1923b

Observation 5609b157-c88b-48ac-a63b-e511b5f6377f · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.272360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e02fb8407d02148ba9e923d771d211e8154bd1302a51b04349cc45cc4ad7d75a

Observation c04100e3-cd8e-4ba3-a284-21d5c58c1a73 · outbound

This paper cites Context-awarevision-languagemodelagentenriched with domain-specific ontology for construction site safety monitoring.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Context-awarevision-languagemodelagentenriched with domain-specific ontology for construction site safety monitoring

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.267250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:a86ceb3e3d9696f96bc7da18b943baa41917c4a8e848cfc4fc37266e3d59e973

Observation db022c64-081e-45f2-8731-59aa3cee9036 · outbound

This paper cites Enhancing vision-language model for construction safety inspection via visually grounded reasoning and reinforcement learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Enhancing vision-language model for construction safety inspection via visually grounded reasoning and reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.274283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b7cfa91a98b54bbc7ca862b4cba9eaec228cbdeb323817008d18efc55d764817

Observation 15e6ae3a-8d24-4bda-a6f4-97f1cf8b41fa · outbound

This paper cites Augmented reality, deep learning and vision-language query system for construction worker safety.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Augmented reality, deep learning and vision-language query system for construction worker safety

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.449701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:1eee950d2a7f2784593e4ed2b0fc36e83c93f3e1a42b69cd3d9e94075f813fa3

Observation b97a474a-15c1-478c-a0d9-7f4144f25c6a · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.588253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:089d843d05392cd924e8d30a3b828a4e7e86613a98041d25232fbe07f20eea83

Observation a8b5b3c8-cc53-4915-913d-0bbd6b662b3b · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.573081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:07d5a96903bcb7aaae77a092327e8de1507f0e6d2d8ba0d84eac0102dbedb90b

Observation 50a78136-ca05-4dfb-b2c1-b1ec9f796d1c · outbound

This paper cites Tailored vision-language framework for automated hazard identification and report generation in construc- tion sites.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Tailored vision-language framework for automated hazard identification and report generation in construc- tion sites

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.452922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b30f43c48c1293ffe5473b328b8f995b0d0280b15dba841670abad117ea92cfb

Observation 79dd781b-948c-4f84-b7e2-c20924fd489e · outbound

This paper cites Canmultimodallargelanguagemodelstrulyperform multimodalin-contextlearning?,in:2025IEEE/CVFWinterConferenceonApplicationsofComputerVision(WACV),IEEE.pp.6000–6010.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Canmultimodallargelanguagemodelstrulyperform multimodalin-contextlearning?,in:2025IEEE/CVFWinterConferenceonApplicationsofComputerVision(WACV),IEEE.pp.6000–6010

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.311086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:f9ee582d81353b43ecee902f6fc45e3dcd5d4769adeb4bdc4b85fb2cab270385

Observation 1db83eaa-394d-4813-a22f-8ed71a09fd3d · outbound

This paper cites Are large pre-trained vision language models effective construction safety inspectors.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Are large pre-trained vision language models effective construction safety inspectors

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.439667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:9e9ce0e04c0a88f49cfb46d99abdd735c4b750af01ce2a43d8e18a3bd4719722

Observation e8539243-9771-4b00-9826-177b464188bd · outbound

This paper cites Expandingperformanceboundariesof open-source multimodal models with model, data, and test-time scaling.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Expandingperformanceboundariesof open-source multimodal models with model, data, and test-time scaling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.308515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:2eac76166973e54cd8f1e0e4060cd41a09d88bcf7740cae277f329aa5e2a4122

Observation 41ae7e5f-ee17-45c6-a353-c484f16f95ec · outbound

This paper cites Chain of Thought Prompt Tuning in Vision Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Chain of Thought Prompt Tuning in Vision Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.569827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:f83ad7ef60fabba8d89614ef3c5067586c034c2cda7819964717642800f42f73

Observation 1d220ae5-72a7-43e9-8af5-4a2abd499063 · outbound

This paper cites Intelligent virtual assistants with llm-based process automation.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Intelligent virtual assistants with llm-based process automation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.313135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:86a33e0a3e384536ecbc6e8ac3b92215f9736e7b15acfb6fbd9ff8c1a342db72

Observation 64dbea47-ccd8-4b3e-a774-abadb1139a23 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.595517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:03a5f8aad5c3ef804b161d241edc02eec18ea8afddceff76e18b93deb2368e6c

Observation 34e022b9-5582-4941-837f-beb8b64786bd · outbound

This paper cites an unresolved cited work.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-07-08T22:35:41.306558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:0c996897ae12f9e19107512c7edd7f991778079abe912185e8a40dc683aec99f

Observation 8cd8e49b-6f76-4ef8-9070-523d55ac4409 · outbound

This paper cites Vru-accident: A vision-language benchmark for video question answering and dense captioning for accident scene understanding.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vru-accident: A vision-language benchmark for video question answering and dense captioning for accident scene understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.302910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:59032f00bb1fd9f8252392f668a56e0a46baac6683974e76c0842f86cde245ea

Observation 39d9e36e-2391-4d2c-8153-146d5dd508ea · outbound

This paper cites Safe-llava:Aprivacy-preservingvision-languagedatasetandbenchmarkforbiometricsafety.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Safe-llava:Aprivacy-preservingvision-languagedatasetandbenchmarkforbiometricsafety

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.601075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:2d3ae369382fe8d1e4afa84ec5bd2c39b7e0e0fd1ccaeea856ac37f268eac07a

Observation 33641ca8-8c68-45d1-b53b-fddf2cd063c1 · outbound

This paper cites Res-bench: Benchmarking the robustness of multimodal large language models to dynamic resolution input, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Res-bench: Benchmarking the robustness of multimodal large language models to dynamic resolution input, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.304705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:2c6533d8ec1aea6bdb3e1bbe56367daf1fe99391fb74efceb92a2c457be36edc

Observation ea9338c7-0f61-4527-99cb-c07ee5d7f74b · outbound

This paper cites Construction site fall hazard identification and automated captioning using adapted vision-language models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Construction site fall hazard identification and automated captioning using adapted vision-language models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.467360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:177c404c9a96ed64f53fee4faf514ee4561e4ed81b002a4b68af54c8a9ddbc87

Observation f3830f6e-72a3-44e8-b6b8-88a0da821de7 · outbound

This paper cites AdaptVision: Efficient vision-language models via adaptive visual acquisition.arXiv preprint arXiv:2512.03794, 2025.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring AdaptVision: Efficient vision-language models via adaptive visual acquisition.arXiv preprint arXiv:2512.03794, 2025

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.551322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:321f5e663714bb72c1e580767ed275096c86f5268e0bd7883f30d1e7ba6957bb

Observation 3761a0a6-e163-4383-8877-fd02b169a6b3 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Llava-next: Improved reasoning, ocr, and world knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.314843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e024332ee05843129f730d8e0dbda09e88b938e54549ed12dc29b6eceb834cb7

Observation 40e2f463-7018-460c-b20b-fa0004ecc42c · outbound

This paper cites Visual instruction tuning, in: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visual instruction tuning, in: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.301085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:933576781468fb878e824511f11248e3cdab8afdf365dd1a06251e3ab6b1c077

Observation 0782e0f5-ef7b-4b50-bb0d-c21cd184fb6a · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.580445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:a2797341b0cb113b515a8747d3cb04b959f73d367bf8679ad9a787b67921e22a

Observation 675779f7-775f-42f0-8e63-74351b6ab11e · outbound

This paper cites The Llama 3 Herd of Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring The Llama 3 Herd of Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.575764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:28c86404c6b7f8b3f47511a1f674f2e627316c999aa7be9c739a2fc832c8aac4

Observation 7b142371-113d-4e6a-bd64-fcc9f6097183 · outbound

This paper cites GPT-4 Technical Report.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring GPT-4 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.578047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:25892381ffb70f65bb6720aadb81f78bf3a0d84fa40ec6c38536b860409ce9ab

Observation b911e575-3131-4eb4-a6d0-3ddc7845e535 · outbound

This paper cites Privacy-Aware Visual Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Privacy-Aware Visual Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.590687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:f95b137436bda44913ef0081fc1baa9f6999fe976de27ec1076198eb40e166c3

Observation 81a41807-1877-427e-a3fc-02ad0209459e · outbound

This paper cites Vlm-robustbench: A comprehensive benchmark for robustness of vision-language models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vlm-robustbench: A comprehensive benchmark for robustness of vision-language models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.603788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:425a587a349e0e4754ab5e0457cc3d1a2a4353abbc81a63aa3758fdb51519863

Observation fa13d331-0a3d-4d9b-a0a6-27624befe93f · outbound

This paper cites Vision-language models for edge networks: A comprehensive survey.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vision-language models for edge networks: A comprehensive survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.292231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:176e3da7e60df6a2447b18166887ee22d35fc4d21ec32235ff9d76f31b340285

Observation 4ef34baf-f4b8-43b6-b59e-71e9e1cfc553 · outbound

This paper cites Upop: Unified and progressive pruning for compressing vision-language transformers, in: International Conference on Machine Learning, PMLR.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Upop: Unified and progressive pruning for compressing vision-language transformers, in: International Conference on Machine Learning, PMLR

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.297412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:4bd6356307f9626c7aaa494380b9b427aa3a34f17d75b0d42ad5e3eda901748b

Observation 01db1329-854a-4e32-a16d-f8f63d9592ac · outbound

This paper cites Real-time safety detection on construction sites using a vision-language and nlp- based model.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Real-time safety detection on construction sites using a vision-language and nlp- based model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.446394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:96b8b700e758ac3da123ad67873dfe97c329e29870600193d92246a1ea77989a

Observation ef1d064f-93a1-4846-9930-8abc511a1541 · outbound

This paper cites A double thinking enabled visual language model for open-set construction site safety inspections.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring A double thinking enabled visual language model for open-set construction site safety inspections

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.290434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:491762d7621404aa89fdc8647a4d5db3e72ef83a41f1b696dd9da7f4cf2e9169

Observation 5a347992-f026-482e-9bb5-576ef879123e · outbound

This paper cites Region-level vision-language model for detecting distraction behavior and mobility attributes of vulnerable road users.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Region-level vision-language model for detecting distraction behavior and mobility attributes of vulnerable road users

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.461810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ef635066d9ac7dceef4c41dd0d675d9cc8450c08350769c8c974d652115ee373

Observation a9ca898f-8578-4960-ac15-8a82a00adb27 · outbound

This paper cites Visual question answering-based referring expression segmentation for construction safety analysis.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visual question answering-based referring expression segmentation for construction safety analysis

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-08T22:35:40.458666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:9945cd3fca3ef2fb303891b012488130e4c99368a8c361cb4112c6fd224c5101

Observation 03238cdb-052e-49f1-8880-de628907165f · outbound

This paper cites Selma: A speech-enabled language model for virtual assistant interactions.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Selma: A speech-enabled language model for virtual assistant interactions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.295114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:7b5c6d17ab9cdc4aeeaf6dcbb7ddc7797eb15d0b19e7f28ac03079cb7647c6d7

Observation 34c9128a-0134-4b53-9856-11c20b6f6627 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.288555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:a1c6647b24417d831e9e69cd107a3b601a1e2f483ae712d4bae390b84b342891

Observation b89f6176-1db0-4b14-99ce-d82dc42d1d68 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.582870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:b85e6fb39f3b8d1c52ef3a2b1141b22e4b005e02b881cd9c4062e0e6a7118731

Observation 845780f1-5d67-4024-8896-568df97ed835 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.609039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:5fa3b1e4e7585200495bb53c3a81c30b78cde398ed7e3c0c791e7eb7c183bf2c

Observation 55cfce7d-25f9-42f7-ad8f-44484ae995b2 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T22:35:40.606423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:3cfa6b0cdfc8cd2708ad1e255c77ea85c1de45de7306a2f7a6c8b90ad267161f

Observation 309f9b4b-22ac-4542-9287-bb29ee9c9b1e · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step, in: Proceedings of the IEEE/CVF International Conference on Computer Vision.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Llava-cot: Let vision language models reason step-by-step, in: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.281322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:d19fefd8da4cb128e2fefc03e5c88e1e35ec0d8f8a9a9dc9104b4431602c3173

Observation ba0ffa3c-a517-4530-aa69-323d59955cd3 · outbound

This paper cites Edgevideoanalytics:Asurveyonapplications,systemsandenablingtechniques.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Edgevideoanalytics:Asurveyonapplications,systemsandenablingtechniques

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.282854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:1dcc4db6e3332dda446d5f31ee6c1525e3633f55aecd1463ea11a9c5fb65d7a1

Observation b0423c76-7678-438b-bf77-54c87ea12c95 · outbound

This paper cites Qwen2.5 Technical Report.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Qwen2.5 Technical Report

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.585730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ab325147faa251bcc450867490739e624621690c35e423b8594662356d76103e

Observation 1c8752b1-5b71-4006-b0fb-0b9598995377 · outbound

This paper cites Vision transformer-based visual language understanding of the con- struction process.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Vision transformer-based visual language understanding of the con- struction process

Reference 45

Resolution
verified exact
doi, observed 2026-07-08T22:35:40.464393Z

Source-reported events for the cited work

correction dated 2024-05-21. Source: crossref record 10.1016/j.aej.2024.05.064->10.1016/j.aej.2024.05.015:correction, observed 2026-07-11T03:14:22.632143+00:00. This notice travels one citation hop only.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:ef24dcbd83e6204b3642e7d585b04177374a268573a3a82beb59b4a600b56665

Observation d220c57d-6314-422a-94b9-67342edbee60 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visionzip: Longer is better but not necessary in vision language models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.284722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:09468599fff59a88513192441a60d431aa684aca49fdb158fb41657e861de83f

Observation 3f04773a-9683-4cc4-9277-5bd91c9cf822 · outbound

This paper cites Visionthink: Smart and efficient vision language model via reinforcement learning.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Visionthink: Smart and efficient vision language model via reinforcement learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.279512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:2b125b5028ccd3f98286c006916503544b0d19637d34955863602d1c40f23cdd

Observation aec977bd-d0e7-416b-90e9-f9dca407730b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring BERTScore: Evaluating Text Generation with BERT

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.611532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:7ae47562c7424cef8fba08596e37990ff4468251b6ecc2cbf42ad1ec5ed25400

Observation f908bd59-8cb1-41b8-b779-8b2e0a937fef · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.598163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:31fc02a9dabd3796e9c579926d9bade1a0a71e0316204bf9e865c3c37a631f9b

Observation 01c5da41-714e-464d-9706-0f192a0a4c77 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Multimodal Chain-of-Thought Reasoning in Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.593114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:3ae87dec489dd1fd989d45a5646910b5bb170ebcb8d571edf085ef5d1624373b

Observation 370f17ac-fcac-4490-a6de-21e2d736c31f · outbound

This paper cites Mmicl: Empowering vision-language model with multi-modal in-context learning, in: International Conference on Learning Representations, pp.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Mmicl: Empowering vision-language model with multi-modal in-context learning, in: International Conference on Learning Representations, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.276107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:e6205067f0c58f519eaef31b94be4982cad316e7ec9e34f5e29963c20fc8c1e2

Observation 31197953-97f3-4862-a023-8cb7add18a2a · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models.

AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T22:35:41.269087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:27:52.861022Z digest=sha256:88a68d8e8fdaff95e640f419e85fc939fc06cbfa5455e88ce8cb197b1a9c6c15

Pith citing papers

No inbound Pith citation observations are available.