Pith. sign in

Paper Citation Record · LEDGER

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

As of 11 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2412.20750.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20750 v4

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:16:55.382256Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T12:45:30.225562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a8b3627c-88c6-4762-9d1d-b6980f7ba249 · outbound

This paper cites Grounding language models to images for multimodal inputs and outputs,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Grounding language models to images for multimodal inputs and outputs,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:58.164424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.817302Z digest=sha256:89b3bcf69a46f33c4213a23894bb083705c32f2fcd0bb279958711e23d65761c

Observation 34fb1541-20be-44cf-9e30-fe2173d468fb · outbound

This paper cites Harnessing multi-modal large language models for measuring and interpreting color differences,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Harnessing multi-modal large language models for measuring and interpreting color differences,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:58.125905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.829252Z digest=sha256:a9bfd810d0044499a3ec210b3212bcf70fcb92e9506701aff7135bad77b52a7e

Observation 5cdae950-263f-4a49-b112-ec043aeba294 · outbound

This paper cites 3vl: Using trees to improve vision-language models’ interpretability,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking 3vl: Using trees to improve vision-language models’ interpretability,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:58.092900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.846285Z digest=sha256:be3c9ba3523f505cb6f5e55cfb7409064b8c18cb8011bf5ea35940c7e9da2e56

Observation 34fd085d-92d3-4fa6-9b8d-3374656b2025 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.852567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.852567Z digest=sha256:a8a3cd2f4016202ddfb50053eee7cb5ede3d85cba0761b4eb4a67e0d8312aecd

Observation b9da80b5-c261-4bf3-95d3-b5490be68ef4 · outbound

This paper cites Who, what and where: Composite-semantics instance search for story videos,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Who, what and where: Composite-semantics instance search for story videos,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:58.048530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.860925Z digest=sha256:7b2ee2ca2bee657f290f0f8a7f8e8cf0b8fc33f58152113e5f1140db2267f347

Observation 37ef8cd2-dd4e-4e4a-8328-361fc979194c · outbound

This paper cites Exploring language hierarchy for video grounding,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Exploring language hierarchy for video grounding,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:58.022945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.874996Z digest=sha256:7378073b907b89671b64ed451f105c2475eeb8629cc2732079585dab65188a68

Observation ebbac2db-1188-45b3-8baf-01a70e9ec4dd · outbound

This paper cites Adaptive spatio- temporal graph enhanced vision-language representation for video qa,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Adaptive spatio- temporal graph enhanced vision-language representation for video qa,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.984184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.881065Z digest=sha256:e5d97c7cd7876685c79a9212b43b6b71653300afebe1f79232ee458b8d3f4b93

Observation 1b457a58-b9ae-41d7-981f-999f974956ae · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.889462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.889462Z digest=sha256:26005545e58e97da1e8ecd39a1eec211c32cc0d15109da7f99114ed8c0e6a80e

Observation d40b7757-eb2b-4bda-a045-1ef96cc88d92 · outbound

This paper cites (2024) Hello gpt-4o.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Hello gpt-4o

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.943791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.897037Z digest=sha256:8b07f19e046da1747563d8fc1850a391cea1ac29fa22f0bd79b2d68c107fe810

Observation 232a566b-4aac-48b3-87f6-bbe0d0014d31 · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking GPT-Driver: Learning to Drive with GPT

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.903733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.903733Z digest=sha256:bf548770325386dcda4a3f8b1be79e5a0c3ead4ed082c6d207c9a05612a3f82a

Observation 73ed88e7-3cc0-4bc7-b6d4-ce3883fa8357 · outbound

This paper cites Drivegpt4: Interpretable end-to-end autonomous driving via large language model,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Drivegpt4: Interpretable end-to-end autonomous driving via large language model,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.910154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.910154Z digest=sha256:de9ccfc00754a4a42d50d85f78fc67500dd1595463ac344e73a3887ea7689b7b

Observation 0eaff5e0-3299-4960-b2b8-eab6c88219d4 · outbound

This paper cites VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.921826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.921826Z digest=sha256:b766d91b5bdf79cc06b0c54ffec2f683dfe3064953374186fd76dc2d9ebaede2

Observation 92555851-77a1-4943-8aa5-2e7b98cbab29 · outbound

This paper cites Langloc: Language-driven localization via formatted spatial description genera- tion,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Langloc: Language-driven localization via formatted spatial description genera- tion,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.929466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.929466Z digest=sha256:3fdcaef91c36de35aad4e4a3eaee7cad104e3aee2912adbfffef06b1a1120fdd

Observation 15f3a64d-6f05-45a2-b8ba-1a8f98430e2c · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.937962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.937962Z digest=sha256:1ede7050d84c3d5b79f6f5238d333b6cc3ea65cd23817726ab76f0cf6c98569b

Observation 32f1cd6b-e19c-4de5-ba22-83560400bd9f · outbound

This paper cites Trafficvlm: A controllable visual language model for traffic video captioning,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Trafficvlm: A controllable visual language model for traffic video captioning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.861854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.945280Z digest=sha256:0356d15ae6abe261a953eabf118caf6ade09d646a767be083cc7e4c742cc3469

Observation 1593e766-91e6-4648-9c6d-3ee1a24a3155 · outbound

This paper cites Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:16:56.461973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.951087Z digest=sha256:2dfe5d91787b3239230b061acfd11747caf8d1a3054eb04a0bf8c10d57e438c4

Observation 8f4e9953-c77b-4c0a-8e50-446a82309b2f · outbound

This paper cites Physically grounded vision-language models for robotic manipulation,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Physically grounded vision-language models for robotic manipulation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.960802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.960802Z digest=sha256:21ce5af0a765d94803904a62901da59b3a6a6c17bd22c647f405ddaffa9faf10

Observation 331daaf8-c9f6-4a17-bcdd-dc080a59cf89 · outbound

This paper cites Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.816285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.968566Z digest=sha256:ac701d5a9070d135a16b828255e05b488bbe44b8ab15d400bc757a54de2742bb

Observation 8565937f-d623-4d14-a8b7-02b0711b903a · outbound

This paper cites Aha: A vision-language- model for detecting and reasoning over failures in robotic manipulation,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Aha: A vision-language- model for detecting and reasoning over failures in robotic manipulation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.796007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:54.976759Z digest=sha256:d93c5908ccf196768f27ec3bc13629aac7381569542a01b5bc12dd748d6c6619

Observation 1c5ee29e-1ccd-49a0-9ad1-fff95276bad7 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.991648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.991648Z digest=sha256:8b40e0fcbd72bce7fe5d0448a91041a2db3b00bb2355fccb985954d583592e5f

Observation 12cf8a70-2080-4615-be78-b547d5fe1f87 · outbound

This paper cites A3VLM: Actionable Articulation-Aware Vision Language Model.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking A3VLM: Actionable Articulation-Aware Vision Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.998856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.998856Z digest=sha256:a4fe483a6122da21a569624b5095e4983046889a2b6fc62bfbce923570bca957

Observation 26bccfe0-b39e-4144-b01e-81b6c84afce5 · outbound

This paper cites (2024) Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.769303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.005244Z digest=sha256:79826f807f808e8f7333047ed62652fc732db5cb08303e9487009ffc20e0dc4a

Observation 2a98c0c7-f0df-4744-8d4c-1d736d6dbbae · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.011356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.011356Z digest=sha256:1fc2e32e97c4eb45a382bae9f1334933087335426ffef5dc9043bbf1ab32d078

Observation 097e5731-e685-468c-baa7-10263cb9ef6f · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.019857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.019857Z digest=sha256:7b5880478216deb43d016fc4ea76721adc12baea4485b54c5eee2a5277a156c3

Observation 6f6e5f1d-17d1-497d-bb4a-65b655c0da63 · outbound

This paper cites Onellm: One framework to align all modalities with language,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Onellm: One framework to align all modalities with language,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.722290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.029173Z digest=sha256:7905ddeffbcf488b0de55302eddf61e4f7c3321f935693f0916b8c06b0d5eb79

Observation 2caa71ea-ad37-41d4-b6ec-05f4594289c4 · outbound

This paper cites XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.036897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.036897Z digest=sha256:4e38cbd551c57d3a5ca7d3706e560f498448faba267d661d727b2588bb7dcda2

Observation 56071ecc-ee84-4dbc-935a-2112c96abf9d · outbound

This paper cites LLMs Can Evolve Continually on Modality for X-Modal Reasoning.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking LLMs Can Evolve Continually on Modality for X-Modal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.047443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.047443Z digest=sha256:0c5fc16bdcc91684badb811bfec685bb204a52cfc1607fcec15c99cead9fb230

Observation 011b1509-cc5e-4774-93f0-fac0bc8c434e · outbound

This paper cites Visual instruction tuning,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Visual instruction tuning,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.057153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.057153Z digest=sha256:5ff0f2b4666ce35f159eb4d8d3b1262e137a0008a551bfa272a77cff7f992a13

Observation 9a910b6a-eb0c-4b64-9019-5d0b62e8c8c7 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.062832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.062832Z digest=sha256:93b35591d1bddb61ee35844771c2fff044ab7e116330ab3591b26efc2043efb6

Observation a766d814-6adb-49c3-8ac9-4fb8d8230c06 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.070096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.070096Z digest=sha256:66dd068202cf7d877746110e702492ab8c494deedf5cbea084e6330c06b18b21

Observation d654cdf2-d889-4af8-a1f9-70ab78727303 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.078302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.078302Z digest=sha256:a7c2b6c2424e5596f20caa95b9a4942925c5c038dd69211c89675886f0924b8e

Observation 2d55d716-a0c4-454c-a509-2261670d070e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.085117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.085117Z digest=sha256:3ff34304165968f5a3e2fc219ee2233fe06b31dc73d2d90f457c96d3c9f3a3ae

Observation d3a49532-c2e3-4eab-9ce9-20031c095353 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.090594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.090594Z digest=sha256:967ec894ab1fbbb912307a16513c5bb0c02955aa239e0e099301ef78316775c2

Observation 9a5667f3-d22b-4c02-9847-bcde6d26da28 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Imagebind: One embedding space to bind them all,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.101354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.101354Z digest=sha256:b1c398dc8f03b0d882b7f287af81921346c314de15e3b34f2e3316c35431255b

Observation 8947ba36-54a6-4d2a-9a9f-eba3a8c9aff1 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking PandaGPT: One Model To Instruction-Follow Them All

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.107696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.107696Z digest=sha256:d5ddb039de5f2b0f14b52770fdfaf90640e80c66629ec7924b8d7bc7f9ab7706

Observation 8f326b97-80ef-48a9-810a-cb0f4b893287 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.114595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.114595Z digest=sha256:c5337a45ece3a097081b2a549da00037764f53fde8f184a3501c26a52d74a5f7

Observation 8dae310a-f67a-4815-8c63-57486054489f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.121002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.121002Z digest=sha256:3cd9d48a0ef4d2fc1276ae4f6d106893f136fcf5f393f08c44d3e3060c81275d

Observation 437c7b22-36b9-4d91-ac02-80a5c55bd0f8 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.131639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.131639Z digest=sha256:5be27ce02a4400d26968ddfa236f839d02bc905410dfdb9ba06803a89f5a081f

Observation 75da596b-52c8-430a-b37f-a9537e2fe48a · outbound

This paper cites Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.136845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.136845Z digest=sha256:88a54a88b4e2f4094d341d16687cd169c703e9b697e099b5affdf5936e4510bd

Observation dc8e5330-580e-4efd-b83b-9dd62650d1d2 · outbound

This paper cites Q-bench ++: A benchmark for multi-modal foundation models on low-level vision from single images to pairs,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Q-bench ++: A benchmark for multi-modal foundation models on low-level vision from single images to pairs,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.564214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.144635Z digest=sha256:97007aa5de24386bfb44e6c6f936919d31043dcfeeeb39dd8d321a62310c58d6

Observation f4f689a7-7995-4b97-aad6-241701e7b7cf · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Mmbench: Is your multi-modal model an all-around player?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.525800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.152443Z digest=sha256:80533dcaca9baa2e8c672ec58f8c7074a196b6f7831fe71498640495b0891d0f

Observation 8e21972b-64ef-42e6-a36b-2f7ff72ae059 · outbound

This paper cites The reviewing of object files: Object-specific integration of information,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking The reviewing of object files: Object-specific integration of information,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.492861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.159331Z digest=sha256:356acdbb4e358aa238a43761babcedfae349d4be94867e984d72b924ee53b357

Observation d576d43a-2cf5-4554-b24d-a20e6b4fc5cf · outbound

This paper cites an unresolved cited work.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:16:57.449205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.167390Z digest=sha256:95e7a7892a2d6850d6e6a0471f5c1fe04e07bd0ce112d97d6a2c3adec6b333ee

Observation 3c8d7df6-dae0-47b0-9658-f395710eee33 · outbound

This paper cites Improved baselines with visual instruction tuning,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Improved baselines with visual instruction tuning,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.173811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.173811Z digest=sha256:6cf80f120b6d46a58c45c17557a321205ea87cc448b2ef46c32312d628ac0165

Observation 2b7e5b1f-8e55-450c-9c7b-1ecef7080308 · outbound

This paper cites Phantom of Latent for Large Language and Vision Models.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Phantom of Latent for Large Language and Vision Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.180265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.180265Z digest=sha256:680375683c0c58390ace742b79e292777b466562dd43864459fb93b54bef02fe

Observation ffb088be-550e-4ff7-bf85-7235199d2b28 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.186893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.186893Z digest=sha256:c8b1aadc93be7514e1f71f65565c5317f224f906daf92dad500fdda1b146eab0

Observation 417d6d1b-aa34-4670-bfa6-9e1e59370ffc · outbound

This paper cites (2024) Claude 3.5 sonnet.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Claude 3.5 sonnet

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.374510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.192741Z digest=sha256:5d927838b3f2e3e7a68038dae8be33f6e8b7576e8a3b36862aac0d99b753c7bf

Observation 6ee7360f-e93a-479d-b2d2-b9cec25375d7 · outbound

This paper cites Prolific,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Prolific,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.321146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.198233Z digest=sha256:f3960a516702201bba06870fce6e8ffe0a168904aea833e56646c3bff39fb6a8

Observation f6b04e6f-437a-46e7-bf60-a6bfe513fb61 · outbound

This paper cites Rank analysis of incomplete block designs,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Rank analysis of incomplete block designs,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.269156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.203398Z digest=sha256:f5f529b15d1e322427fc767f4c99bd9a8d2af8c3ac28f5e50dc97807e4dc99b3

Observation c37c8ccc-8eef-4408-8042-c81a1192b376 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Direct preference optimization: Your language model is secretly a reward model,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.209582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.209582Z digest=sha256:7ab0d6e66f4778b558c62ae31ee1388c59ed044b7d1f458a6294fa271dfbe625

Observation c296a579-0808-42f1-af43-704d1dc5b6ef · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Facenet: A unified embedding for face recognition and clustering,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.216406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.216406Z digest=sha256:b786a949750369980a81523e36fcd15a43c820a772e39b06bc08cd428a137e5e

Observation 6595076d-124d-4d5b-b1e0-105fc15c5ad4 · outbound

This paper cites Target-aware dual adversarial learning and a multi-scenario multi- modality benchmark to fuse infrared and visible for object detection,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Target-aware dual adversarial learning and a multi-scenario multi- modality benchmark to fuse infrared and visible for object detection,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.200496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.222315Z digest=sha256:d579251313d986083bd89b897bf74fcc6905145d28d174b8cf14e6877977a1c2

Observation 827cef51-6cd5-4e73-a4c6-5581618800d3 · outbound

This paper cites thermal dogs and people x6ejw dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking thermal dogs and people x6ejw dataset,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.160227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.228568Z digest=sha256:93d21905f558223e184a3b0b253c77714fc8f0dbfb133dbd354d1146052ed384

Observation cb195b62-1dd8-49a4-9e2d-5f30423b4bdc · outbound

This paper cites pet dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking pet dataset,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.122467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.237447Z digest=sha256:697ae77a41f82c421811aa1e0b2f8e2aafc950ab4c056e5ba6c4e458a626b3e9

Observation 3b3d3338-35a0-468a-813d-bdb3a674873e · outbound

This paper cites Thermal dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Thermal dataset,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.085566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.249014Z digest=sha256:cd3b161899bbd6c757820d2d6ba6b70176b6f3c5321c634284852f3fcc464144

Observation 085270e7-7557-435a-9082-858b224f95e7 · outbound

This paper cites Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle-based object detection,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle-based object detection,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:57.031211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.256113Z digest=sha256:f0bb4c0060529642a7da1b78b678359fd6e94a454253a66e094d63eb3411c6d6

Observation 7f013fc6-b02e-438a-8d2d-ec869489167d · outbound

This paper cites animal-detection-flir-extra dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking animal-detection-flir-extra dataset,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.997919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.264251Z digest=sha256:23d93432942ea1e97aab7963186378b1224cf63d28864e910cbd72fd08ba1619

Observation 2a9907d9-f032-4e54-a59b-4b17e12e7ab3 · outbound

This paper cites chips-thermal-face-dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking chips-thermal-face-dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.910171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.279065Z digest=sha256:d23124f75db3925a6a8440add94627d69709dd355b8ad514c4c4e501fa2030ce

Observation d245561d-3894-4cac-a36b-f8ce905dad9a · outbound

This paper cites Available: https://universe.roboflow.com/one-rphct/animal detection flir extra.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Available: https://universe.roboflow.com/one-rphct/animal detection flir extra

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.950577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.272591Z digest=sha256:90b6bde59747936f6e850a421166e61714ea7712e054af28d4e6d7008a30a646

Observation 7495a38f-2956-49d0-a2e0-08aa4fcb7e17 · outbound

This paper cites DIODE: A Dense Indoor and Outdoor DEpth Dataset.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking DIODE: A Dense Indoor and Outdoor DEpth Dataset

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.294746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.294746Z digest=sha256:c1d8fa22f80f56eed53c8b8ed2846f960a15fdbf632685b5671c0b7300620656

Observation c7781007-900b-457b-a252-e5860655d16d · outbound

This paper cites Ifsod dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Ifsod dataset,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.876035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.285182Z digest=sha256:78258e8feac95a3d07e810c7bd6066bc6ac79c5ac5a703c49512e3018f268f08

Observation 899ecfa9-da76-4845-95db-fd153375aa7a · outbound

This paper cites DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.307903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.307903Z digest=sha256:a05c70efe4c3c4f77a00d047a8ce769966b0fd3f68cb6e7cc3f4bb86dd7f8017

Observation 79fe59f6-426b-4a08-9e9a-e187715d8902 · outbound

This paper cites Indoor segmen- tation and support inference from rgbd images,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Indoor segmen- tation and support inference from rgbd images,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.837748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.301879Z digest=sha256:678708c7af07c4a7f14b0012a800fa1cca7c63987f29bff74cccf34bdf813a41

Observation 834b060e-a989-47c2-859c-a48e126ab185 · outbound

This paper cites X-ray baggage detection dataset,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking X-ray baggage detection dataset,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.764582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.321174Z digest=sha256:120614cf232c26e4f3a9c155bae68ab28a9a6cc953bead3824fec972fb8ba2ef

Observation 51910820-87e8-4fad-be47-580911fab010 · outbound

This paper cites Unifesp x-ray body part classifier compe- tition,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Unifesp x-ray body part classifier compe- tition,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.807204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.314125Z digest=sha256:e11e2c3e6a6d389dba2a5e5bfe9d60db99676208d691ed13d722fa5b8884f351

Observation 7124f7bc-a881-465d-93bd-b0fae4144252 · outbound

This paper cites Decoupled weight decay regularization,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Decoupled weight decay regularization,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.338072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.338072Z digest=sha256:9eacdbaaf9ab4db3ebd17580e89ef4b8a3779ef84a7f2f6ee230cdd3e0dead22

Observation 82c877dd-628b-43d0-8e0f-de0e4f9a1a34 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Qlora: Efficient finetuning of quantized llms,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.327101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.327101Z digest=sha256:6f730ac20044ea873d784686ba7983d4fb3c526cef117c44ab0276004946a897

Observation bfb8b9a2-dfad-482d-80c0-6fe951c4ab8f · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.365253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.365253Z digest=sha256:5c1faeace9e0d91688f73f1cc689bca4f2805062a21a0c6acdf53195b832eb63

Observation 7dc168b1-45e2-415a-ab31-3d680cd248b3 · outbound

This paper cites TE-SQ1 Thermal Camera,.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking TE-SQ1 Thermal Camera,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.667129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.372387Z digest=sha256:e9ec85c1fac2017187fb4e9498ece6ce40d8a9fd3c22822ba28518d35d3d9bf7

Observation 271cf759-74ff-4e8c-bbca-483ebcf9d94b · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.356697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.356697Z digest=sha256:c2d23ed75f76118050fed76bb3fe8614275a8d7c395a9d1f451f63f05a1c88fc

Observation bcc74ea9-23cb-411f-8f15-356748985b2f · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Decoupled Weight Decay Regularization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:55.350018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:55.350018Z digest=sha256:90131057ab9a9dc2c0a9355666c87026c68278b5051b3bd3aceda484449ea406

Observation 8b0f47f1-e734-49e7-9a52-9a4dd10329cd · outbound

This paper cites AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T23:16:54.983609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:16:54.983609Z digest=sha256:4aa972131c1b6a9e776d8539ee0f43b9f920010e776aa27c57312076a4e64a13

Observation 8d5cceea-2470-40a4-93a4-58f8a999e59d · outbound

This paper cites His research interests include deep learning and multimodal large language mod- els.

Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking His research interests include deep learning and multimodal large language mod- els

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:16:56.632139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:16:55.382256Z digest=sha256:d1c8b05517c1a944ecaf0f19b91b5e6d2cf544c6ac9478fa51e17715622124a3

Pith citing papers

Observation 6f1b9142-81fb-415b-9c86-3ec3c2174556 · inbound

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis cites this paper.

Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T12:45:30.225562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:45:30.225562Z digest=sha256:441ff7cb4d8a853703197ce571303455efcf8f6f55457ab7c0e9af2589d65bf8