Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:16:55.382256Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2412.20750.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:16:55.382256Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T12:45:30.225562Z
A source-named dated measurement, never combined with another source.
Source: cited_works
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a8b3627c-88c6-4762-9d1d-b6980f7ba249 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Grounding language models to images for multimodal inputs and outputs,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34fb1541-20be-44cf-9e30-fe2173d468fb · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Harnessing multi-modal large language models for measuring and interpreting color differences,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5cdae950-263f-4a49-b112-ec043aeba294 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking 3vl: Using trees to improve vision-language models’ interpretability,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 34fd085d-92d3-4fa6-9b8d-3374656b2025 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Timechat: A time-sensitive multimodal large language model for long video understanding,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9da80b5-c261-4bf3-95d3-b5490be68ef4 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Who, what and where: Composite-semantics instance search for story videos,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 37ef8cd2-dd4e-4e4a-8328-361fc979194c · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Exploring language hierarchy for video grounding,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ebbac2db-1188-45b3-8baf-01a70e9ec4dd · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Adaptive spatio- temporal graph enhanced vision-language representation for video qa,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b457a58-b9ae-41d7-981f-999f974956ae · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d40b7757-eb2b-4bda-a045-1ef96cc88d92 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Hello gpt-4o
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 232a566b-4aac-48b3-87f6-bbe0d0014d31 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking GPT-Driver: Learning to Drive with GPT
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ed88e7-3cc0-4bc7-b6d4-ce3883fa8357 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Drivegpt4: Interpretable end-to-end autonomous driving via large language model,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eaff5e0-3299-4960-b2b8-eab6c88219d4 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking VLM-Auto: VLM-based Autonomous Driving Assistant with Human-like Behavior and Understanding for Complex Road Scenes
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92555851-77a1-4943-8aa5-2e7b98cbab29 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Langloc: Language-driven localization via formatted spatial description genera- tion,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f3a64d-6f05-45a2-b8ba-1a8f98430e2c · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f1cd6b-e19c-4de5-ba22-83560400bd9f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Trafficvlm: A controllable visual language model for traffic video captioning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1593e766-91e6-4648-9c6d-3ee1a24a3155 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f4e9953-c77b-4c0a-8e50-446a82309b2f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Physically grounded vision-language models for robotic manipulation,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 331daaf8-c9f6-4a17-bcdd-dc080a59cf89 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8565937f-d623-4d14-a8b7-02b0711b903a · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Aha: A vision-language- model for detecting and reasoning over failures in robotic manipulation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1c5ee29e-1ccd-49a0-9ad1-fff95276bad7 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12cf8a70-2080-4615-be78-b547d5fe1f87 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking A3VLM: Actionable Articulation-Aware Vision Language Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26bccfe0-b39e-4144-b01e-81b6c84afce5 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a98c0c7-f0df-4744-8d4c-1d736d6dbbae · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SpatialBot: Precise Spatial Understanding with Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 097e5731-e685-468c-baa7-10263cb9ef6f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6e5f1d-17d1-497d-bb4a-65b655c0da63 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Onellm: One framework to align all modalities with language,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2caa71ea-ad37-41d4-b6ec-05f4594289c4 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56071ecc-ee84-4dbc-935a-2112c96abf9d · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011b1509-cc5e-4774-93f0-fac0bc8c434e · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Visual instruction tuning,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a910b6a-eb0c-4b64-9019-5d0b62e8c8c7 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Llava-next: Improved reasoning, ocr, and world knowledge,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a766d814-6adb-49c3-8ac9-4fb8d8230c06 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d654cdf2-d889-4af8-a1f9-70ab78727303 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d55d716-a0c4-454c-a509-2261670d070e · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a49532-c2e3-4eab-9ce9-20031c095353 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a5667f3-d22b-4c02-9847-bcde6d26da28 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Imagebind: One embedding space to bind them all,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8947ba36-54a6-4d2a-9a9f-eba3a8c9aff1 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking PandaGPT: One Model To Instruction-Follow Them All
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f326b97-80ef-48a9-810a-cb0f4b893287 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dae310a-f67a-4815-8c63-57486054489f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 437c7b22-36b9-4d91-ac02-80a5c55bd0f8 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75da596b-52c8-430a-b37f-a9537e2fe48a · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8e5330-580e-4efd-b83b-9dd62650d1d2 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Q-bench ++: A benchmark for multi-modal foundation models on low-level vision from single images to pairs,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f4f689a7-7995-4b97-aad6-241701e7b7cf · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Mmbench: Is your multi-modal model an all-around player?
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8e21972b-64ef-42e6-a36b-2f7ff72ae059 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking The reviewing of object files: Object-specific integration of information,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d576d43a-2cf5-4554-b24d-a20e6b4fc5cf · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c8d7df6-dae0-47b0-9658-f395710eee33 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Improved baselines with visual instruction tuning,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b7e5b1f-8e55-450c-9c7b-1ecef7080308 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Phantom of Latent for Large Language and Vision Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb088be-550e-4ff7-bf85-7235199d2b28 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417d6d1b-aa34-4670-bfa6-9e1e59370ffc · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking (2024) Claude 3.5 sonnet
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ee7360f-e93a-479d-b2d2-b9cec25375d7 · outbound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f6b04e6f-437a-46e7-bf60-a6bfe513fb61 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Rank analysis of incomplete block designs,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c37c8ccc-8eef-4408-8042-c81a1192b376 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Direct preference optimization: Your language model is secretly a reward model,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c296a579-0808-42f1-af43-704d1dc5b6ef · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Facenet: A unified embedding for face recognition and clustering,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6595076d-124d-4d5b-b1e0-105fc15c5ad4 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Target-aware dual adversarial learning and a multi-scenario multi- modality benchmark to fuse infrared and visible for object detection,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 827cef51-6cd5-4e73-a4c6-5581618800d3 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking thermal dogs and people x6ejw dataset,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb195b62-1dd8-49a4-9e2d-5f30423b4bdc · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking pet dataset,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3b3d3338-35a0-468a-813d-bdb3a674873e · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Thermal dataset,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 085270e7-7557-435a-9082-858b224f95e7 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Hit-uav: A high-altitude infrared thermal dataset for unmanned aerial vehicle-based object detection,
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f013fc6-b02e-438a-8d2d-ec869489167d · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking animal-detection-flir-extra dataset,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2a9907d9-f032-4e54-a59b-4b17e12e7ab3 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking chips-thermal-face-dataset,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d245561d-3894-4cac-a36b-f8ce905dad9a · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Available: https://universe.roboflow.com/one-rphct/animal detection flir extra
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7495a38f-2956-49d0-a2e0-08aa4fcb7e17 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking DIODE: A Dense Indoor and Outdoor DEpth Dataset
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7781007-900b-457b-a252-e5860655d16d · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Ifsod dataset,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 899ecfa9-da76-4845-95db-fd153375aa7a · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79fe59f6-426b-4a08-9e9a-e187715d8902 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Indoor segmen- tation and support inference from rgbd images,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 834b060e-a989-47c2-859c-a48e126ab185 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking X-ray baggage detection dataset,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51910820-87e8-4fad-be47-580911fab010 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Unifesp x-ray body part classifier compe- tition,
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7124f7bc-a881-465d-93bd-b0fae4144252 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Decoupled weight decay regularization,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c877dd-628b-43d0-8e0f-de0e4f9a1a34 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Qlora: Efficient finetuning of quantized llms,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb8b9a2-dfad-482d-80c0-6fe951c4ab8f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc168b1-45e2-415a-ab31-3d680cd248b3 · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking TE-SQ1 Thermal Camera,
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 271cf759-74ff-4e8c-bbca-483ebcf9d94b · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc74ea9-23cb-411f-8f15-356748985b2f · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking Decoupled Weight Decay Regularization
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0f47f1-e734-49e7-9a52-9a4dd10329cd · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d5cceea-2470-40a4-93a4-58f8a999e59d · outbound
Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking His research interests include deep learning and multimodal large language mod- els
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f1b9142-81fb-415b-9c86-3ec3c2174556 · inbound
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis Enhanced Vision-Language Models for Diverse Sensor Understanding: Cost-Efficient Optimization and Benchmarking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.