Pith. sign in

Paper Citation Record · LEDGER

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI

As of 22 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2505.05895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05895 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:57:24.615262Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:11:55.120740Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d14df83-639f-409c-b371-9b1122964a11 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.612849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.375287Z digest=sha256:9a83e47b7f38713ad632ef382cef5399abaea93aafbf8bfed617d1104d173892

Observation f33b714b-ea48-46cd-b218-07e279888710 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.379811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.379811Z digest=sha256:fd5f3e87f92c7c748a9e26ed8275d96a7f7017dc37f6904a60c4a8bf2f591805

Observation eca134e8-08b3-4b1d-8931-23c829e30ca2 · outbound

This paper cites Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.599563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.385323Z digest=sha256:1ccd991f2681400e648796af63fc353569c549b8b46ee8f67cbe5b60519be49b

Observation cf86736b-96db-4d44-940f-4cb5913c7fc9 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.390797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.390797Z digest=sha256:cbb3036aadd605dc6e0f33b8523c6d2a7d6f8cf6d28bd7b3bd07c41ed4ec0284

Observation 254a10d6-d033-44d1-b5e0-c10a68a380b0 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.395290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.395290Z digest=sha256:e5e22c5bfc59f2f7805fbb931f2281b62721f80ce241c124a9a5bc0b18132ca7

Observation cd92d394-01eb-470c-9524-f15f133cb55c · outbound

This paper cites Website screenshots dataset.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Website screenshots dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.586593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.399652Z digest=sha256:316e0105741b2b8ea162d0ff5c770c9eeaf2b45de514338becaead67556d2074

Observation 73273d8e-bae7-42cf-a991-88a76ec10e1d · outbound

This paper cites Understanding mobile gui: From pixel-words to screen-sentences.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Understanding mobile gui: From pixel-words to screen-sentences

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.573820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.404211Z digest=sha256:aa3b504f8423b8ee9cfe45dbd06436dc6c2ee31ed92177c7df44771a5e0f5a5a

Observation 064b9377-cff7-42d0-a17a-c927b6daa505 · outbound

This paper cites Mobileviews: A large-scale mobile gui dataset.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Mobileviews: A large-scale mobile gui dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.408687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.408687Z digest=sha256:bdf8f5aab856a92142fc3fb94767a1192a5e37e00204506ecc4de9a565be241a

Observation a003114e-61d6-458c-8b3c-dde41d8b6a43 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Exploring the frontier of vision-language models: A survey of current methodologies and future directions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.412883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.412883Z digest=sha256:916f3167a5e4f2c3b8498ff48f7267fa4ab68dab409ba3bd732a27c957803936

Observation 5652d7c6-59b6-4c12-8cac-77ccd8c0cc97 · outbound

This paper cites Navi- gating the digital world as humans do: Universal visual grounding for GUI agents.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Navi- gating the digital world as humans do: Universal visual grounding for GUI agents

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.560154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.417207Z digest=sha256:ecf7904fcaf6e772597f89e9efa267dca45589db673864df601c7df8497ecc49

Observation 50030e4d-fece-4685-8af7-e086e24ac74f · outbound

This paper cites UGround, Feb.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI UGround, Feb

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.546512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.421234Z digest=sha256:d5c083742ffdfae50ff05a2f4c59385037834e69389cd7b14bee4ba7f2e911ef

Observation 7d3be1ee-c14e-4af1-b875-2eda72894c11 · outbound

This paper cites J., S HEN , Y., WALLIS , P., A LLEN -ZHU, Z., L I, Y., WANG , S., W ANG , L., C HEN , W., ET AL.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI J., S HEN , Y., WALLIS , P., A LLEN -ZHU, Z., L I, Y., WANG , S., W ANG , L., C HEN , W., ET AL

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.533392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.425821Z digest=sha256:b4f63b743a1a0ec5f19e4cc4228d5525ea883bb3b6a1a11a2be34093c633de35

Observation ef643501-456b-420f-8926-d2fb645586d6 · outbound

This paper cites Ultralytics YOLO, Jan.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Ultralytics YOLO, Jan

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.519939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.429766Z digest=sha256:b3b0fcfcd97defd871ff27f53553c12c7a5d4541d722ed5c8b86b00b8e8dcee6

Observation 49d19b93-622d-44b3-a238-fddb11759e5a · outbound

This paper cites E., AND KHAN , F.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI E., AND KHAN , F

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.505662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.433768Z digest=sha256:f8ffb8b2bd80a01d9d8d954e7f11693dceca57dcb4aa2e19f91bcb279eb93619

Observation d575d4af-26dd-4e85-b2f5-5f985b430d13 · outbound

This paper cites Hardware in the loop for automotive vehicle control systems develop- ment.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Hardware in the loop for automotive vehicle control systems develop- ment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.491417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.437419Z digest=sha256:ccb9e82e477a603518ef61cf1393954f4eb0ced94194bd1b57eae590f9c1b8d9

Observation 1e4a718f-5dbb-41c2-8d74-dc129ed72726 · outbound

This paper cites J., S AXENA , M., AND KUMAR , A.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI J., S AXENA , M., AND KUMAR , A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.477813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.442240Z digest=sha256:f4ec00284fe52947f73e11e5f052076fadb063910af8cbbd3d7ea5655c89239a

Observation 59927e13-07cc-4814-81a2-b96dd379a543 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.465765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.446903Z digest=sha256:64a7c805007fd908199ac5c9f10f9e5e93661a55363edbb2da84c3d29afe65ea

Observation 0c8cbc0f-10f9-40eb-96a8-9e0ad495669b · outbound

This paper cites E., L I, A., R AWLES , C., C AMPBELL -AJALA , F., T YAMAGUNDLU , D., AND RIVA, O.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI E., L I, A., R AWLES , C., C AMPBELL -AJALA , F., T YAMAGUNDLU , D., AND RIVA, O

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.453355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.451200Z digest=sha256:65fe130585fd8cf3f483662c72b15bf3458efd7d8253d4d6f79ab8ba10dc07a7

Observation c5f5653d-de91-4bc9-80e4-d32f723fd76a · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.455088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.455088Z digest=sha256:f7310b18f78e713da53d5880fcd7a877ff32005b5c129f35785a0c1afb24c534

Observation 1d7bb569-2222-4796-8648-636a60406375 · outbound

This paper cites Q., L I, L., G AO, D., Y ANG , Z., W U, S., B AI, Z., L EI, W., W ANG , L., AND SHOU , M.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Q., L I, L., G AO, D., Y ANG , Z., W U, S., B AI, Z., L EI, W., W ANG , L., AND SHOU , M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.440385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.459234Z digest=sha256:5b6746d4b37ca5088e4aa753a167c20b2c7ff6c62553c0df28dfefe45ad807b0

Observation 141936cc-6bdf-47af-b05b-1d6ef5434e5b · outbound

This paper cites an unresolved cited work.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:57:25.428099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.462909Z digest=sha256:9657314e8fe16f3c2a6eb2bb3690a24ca145a2b975b8fc3b3451e92254ec6e99

Observation 6f990c30-9dbe-40dc-9b44-4083fcea05b3 · outbound

This paper cites Omniparser for pure vision based gui agent, 2024.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Omniparser for pure vision based gui agent, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.415345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.467388Z digest=sha256:3b04578b18609e9bd8a24b94ad112df5a89b082952d12128f088ae79939d8e29

Observation c4cc99a2-aa1b-4767-83fc-16df54f17622 · outbound

This paper cites Automotive User Interfaces: Creating Interactive Experiences in the Car.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Automotive User Interfaces: Creating Interactive Experiences in the Car

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.402301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.472215Z digest=sha256:0857d077a2e4a33e23ad0ce9ffcde77d7650642aa100abbd4ec22fe4f76b98bf

Observation 42050c35-10bb-4ced-bec8-09621970f540 · outbound

This paper cites Functional gui testing of in-vehicle infotainment systems in virtual and real environments.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Functional gui testing of in-vehicle infotainment systems in virtual and real environments

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.389098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.476063Z digest=sha256:ac486c5af75a15d02e70093f46317ca3d6f4dce702a9878ef11deeb88c87f676

Observation 0a59b328-85ed-4d0f-a9cf-9a823da3a424 · outbound

This paper cites Gui agents: A survey.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Gui agents: A survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.480088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.480088Z digest=sha256:51cb500b82605b528777238768c25b3c687b8d1b9d4ef8b83fa618932a8ce081

Observation 3bf4db7e-b6ac-4d92-98c6-f78d36aceee6 · outbound

This paper cites Nomic Embed: Training a Reproducible Long Context Text Embedder.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Nomic Embed: Training a Reproducible Long Context Text Embedder

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.484303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.484303Z digest=sha256:bf8008d190c9838f65915b1c150d9d1b6f54be4d86a9719e39ba7dcc93caba07

Observation 7c050323-7a8a-4314-8b57-cb775839ca07 · outbound

This paper cites TinyClick: Single-Turn Agent for Empowering GUI Automation.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI TinyClick: Single-Turn Agent for Empowering GUI Automation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.488711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.488711Z digest=sha256:0bd31827be892a9aadd0ddd42ee7e92178d803ea8830a873c57020ad62d037d6

Observation fd3d2435-190f-4208-8abd-e546b6158353 · outbound

This paper cites W., H ALLACY , C., R AMESH , A., G OH, G., A GARWAL , S., S ASTRY, G., A SKELL , A., M ISHKIN , P., C LARK , J., ET AL.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI W., H ALLACY , C., R AMESH , A., G OH, G., A GARWAL , S., S ASTRY, G., A SKELL , A., M ISHKIN , P., C LARK , J., ET AL

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.375262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.493676Z digest=sha256:b2f0e7f4dd89dc2962531fe772282a7cd8776140327eb9a939a16ba83f3bd145

Observation 891ac202-ed1d-412a-a344-beeb93fccef6 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control.Advances in Neural Information Processing Systems 36 (2023), 59708–59728.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Androidinthewild: A large-scale dataset for android device control.Advances in Neural Information Processing Systems 36 (2023), 59708–59728

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.362714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.497775Z digest=sha256:7bff39616cdbab048b0ff41faf2ee1e7d40b203496d4ede613cab98cf88d9e43

Observation 3d405db2-88d2-4281-bd39-6147d5ef31dc · outbound

This paper cites An overview of the tesseract ocr engine.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI An overview of the tesseract ocr engine

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.349970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.501605Z digest=sha256:88d77d68cc44b08e454ac59a3d856be3d46b2ee11d3081384c81544f6d9829eb

Observation 2e6b6d85-98cb-40f6-9707-a878e44e73a8 · outbound

This paper cites Visualizing data using t-sne.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Visualizing data using t-sne

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.336558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.505552Z digest=sha256:89dd827d5560b29c740e26ec950871d521751dc8f788d2eecc80cbcda8b94598

Observation 7824d42f-9f11-41c9-b848-cd4ae836aac4 · outbound

This paper cites Canoe — hil & sil tools across multiple industries — simulation, 2025.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Canoe — hil & sil tools across multiple industries — simulation, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.323830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.509517Z digest=sha256:f49c28e2db7246c1ccb44357537593c876e419a8be6b792aa61b926a9ef7bcc0

Observation b2b84fae-3894-4006-8ed6-014d8eb2bfb7 · outbound

This paper cites Large Action Models: From Inception to Implementation.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Large Action Models: From Inception to Implementation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.513182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.513182Z digest=sha256:28ad5e1e340a5052b5e492fdabe4f73eb6d760b6016265d9df5b67c73163cb19

Observation 1356ea50-503f-4578-9c93-f7fb9b507162 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.517277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.517277Z digest=sha256:130e80767867aaeb528f8d89f602c7fd89ab342733b54ed446f8bd96fc6ce1f0

Observation 6a8e7663-b871-4744-ad83-accac0b7532b · outbound

This paper cites GUI Agents with Foundation Models: A Comprehensive Survey.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI GUI Agents with Foundation Models: A Comprehensive Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.521396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.521396Z digest=sha256:17bf4b6e17f7d0c1c80f1cc56e764fdc1c0a5595e58c49d8133dfc38ad593e0e

Observation 8a327eca-7943-48f5-8b17-0b6ae68557a6 · outbound

This paper cites Webui: A dataset for enhancing visual ui understanding with web semantics.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Webui: A dataset for enhancing visual ui understanding with web semantics

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.311324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.525480Z digest=sha256:8087167566f31d180bb20259eaded1fa5829919503b7c39a8f0fe937d3bbb4c7

Observation a1bb676c-8f0d-467c-a292-452c979a03b0 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.529413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.529413Z digest=sha256:1de4397007d8c5ce774b3f469611cabbaabd0f2fea0ca6a8f4ab926b96619d4e

Observation e063f5eb-1a28-493b-8569-c054c7370469 · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.533628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.533628Z digest=sha256:d43fd19d1b7580e4190a837d17d52faf7f6e17db4fef11829567eb812e2592c2

Observation bc089ea0-3488-464b-9acb-831476845143 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.537694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.537694Z digest=sha256:4ce6fca77e97f41b2bef873f2090db4cb19d0dd88d254934f47d8859a22fa325

Observation 8713b0e4-c0fb-4ad4-af36-0c3804aeafac · outbound

This paper cites Automated testing for automotive infotainment systems.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Automated testing for automotive infotainment systems

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.298963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.542073Z digest=sha256:b4cfc2daa37f200dab1f23d2e3484fa09a167e07d81af13708c65286c525be02

Observation 316f154e-1bff-404b-848b-1104a33dc176 · outbound

This paper cites Large Language Model-Brained GUI Agents: A Survey.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Large Language Model-Brained GUI Agents: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.546651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.546651Z digest=sha256:d9ec97bc19bbf3f85b035889a75348c779382a4bc3247943cb9ef1704de23f2a

Observation 4bb2817f-6b9a-4d99-a2c9-e1db882fa6e4 · outbound

This paper cites Swift:a scalable lightweight infrastructure for fine-tuning, 2024.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Swift:a scalable lightweight infrastructure for fine-tuning, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.285806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.551580Z digest=sha256:58891f8dac7e0a0abc8936fcb220ecd95b5a4f4f6bdab4e05c038834d2eadd54

Observation a7faab3e-e77b-483a-a52e-cdb3a0225788 · outbound

This paper cites AgentStudio: A Toolkit for Building General Virtual Agents.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI AgentStudio: A Toolkit for Building General Virtual Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:57:24.555597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:57:24.555597Z digest=sha256:3f041f86e285795ae11899c280f98e3250e2384ed00d89282a2ab326609f05db

Observation 19e50659-aa0e-413a-98e4-e6536c6118d1 · outbound

This paper cites select kWh/100mi as electric consumption unit.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI select kWh/100mi as electric consumption unit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.272230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.560091Z digest=sha256:0fb16b2916a2c611a3116ec73337b49962a1910f85f2fc1c40e487e6b9541cfe

Observation 79f6e6aa-a4ed-4cc7-9027-3b1798c1bea1 · outbound

This paper cites Select the burger button.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Select the burger button

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.258948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.564214Z digest=sha256:a6c74914940ac906905bf1b2ed55c5cf3c78fc5fb96a3e6d10d2565969aee2a5

Observation fb593c0d-44ef-47a5-9fcd-a3089802b73f · outbound

This paper cites turn steering wheel heating on.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI turn steering wheel heating on

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.246411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.568395Z digest=sha256:d37cc3824728ac97ee3f2e747894761cd2adc608f537d86240cc6508e1649aea

Observation 383b230e-793e-4ca7-9dc2-b97cffc3cacf · outbound

This paper cites deactivate the reminder signal for mobile phone.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI deactivate the reminder signal for mobile phone

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.233341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.572491Z digest=sha256:105eb995e43c8919b1508e17ccab05c3d668290c5041b54c4ef14dde97a40675

Observation 6741ee9b-956b-457f-89a2-636e24df5470 · outbound

This paper cites add phone to Favourites.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI add phone to Favourites

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.220457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.576584Z digest=sha256:50443eb9b17badbad00344c14ef51bc13d5e26e2d0d27e6f300d7b1dc5f56d09

Observation 67d14532-d8d6-4d6d-8897-c105c9b97678 · outbound

This paper cites increase the right temperature setting with the plus button.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI increase the right temperature setting with the plus button

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.207197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.581292Z digest=sha256:7795246253930c3f0c212d216a40804ef0a595a62736e19e373511ddd6a96635

Observation 05d56554-e46f-4915-b751-13a23459bd0d · outbound

This paper cites Activate the first weekly item from the charging list.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Activate the first weekly item from the charging list

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.194787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.586345Z digest=sha256:b0c5835032e72976dc79015261f26593d23f892d92f8b7ec4dce46f7b107715f

Observation 0b96107b-8705-4c59-8156-8031ff2d3fc6 · outbound

This paper cites Go to main menu.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Go to main menu

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.182094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.590507Z digest=sha256:6019ee5416512ed45c604e3a0c434a6379d949663685b8a59a10a1e5e1a10b7a

Observation 47362b38-3962-4802-9a5e-c717ad857756 · outbound

This paper cites Audio quality is set to low.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Audio quality is set to low

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.168527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.594659Z digest=sha256:0445d3d385d351a60512ccb030f884d19ade6a2a5181f0be45a267a867efa8e9

Observation 2848ad3f-b212-41a6-9e28-49047c0ed883 · outbound

This paper cites Sound settings are displayed.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Sound settings are displayed

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.154134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.599483Z digest=sha256:70f3b0fcb8faa7355f361b4d3cd486767218487ebe1490723757792870d07cd1

Observation cc305e1a-d7da-4105-a9f7-3e3bd3191ec9 · outbound

This paper cites Driver assistance options are dis- played.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Driver assistance options are dis- played

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.141374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.603547Z digest=sha256:be1c3d171a89b2790817a8c036a0ce815ee9ef6b1a878369c9fe21e45a1d7828

Observation d77b9c8e-8757-4382-871b-b56293aedd8e · outbound

This paper cites The medium sensitivity of the distance control was se- lected.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI The medium sensitivity of the distance control was se- lected

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.128237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.607390Z digest=sha256:7cf16fb323c9dbc0c56ded5b453c6164c490f4d79ed49e4cde14a8cbf470484c

Observation 001d0416-cc4e-47d2-bf21-e9ee5f468803 · outbound

This paper cites Noise reduction is disabled.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI Noise reduction is disabled

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.114654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.611357Z digest=sha256:5e8008cbfb649d105481d690d5931327ff6a721e6d52dd0edcbb2a640b8b2500

Observation 976708f6-9228-49ab-b482-cbea7688d00e · outbound

This paper cites E-Call settings are shown.

Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI E-Call settings are shown

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:57:25.101192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T22:57:24.615262Z digest=sha256:b69642f4c6fb3ac7539923b96e087c006f8f2440d14294f882ea5b8c71a2cf8a

Pith citing papers

Observation 549febca-045c-4126-9865-9366524671b0 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:55.120740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:55.120740Z digest=sha256:a8906ef4e6eed0052e09a59116cc5f90796c5050e7c41a6e75252e6a426faab2