Pith. sign in

Paper Citation Record · LEDGER

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems

As of 15 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2608.07861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07861 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:19.424861Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact6
  • verified fuzzy62
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71d47ad8-bae6-495f-9981-555a93e09997 · outbound

This paper cites VQA: Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA: Visual Question Answering,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.017635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.017635Z digest=sha256:56d390cbd90cd1cfa8c929179253cfedb49203bc7c2c916437ccb9d807dec4e9

Observation 8d7bdee0-3f56-4293-b5f7-efcfc3389f5c · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in visual question answering,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.022230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.022230Z digest=sha256:ebf51e75be6ecdfd24bb91116c4299ac984c119f28ff67c8483c433fb96d11a0

Observation aeee4825-6bf1-400e-b230-b03f9f69aee0 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.026832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.026832Z digest=sha256:ba0eee216f566dcef8dd5f953d9bbe22348553049f22a1a323521fe85618ca0a

Observation 51d95b2f-c7d9-4ea4-819e-9c7989f2f0c0 · outbound

This paper cites Ray-Ban Meta AI glasses gen 2 & gen 1.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Ray-Ban Meta AI glasses gen 2 & gen 1

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.031728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.031728Z digest=sha256:1c983c270ce1fd8c4571bb25c087496ddcac51e8ac7270aa83c983709d619e2b

Observation ca9a5527-1242-4b86-bd1c-85c3fadb4b4f · outbound

This paper cites Vision-language models for vision tasks: A survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision-language models for vision tasks: A survey,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.036101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.036101Z digest=sha256:f430065f250f38af85c05981adcda4932f4cbd00c638592792578589221d81b1

Observation c423ef81-b965-4ded-9423-26bd29476cc2 · outbound

This paper cites A survey on multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A survey on multimodal large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.040708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.040708Z digest=sha256:e6f8017ddcd5ce1edd0ccd49cbe17a98321aff9f332e24ef33deeed4de9b76b6

Observation 5f6c1989-ff9c-411f-9b78-d5a98caf6067 · outbound

This paper cites VizWiz grand challenge: Answering visual questions from blind people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VizWiz grand challenge: Answering visual questions from blind people,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.045319Z digest=sha256:9d15c2e2b38540fe26029187f2e72acf2ace3c0f00e4d90be15e85dae2cf019d

Observation 2ec2a70d-2d04-40e7-9f53-bd449a390a50 · outbound

This paper cites Be My Eyes: Lend your eyes to the blind.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Be My Eyes: Lend your eyes to the blind

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.049342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.049342Z digest=sha256:4a87f5b4a2490534777caf75b2cb290e824bb5828acd197678b62beefbde29d9

Observation c7a0d083-22a2-4034-b581-b6528d7cdf17 · outbound

This paper cites Long-form answers to visual questions from blind and low vision people,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Long-form answers to visual questions from blind and low vision people,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.053382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.053382Z digest=sha256:cab8a11c382180fab353a40a803c4bbf0838e3ec174cdf6d94b7592c2eeb6d42

Observation f809fab8-781e-4a98-a07b-14dbb91febbb · outbound

This paper cites Augmented reality anatomy visualization for surgery assistance with HoloLens,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Augmented reality anatomy visualization for surgery assistance with HoloLens,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.057368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.057368Z digest=sha256:b090c9e2a135f8ba2e9e4333cdb94b5d268a8c3b2e04307e0425276494ffe050

Observation f4e53888-1f74-4bcb-8700-525f7fbc0d40 · outbound

This paper cites LLMs Enable Context-Aware Augmented Reality in Surgical Navigation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLMs Enable Context-Aware Augmented Reality in Surgical Navigation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.847198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.061282Z digest=sha256:21a1e9f683e28eaf1b4f787a6863ff658e84c7fe682a492b1e9d7d815a95f326

Observation 1f4ba1a6-b0cd-47fd-affe-6011636153dd · outbound

This paper cites DriVQA: A gaze- based dataset for visual question answering in driving scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems DriVQA: A gaze- based dataset for visual question answering in driving scenarios,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.065974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.065974Z digest=sha256:acd89c7576f5b759e88b026e4f47574746620fd172a6c8eb65552ae15debbe9a

Observation f23872a6-04d1-484e-93c5-fa11a7b911cf · outbound

This paper cites Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mimicking human attention in driving scenarios for enhanced visual question answering: Insights from eye-tracking and the human attention filter,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.069805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.069805Z digest=sha256:b00af17c0f2cf1f4dd0b6727040514c566af585408aa1e6e671fd228e0ee6e23

Observation 0e174cce-37bf-44c7-a1d0-9ecab01484fb · outbound

This paper cites Demystifying small language models for edge deployment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Demystifying small language models for edge deployment,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.073724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.073724Z digest=sha256:5958629730cdb3bd75e0d614c309a49f8edaed60b13e42434e64d50056287844

Observation 4f3be074-5d14-4261-9b0f-284edc6d3958 · outbound

This paper cites Efficient processing of deep neural networks: A tutorial and survey,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Efficient processing of deep neural networks: A tutorial and survey,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.077650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.077650Z digest=sha256:2fcbee5908327713f5b0e439f8c1ee447e6a2893e5b7434fcc9dbbd39b08f912

Observation 3f716f81-e7ef-4153-bce3-1a2d210cacd2 · outbound

This paper cites LLM in a flash: Efficient large language model inference with limited memory,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems LLM in a flash: Efficient large language model inference with limited memory,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.082137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.082137Z digest=sha256:387ad2404cd40b425e28080c69a8398b5460eafa045662661a52e8c729e7ec85

Observation 093f752c-10f3-48fc-aea0-9e19b5827545 · outbound

This paper cites Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mo- bileLLM: Optimizing sub-billion parameter language models for on- device use cases,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.086239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.086239Z digest=sha256:6e8a92df93fb9e68aadb61f7e9a54928528b2ab69a9557b32d55e34a59c9b47d

Observation ed140e8c-6bb7-4a98-bae5-f052bb71ac1e · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.090469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.090469Z digest=sha256:b404677d31e8d1fa57af676a993df6b58790b9de5eff17b51b05418422788605

Observation 5591e1b1-915d-433f-9465-befe03319a4a · outbound

This paper cites Bench- marking tinyml systems: Challenges and direction,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking tinyml systems: Challenges and direction,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.095087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.095087Z digest=sha256:080f72364e857b3e25256792f5fe3f3d48796194e3958a83f8f50968adcafaf1

Observation e0624723-5114-4f44-ba70-eaa0a743094d · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.099492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.099492Z digest=sha256:85f1baace953a6d91742b065538e634391e535362f63de12c7aea9e53190185e

Observation 030b6d04-1204-49be-9816-dcd0f2bb3b2d · outbound

This paper cites Rokid AI & AR glasses—redefining reality.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Rokid AI & AR glasses—redefining reality

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.103858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.103858Z digest=sha256:9903bf866d693e2218e7a4b74071f0548b7de211fd6c0c7cb1df73cbd2b0343d

Observation 33c982d4-3c83-47e6-acb7-fa530e2c4d11 · outbound

This paper cites Project Astra: A research prototype exploring the future of AI assistants.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Project Astra: A research prototype exploring the future of AI assistants

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.108144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.108144Z digest=sha256:0895fc2a6722fbc6d7391907e9f2513c543ee3ee9abde76d138b7efbafc8f994

Observation 43286f77-d8ae-43c3-9bf8-3483d41a9953 · outbound

This paper cites Google vs meta smart glasses: Which AI frames are better.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Google vs meta smart glasses: Which AI frames are better

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.719204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.111993Z digest=sha256:0142caaf579b376da1c907116b44150f8e141af1eb8409ac16ce034a225f217e

Observation f4719216-9ab0-4982-9300-c7e5fadb17c7 · outbound

This paper cites Edge cloud offloading algorithms: Issues, methods, and perspectives,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge cloud offloading algorithms: Issues, methods, and perspectives,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.705801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.115924Z digest=sha256:0bed229459183e8e06aee12745c5625f5330eacda37ad347b79b4bb9c6307112

Observation e2197941-1bf2-423d-afa5-c6c52a9dda66 · outbound

This paper cites Cosmos:computation offloading as a service for mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cosmos:computation offloading as a service for mobile devices,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.690590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.119812Z digest=sha256:4c79709ea9744427cad7199efc10277fb5f11e64b6b4a9335d3127ee5dab2f50

Observation b49ec1be-5003-4b1d-8980-a2abd19fa10c · outbound

This paper cites To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems To offload or not to offload? the bandwidth and energy costs of mobile cloud comput- ing,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.675452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.123655Z digest=sha256:fca5d68f866c209e7693878423d6603eb19b1f9885eeb28b5f5e30d97f5ca64c

Observation 46a2ff98-5f5e-4710-8ec2-ac5fa2d36e6f · outbound

This paper cites Images and vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.662673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.127468Z digest=sha256:218b67784fde5e342642a57b0430fdd901291dd324908c6a03c0bdbbf01d67fa

Observation e1199267-7777-484a-be17-82f4ba020a32 · outbound

This paper cites Understand and count tokens — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understand and count tokens — gemini api,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.650544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.131426Z digest=sha256:5226b706703010dc44b1adc472600f3988834a9119508675b45b44eddddf55cf

Observation e2f21d7e-8e38-48c8-a39c-c1f0178b018b · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems A-vit: Adaptive tokens for efficient vision transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.636667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.135656Z digest=sha256:5d39627affd0a2d13abf30db19d07089704cbc5963e4995da11df0e066bb2011

Observation 6d2e4ea8-265e-463d-9465-ec7bf0e3941a · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Not all patches are what you need: Expediting vision transformers via token reorganiza- tions,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.139426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.139426Z digest=sha256:ce7a7b0d07e9abb227e0302e851247b3d3bc3bb112ae9890ee7f38d97e7b7299

Observation b1f216e8-d0ab-496d-9366-3b51981bba7a · outbound

This paper cites Token merging: Your vit but faster,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Token merging: Your vit but faster,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.615091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.143097Z digest=sha256:1ade6d29b8512788492f716ba22babf5db480889666d9b6a582d3727f8b644a6

Observation 6fb6ed1d-b647-4032-bb55-085660552667 · outbound

This paper cites Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Elf: Accelerate high-resolution mobile deep vision with content-aware parallel offloading,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.601893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.146721Z digest=sha256:d1ae5a3956d287cb91bbb1cefcea97a411b0e0a0f053506e481e3bad5a764e01

Observation c815e248-ed3e-4b28-ba44-419e630d7ae0 · outbound

This paper cites Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accumo: Accuracy- centric multitask offloading in edge-assisted mobile augmented reality,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.589138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.150710Z digest=sha256:30ab1d692c00f4aade3f6558402f5ae2b162f77e33adb6e1baddf384e9a48af8

Observation db119f59-071d-4edf-b611-a1ad706dad96 · outbound

This paper cites Deep contextualized compressive offloading for images,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep contextualized compressive offloading for images,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.575387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.154686Z digest=sha256:9f6d383b652eae3b7877c91eb12e6befac9cbed7ef55f4d9a734ccb02b60911b

Observation 62d30240-4617-49fe-8b99-a546eb9e352a · outbound

This paper cites Towards wearable cognitive assistance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Towards wearable cognitive assistance,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.561858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.158591Z digest=sha256:0947c342b2a2a176a2ebdf17d3099c70ace02479b03eeffac4637f94c4c42f0c

Observation e1736367-04fd-4dca-81ad-2fa144d30a94 · outbound

This paper cites Deep learning with edge computing: A review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning with edge computing: A review,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.162754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.162754Z digest=sha256:2f0546f37c4f220784ad0da6604b4713c052c2a5c717ef1633a6f40460843c8e

Observation b611c525-e955-40b7-9ec3-84c34dd7569d · outbound

This paper cites Color-to-grayscale: Does the method matter in image recognition?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Color-to-grayscale: Does the method matter in image recognition?,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.540375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.167041Z digest=sha256:be5bf69be54cc17fd5f0af6dbca633e44ee62c48f77f27e4318a912ef76e4b99

Observation 82072b31-de42-4ac8-bb17-d7543c2d9d70 · outbound

This paper cites The jpeg still picture compression standard,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems The jpeg still picture compression standard,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.170915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.170915Z digest=sha256:163a31ee7c12d2e4c6135a38e5fafe114b2d6318b589b6cfc23e4bef15cd243f

Observation a4b2b770-2d6d-4fd1-af7e-2750e9d6511e · outbound

This paper cites Learning to resize images for computer vision tasks,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Learning to resize images for computer vision tasks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.520409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.174922Z digest=sha256:68be9d73531da93720d6a86e9416a11858319faebebb54aa02400c7c96398ce8

Observation 6e7aa746-6bc4-4004-a37f-d48474ab831a · outbound

This paper cites VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems VQA-MHUG: A gaze dataset to study multimodal neural attention in visual question answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.509216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.178967Z digest=sha256:7c624f5612dcc9fd80e02bc1ebd3e13a1018726b388dcb3a84dbd8a71ad1b62c

Observation 8c632c78-07be-4139-a733-c27e553b0ce6 · outbound

This paper cites Eye gaze tells you where to compute: Gaze-driven efficient vlms,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Eye gaze tells you where to compute: Gaze-driven efficient vlms,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.497438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.182640Z digest=sha256:5eba4c10bd81f7ec8c7f516e4229d668c15434f1c01f3765d973582fb71b0cf8

Observation ea0c5c9b-871a-45d3-924c-b67571e48b83 · outbound

This paper cites Saliency detection: A spectral residual approach,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency detection: A spectral residual approach,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.485099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.186564Z digest=sha256:847eccdf151e94f3f904b6f2d6221ec16cb5986eb3180c4266be51c40229adca

Observation 2984d488-75ea-4a56-b6c8-3cfd71b74027 · outbound

This paper cites Saliency driven perceptual image compression,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency driven perceptual image compression,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.473103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.190484Z digest=sha256:a06379d9a09c4c375b48fe8e8b54bdd50d4fbd46144c975c2fd6e90276819e24

Observation 46657777-0be2-4c24-a490-16df0595a874 · outbound

This paper cites Visual cropping improves zero-shot question answering of multimodal large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Visual cropping improves zero-shot question answering of multimodal large language models,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.461376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.195307Z digest=sha256:c40e29d1ebd6bfce11b085e38230b2c019aad47b0f0ef12fa4e33140dedb2a8d

Observation 0661ec84-8aed-4a0c-9051-4f3a6796db86 · outbound

This paper cites Edge assisted real-time object detection for mobile augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Edge assisted real-time object detection for mobile augmented reality,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.450028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.199136Z digest=sha256:8e1fabf8517a53d25a5a29a4f17c2a256bdfb3d86c61e3e7bcb21b5a1f9c27e1

Observation 6d1ff298-076a-42f4-9024-d3724fdc2cd0 · outbound

This paper cites Glimpse: Continuous, real-time object recognition on mobile devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Glimpse: Continuous, real-time object recognition on mobile devices,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.438645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.202934Z digest=sha256:2f24af784f0c1cad3d3a77460830cbcc2a65ad616586130e4eecc0e377782f04

Observation bb99a795-38e0-4c1d-834e-d84387269ab0 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.426348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.207129Z digest=sha256:fb3997c7fb67a4b056b61e89f7ed471c5a9d5729cf57ebf17c77f5110c97ad30

Observation a4b4c0c5-25ef-4e90-8c8c-a85746a2e407 · outbound

This paper cites MMBench: Is your multi- modal model an all-around player?,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMBench: Is your multi- modal model an all-around player?,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.414074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.211247Z digest=sha256:b12fab0f5f5ed56dea32dd78e9c9a96e38c2f37882b84a2b1bb89626e6f09e0a

Observation af7e46f8-ecda-49b8-af04-3f441fd59c3e · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.401510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.215889Z digest=sha256:33dc3ca98a5253e5f26214701c7483373cdeb0a687ac46ae88bb97401e73de60

Observation 2659ab9f-0bc5-4bfb-bbd5-c13dde02c5d8 · outbound

This paper cites MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MathVista: Evaluating mathematical reasoning of foundation models in visual contexts,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.389142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.219803Z digest=sha256:6b38509dfeb0b367808a9f650f9df8086e3d23126e041760cb9316f6de8a6fb3

Observation fab0333c-8dcf-4f0b-a835-2e2417777199 · outbound

This paper cites HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems HoloAssist: An egocentric human interaction dataset for interactive AI assistants in the real world,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.376588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.223759Z digest=sha256:a34e2f07a6010268c81b1f728876c1b5338e4793aab38c81df0cf2b62115fb06

Observation 92215827-6b85-4a43-9454-1c9eaaba8522 · outbound

This paper cites WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems WearVQA: A visual question answering benchmark for wearables in egocentric authentic real-world scenarios,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.364074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.228109Z digest=sha256:845cd9ebed58f4ad8ac53452717726b60ebffed6e31d380f28fb5ee97752dd2b

Observation ea3fede4-c24d-4cf7-a57f-7641895da963 · outbound

This paper cites SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.232005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.232005Z digest=sha256:15a56edb0a1d946bd842f2c218a936a6d1411a2b300da34a4490168a5bd7995c

Observation 9510abb9-44a5-4dd3-aab4-f6334dcd9568 · outbound

This paper cites CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems CRAG-MM: Multi-modal multi-turn comprehen- sive RAG benchmark,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.238653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.238653Z digest=sha256:c7afda35797b72039f41aeb6150fba4264c958fc27b7ba81929077c72130ed12

Observation d117e68e-2597-4511-8a8f-cd4b438c1b90 · outbound

This paper cites OK-VQA: A visual question answering benchmark requiring external knowledge,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems OK-VQA: A visual question answering benchmark requiring external knowledge,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.351783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.242494Z digest=sha256:52e2807ff56d32fb06c9517b53828a50fae8bbcdd7ad80680124010c99655d38

Observation a41e1861-e52a-4ab1-ba2f-2da8df2db08a · outbound

This paper cites MISAR: A multimodal instructional system with augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems MISAR: A multimodal instructional system with augmented reality,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.339738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.246484Z digest=sha256:9c6f3a6f1398d1351417ec1988ae4bb3d88f1773766f28e0c1faf39a32cd3ed2

Observation e10d0e33-8929-4845-b0c5-9b401e07c111 · outbound

This paper cites Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Guided Reality: Generating visually-enriched AR task guidance with LLMs and vision models,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.328278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.250850Z digest=sha256:0fbb56f68bdec0dd5cb96df07808b7488b83a817aa5b15feda3d35885ffcf4b8

Observation 54592885-ffbf-410c-9493-1fea13cc1cd7 · outbound

This paper cites EmBARDiment: an Embodied AI Agent for Productivity in XR.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EmBARDiment: an Embodied AI Agent for Productivity in XR

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.254659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.254659Z digest=sha256:88e07c4d7775a7bb44d6a3a0cca8bc1232ea9f9c6c195bb41ea1eeb7c8bf8ca6

Observation ef168b0b-05ee-4c8c-acfc-835114242bdf · outbound

This paper cites GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems GazePointAR: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.315950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.261066Z digest=sha256:0654f8a5809d8bb452b282982bfbe2d7b2b2a509ce09f2fa71a7b0e4a7a37c9f

Observation 743d7ba5-8ffe-416b-9a38-98f7ed88fb16 · outbound

This paper cites Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.701778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.265562Z digest=sha256:87a5d07133bb83fffce6f95c0390e4ac45ad6f7b7b492d6e3cd04e4cd43b644d

Observation 01dd9cdc-092c-4bb0-bc95-48c813e8338f · outbound

This paper cites Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.676748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.270631Z digest=sha256:f503f6deaa3e60c8cd513602d4470d4d8b715fd3fe148a76cc910fd99648f42c

Observation 6676f53c-2d73-4ea9-8c37-0990526deebc · outbound

This paper cites Exploring the use of VLMs for navigation assistance for people with blindness and low vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Exploring the use of VLMs for navigation assistance for people with blindness and low vision,

Reference 62

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:50:19.657100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.275112Z digest=sha256:63b34f51cea3acabe1aa4117f4f8dbd64ba34837e3c84117167aa19da060363f

Observation ed1391d9-df16-421e-b77a-79a1b2965b96 · outbound

This paper cites BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.569534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.280108Z digest=sha256:43346516e01eaeecdb957f662898dcf8620067d277d3e613208c340ff4ad8c02

Observation 22753e43-642f-43cd-833a-286ca03f7fa5 · outbound

This paper cites Objective assessment of the webp image coding algorithm,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Objective assessment of the webp image coding algorithm,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.304358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.284581Z digest=sha256:cb41af9e9097543f48cad696ea35d581c3766d442090590a6010e370c62f88d6

Observation 01bace46-f84d-4361-85c5-794780a21227 · outbound

This paper cites Vision - claude api docs,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Vision - claude api docs,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.292984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.288751Z digest=sha256:bf9587ab82b518b09ca26df60e40c6603368bec2d0adc83f75134b4a2266de81

Observation fa6fb7df-483a-4d21-956c-805c2a386197 · outbound

This paper cites Image understanding — gemini api,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image understanding — gemini api,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.281030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.292821Z digest=sha256:b5d5d16d051f8a3fce6149b1e85455bb54f15fbaf151630979df7326122cb41f

Observation 41af0a65-9a0f-4dd8-8072-cab64b6cb2a6 · outbound

This paper cites Image interpolation and resam- pling,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Image interpolation and resam- pling,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.268208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.297025Z digest=sha256:d2e84c5289a2e6073d5b5fac9eed289581abbd76ff31553929d0988c2968a648

Observation f99b59f9-f810-4d14-8e34-e9a36a3fd0cd · outbound

This paper cites Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Colorbench: Can vlms see and understand the colorful world? a comprehensive benchmark for color perception, reasoning, and robustness,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.256188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.301374Z digest=sha256:412a7ca65fb2bfb5fa7781e1649fb4f6c74861ea92c2f905073dc6814b225e29

Observation 6c8b19a3-96f7-42bb-86d6-9061eba41935 · outbound

This paper cites V oila-a: Aligning vision-language models with user's gaze attention,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems V oila-a: Aligning vision-language models with user's gaze attention,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.242817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.306560Z digest=sha256:90e1f3dd24d4ec589bdebeb9859e4756c300a7385a83cef1ff5fcc9f0093d9aa

Observation f87e2bc5-b3fd-4be5-8c95-6ba2989c591d · outbound

This paper cites EgoVQA—an egocentric video question answering benchmark dataset,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EgoVQA—an egocentric video question answering benchmark dataset,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.228121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.310732Z digest=sha256:7625baffffdb843d4d98ab378173138355e5efe714331d8ea12348f52369e83f

Observation 9bf1c15d-5ff4-405f-8e17-d64dfc0701e2 · outbound

This paper cites Neural machine translation of rare words with subword units,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Neural machine translation of rare words with subword units,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.216201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.314944Z digest=sha256:8de912f4b695ce883507d70845dbd5033f92f7b882b8194a05b5adb5ca13c28a

Observation 2fd64ec6-85d1-4967-b5d8-38c431fa085d · outbound

This paper cites What are tokens and how to count them?.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems What are tokens and how to count them?

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.204919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.319134Z digest=sha256:935fc703cd3252192c88a561a205e738ddc34aaf0af2fcfa70b2b8fd78d723ef

Observation 5da132dd-8582-4cbc-8c93-f29ca696bc7f · outbound

This paper cites Saliency based image crop- ping,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Saliency based image crop- ping,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.192706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.323268Z digest=sha256:9f912daa0fe2216704be28c9099dc9eb076144878693e58dbcc6c3adffdbe076

Observation 2794bdd8-29c7-47b4-a08d-47226a09f943 · outbound

This paper cites Benchmarking deep learning models for object detection on edge computing devices,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Benchmarking deep learning models for object detection on edge computing devices,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.179180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.327140Z digest=sha256:87d836253ec801e01f8bdacf36833cfc272d4fd925a100552e182931f327c75e

Observation 89aca3b1-210a-4f3c-a678-ffd30360e0e6 · outbound

This paper cites Region-of-interest extraction method to increase object-detection performance in remote monitoring system,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Region-of-interest extraction method to increase object-detection performance in remote monitoring system,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.166519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.330909Z digest=sha256:04cd6a8bec645c98c6158f5d95648cd6d8b6354413156f61a8609c22bc102dde

Observation d50df610-97c5-461a-ad07-33e8fbcfb191 · outbound

This paper cites Improving automatic VQA evaluation using large language models,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Improving automatic VQA evaluation using large language models,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.154479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.335462Z digest=sha256:604c2e7463b6bf44c7ba993256f460c5285254d93764733a905c5a824500ff2d

Observation 33234788-e51a-4184-b466-9825241e5b41 · outbound

This paper cites Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Mind the uncertainty in human disagreement: Evaluating discrepancies between model predictions and human responses in vqa,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.142051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.339617Z digest=sha256:718c4e9748ee71365d102fc1342d6eab2398606975a11e83c43602c9160730fe

Observation 0f4ff4da-137d-44ba-8cf0-faab0953198f · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.128003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.344182Z digest=sha256:a449abb1309b8b36699f7918acf95cce0b40032301d017f31e249a9593dcc8e4

Observation afd1f3e0-03c3-4560-98cf-b74131d5675a · outbound

This paper cites G-Eval: NLG evaluation using GPT-4 with better human alignment,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems G-Eval: NLG evaluation using GPT-4 with better human alignment,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.113194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.348271Z digest=sha256:b0e10b9497f78472dabd8c2d7ed3c1b2b90b325816bb71e2bad4b9b448084ba9

Observation 57d015c5-2614-4081-8fbc-ef1c06409793 · outbound

This paper cites Chatar: Conversation support using large language model and augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Chatar: Conversation support using large language model and augmented reality,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.099459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.352344Z digest=sha256:39efe21a14879dc179a9ae0aad5735bbe5b922be294fa263a34f3926948e29e2

Observation 0db091eb-cd22-4e25-9b2a-5515d36fb4b3 · outbound

This paper cites Next-generation networking and edge computing for mixed reality real-time interactive systems,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Next-generation networking and edge computing for mixed reality real-time interactive systems,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.086379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.356200Z digest=sha256:aa2566ede27ecdc441287f97086f404911127c1e3c732b19193a578ea4567c74

Observation 43a2847e-f2a4-4cdc-866c-3ed4e713706d · outbound

This paper cites An egocentric vision-language model based portable real-time smart assistant,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems An egocentric vision-language model based portable real-time smart assistant,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.071213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.360215Z digest=sha256:7a91273bb95b9012f4375b2ec5728f9b7e3a70c191d3c7e93c13d9c77f1e4821

Observation 4cbd87af-581d-45c3-a2de-fb96d7c5dbfa · outbound

This paper cites Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Teaching LLMs to see and guide: Context-aware real-time assistance in augmented reality,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.364187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.364187Z digest=sha256:a863d56b30fc0323321b5599e4ba91a389f7681284c4ce533da734ef3c8fd772

Observation 3999f9fc-18db-49bc-85a5-767882c1d7fb · outbound

This paper cites Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gartner predicts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.057437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.368841Z digest=sha256:3becd6f3082471b19ee3042488be05accb4221d2c4d9e42d35a5e81a28052170

Observation 9ad2d767-1a29-49bb-8fc9-b5b4f177566a · outbound

This paper cites Deep learning-based object detection in augmented reality: A systematic review,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Deep learning-based object detection in augmented reality: A systematic review,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T00:50:19.372822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:50:19.372822Z digest=sha256:9e1a12085f0fb750b7c9fdde7669279e61297726097bc1c64040d7e69bc678f7

Observation b3c695f9-5789-4416-885c-8d6bf76902b8 · outbound

This paper cites Realtime API with WebRTC,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebRTC,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.036424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.376497Z digest=sha256:25a6ada089eda953f8c38197c07c33f2330523290c1323cdd04b3a90a81223e4

Observation 2437408b-3e1a-4cee-9a4e-18cefaaabc6c · outbound

This paper cites Realtime API with WebSocket,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Realtime API with WebSocket,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.022217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.380300Z digest=sha256:171c37a73c6e1547ed31ca039d6c55d63ff50e9990a9e28dd442e43f189f2194

Observation 0b4103ae-d36f-4f62-b6d0-c6c63b65fa68 · outbound

This paper cites Bench- marking the effects of operating system interference on extreme-scale parallel machines,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Bench- marking the effects of operating system interference on extreme-scale parallel machines,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:20.006413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.384617Z digest=sha256:45f7b916e66a0689b182d44a14ed246ba340ea87f8a2b2d71af0c30f8aa7f4dd

Observation 1f047041-fa3f-4727-9554-0090e9e3532e · outbound

This paper cites Understanding the causes of performance variability in HPC workloads,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Understanding the causes of performance variability in HPC workloads,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.990984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.389165Z digest=sha256:4927a359eb964d1f65f0aba2072770e32c9e4f4d64b33e3f78771981662a302f

Observation 129c06b6-3e73-46a7-b6b0-c9eac695ac01 · outbound

This paper cites Overhead Measurement Noise in Different Runtime Environments.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Overhead Measurement Noise in Different Runtime Environments

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.466260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.392944Z digest=sha256:e454747505df6ca3bdd21c211ba713ad920d13ec65043e26fcbafa6e1d6e6c9e

Observation b5e9adb5-74f6-44ab-8670-c8f301c10ab1 · outbound

This paper cites Using microbenchmarks to evaluate system performance,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Using microbenchmarks to evaluate system performance,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.976627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.397014Z digest=sha256:b21330b0f0aae7166ee9aaadd05f2b37d428bc61d132ec7be5e0d1fa8b4a6168

Observation 308695cf-3a57-4dd8-945a-a254dd5502f4 · outbound

This paper cites Beyond inference: Performance analysis of DNN server overheads for computer vision,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Beyond inference: Performance analysis of DNN server overheads for computer vision,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.962037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.400903Z digest=sha256:c29bac2655887e1f85cf9c2a81aff1ba377feabcb46c73d7514d33f05c951d3b

Observation 199f6712-add6-4936-95dd-b41d3f3fd1a6 · outbound

This paper cites EdgeYOLO: An edge-real- time object detector,.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems EdgeYOLO: An edge-real- time object detector,

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.948887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.405290Z digest=sha256:794121dc69adf25a563b7e2911022909b45dd2285b1ef7b4c7757efb3a802e55

Observation 464bb5f5-a174-40fb-860b-14cfe70ae457 · outbound

This paper cites Images and vision — calculating image tokens.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Images and vision — calculating image tokens

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.935016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.409358Z digest=sha256:8b9460a2cec9468931a07f0879d6c8979077a4520e0aa41aaf0aed2e1988ae38

Observation 7d4d740a-2e86-46f3-b44f-b283bba3fcf4 · outbound

This paper cites Per-part media resolution (Gemini 3 only).

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Per-part media resolution (Gemini 3 only)

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.920442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.413398Z digest=sha256:1fbe4881e2c0b61b4806b20a7906c5af262a8f8b9a472c316bc741839e61e56b

Observation 279c4101-1146-4c54-a220-213d08828304 · outbound

This paper cites Gemini 3 Flash — model documentation.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 Flash — model documentation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.904265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.417282Z digest=sha256:a842b852d5718ad1079a93405c63bcca5ef1e68c0cd6cbf99046498260f1d13c

Observation 6e963d5a-b3a1-4dab-97b1-0a1dc045532d · outbound

This paper cites Gemini 3 developer guide — generateContent API.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Gemini 3 developer guide — generateContent API

Reference 97

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T00:50:19.877165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.424861Z digest=sha256:f033227cbe7b0dedcceef278d789325620a52eb5edbf0e9eeb0eb333fd57c945

Observation b0674577-89e0-4df8-88f0-db6abe44eb2d · outbound

This paper cites Accessed: 2026-05-28.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems Accessed: 2026-05-28

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:50:19.891144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T00:50:19.421183Z digest=sha256:173bc3475e5ce69929c0b2c300a922d63de11ee7254842aca0d699071f328d04

Pith citing papers

No inbound Pith citation observations are available.