Pith. sign in

Paper Citation Record · LEDGER

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

As of 19 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2607.22708.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22708 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:31:12.584059Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c8e6953-5b31-4ac8-b2ad-e0ac40f05df8 · outbound

This paper cites Flamingo: A Visual Language Model for Few-Shot Learning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Flamingo: A Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.163051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.163051Z digest=sha256:85051c66a8b55fc966f9cfb498d8d5a2bd19d6d4666ba9077d7f2f9fe30fed99

Observation d623437d-5790-4aef-a9b7-74a8ebfaa77f · outbound

This paper cites Visual Instruction Tuning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Visual Instruction Tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.167784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.167784Z digest=sha256:6b9d7b8e7f32d83d7a14e9274a87c2e365b958453ef6431d502db4e4e606ed6f

Observation 2af6481b-f764-47d3-90f7-9022fa122675 · outbound

This paper cites GPT-4V(ision) System Card.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design GPT-4V(ision) System Card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.172041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.172041Z digest=sha256:543a65221f75ee2b675ebd83d4fe5643a1d4c40c77bc651e38a73e83d3b3aad6

Observation f27a1725-7b9c-4872-bfef-6746455e6b8d · outbound

This paper cites Qwen2.5-VL Technical Report.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.175942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.175942Z digest=sha256:64544f692b79969a48293e4206384a8f5ccffd49a38339d48d96325fb122b9f6

Observation aafdf7b7-e115-4a26-a35a-bc270fd35c4b · outbound

This paper cites MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.180838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.180838Z digest=sha256:089b086b407c2b2971dd1c072545d7194fda038399fabacd87ec0287ad17a7ba

Observation 2bf44dd0-c283-44b5-8f90-6bbf260720ad · outbound

This paper cites Gemma 4 Technical Report.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Gemma 4 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.185014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.185014Z digest=sha256:01bc5aade1d6020581e833198289ffd7b92a701c9a512c24c3f3b00055f624a9

Observation 0bcbfa66-c464-45c9-884f-d815ade1a0df · outbound

This paper cites Qwen3.5-Omni Technical Report.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qwen3.5-Omni Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.189450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.189450Z digest=sha256:648bd1f0b229bb3660ca480d5f7e871df586355f1d9bd00331f7047705095858

Observation ee3993c5-ccbf-419a-a548-6bb8096c548c · outbound

This paper cites Mistral 7B.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.193467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.193467Z digest=sha256:343f401d9b9ca9ea003e9d6020aad5489e7590487cf7b4f699cb47a11c53da71

Observation fa1b171c-598d-49d2-950b-888644a0a6d8 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Generating Long Sequences with Sparse Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.196881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.196881Z digest=sha256:a856f7e02f2e278eab6316a2d08f53383b0a1393ac969894a14407bb0b39e659

Observation b44023af-e0e9-48df-b140-4de3eddafa76 · outbound

This paper cites GPTQ: Accurate Post-Training 29 Quantization for Generative Pre-Trained Transformers.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design GPTQ: Accurate Post-Training 29 Quantization for Generative Pre-Trained Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.201139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.201139Z digest=sha256:730309039ac4bc4e5e24d1bcd4faa4ee798b842f628887470a6a357b0a8e9188

Observation 157f9639-fbd3-4060-a9de-e4038f3cb09b · outbound

This paper cites A WQ: Activation-Aware Weight Quantization for LLM Compression and Acceleration.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design A WQ: Activation-Aware Weight Quantization for LLM Compression and Acceleration

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.205033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.205033Z digest=sha256:809baef804991df22671a3f86c2859bc881698899d71547aa8b973929d9e8039

Observation 277d40db-e98c-4a1e-9e62-c75bd536fd96 · outbound

This paper cites SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.208252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.208252Z digest=sha256:0e69d7eeb29eb1f73a4498a36aa2d8a068855e4b44e13abe572e8fd41ebe76a2

Observation 7f5eed63-a1c1-4330-84dc-f15c0dce9e89 · outbound

This paper cites Qwen3 Technical Report.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qwen3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.212067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.212067Z digest=sha256:a8b3199172c5b62fa86f5cb7b90524ca8ab83de21aa78292522fcb8fc7b9019a

Observation 67132d55-5d58-40dd-b621-6cd1a598cd11 · outbound

This paper cites Qwen3-VL Technical Report.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qwen3-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.216113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.216113Z digest=sha256:8d71382c0ca2b38ca4928e82b67681f04b28c7c3fa15b4d3b91c590e1ce01417

Observation 13082b73-57a1-48b9-80ca-8342fc4adae6 · outbound

This paper cites LLaV A-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design LLaV A-UHD: An LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.220775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.220775Z digest=sha256:299b5f69ac16b3e2fe4b6444e1d22a746956276bd2a9dd75399b9b4522155b2f

Observation 8a13f7d4-0371-4349-8911-369b685a8351 · outbound

This paper cites Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.224262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.224262Z digest=sha256:db4ff50a392990e40e8fee08f332166ce457ef11ac96c842ef3927b6587ed74d

Observation a27568de-930e-4054-8503-9238fb08251d · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.227338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.227338Z digest=sha256:321d077cc89911484cfc6e24bb49ada0062d1bc86173fc141e859eda9e723952

Observation b69c3126-9032-4e9a-82fb-0f4fe6d89e35 · outbound

This paper cites ICDAR 2019 Robust Reading Challenge on Large-Scale Street View Text with Partial Labeling (LSVT).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ICDAR 2019 Robust Reading Challenge on Large-Scale Street View Text with Partial Labeling (LSVT)

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.230406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.230406Z digest=sha256:14c511283023b192e763e4f8e46d9a3f309fde06c8958d6f1b78be454612ab57

Observation fecceb71-0ec1-44ba-9f81-a2fd983b0ece · outbound

This paper cites Seven Problems with the Claims Related to the Hubble Tension in arXiv:1810.02595.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Seven Problems with the Claims Related to the Hubble Tension in arXiv:1810.02595

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.233576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.233576Z digest=sha256:28e6c572a76c2a8f73397310182cf67998baef5068dbf479ff3a53fa301e05ca

Observation aedb5422-ff5a-47ce-84cc-1f9660204bbb · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.237847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.237847Z digest=sha256:0cd4d72d08cb866a56c5dd78e47ee55380f4c30740753860f5c2a4287f038aaa

Observation cbf48ff2-8655-48fb-9819-46dc22835696 · outbound

This paper cites Synthetic Data for Text Localisation in Natural Images.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Synthetic Data for Text Localisation in Natural Images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.242136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.242136Z digest=sha256:2c44fc1b3ed1534691fd6e83fec92aecce8c0465a9cb8be014cc359b4d16c3d8

Observation 899e6d8b-e429-41b6-94a5-de61369ac4a6 · outbound

This paper cites Tex- tOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene Text.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Tex- tOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene Text

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.245463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.245463Z digest=sha256:8b4e73be4a456381bb6f603d2141934a27971a5362d4311922802f297312c1c0

Observation 92839a7a-e3be-4c4a-ac1e-0e5462b5b88c · outbound

This paper cites Jawa- har.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Jawa- har

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.249440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.249440Z digest=sha256:caaedfe6aea5a8439682c79bff865ec5eb146111c98e21df2068593348e232ca

Observation ea94e2c1-710c-4187-a9eb-ddc9b654f1a4 · outbound

This paper cites XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.253558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.253558Z digest=sha256:cf3bfd29567655be6688e78618be8e73065f85007c2d78920b6fdcb9eeffaea1

Observation ab0557f8-251c-461b-955f-aadbe7129cdb · outbound

This paper cites FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.256919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.256919Z digest=sha256:6d8dc97ea8f7d8765b66cb318190170283d811fb6395be5e0db90ac660d5db45

Observation 0f45d8ff-9123-4659-91e1-a5216e38cd82 · outbound

This paper cites Image-Based Table Recognition: Data, Model, and Evaluation (PubTabNet).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Image-Based Table Recognition: Data, Model, and Evaluation (PubTabNet)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.260891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.260891Z digest=sha256:29ab07a3465540a4d299043c56f4901c303ed946c4139b2bc6905a9e7c3227c2

Observation 327a0edc-4c53-434a-9e33-80f92e42df3c · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.264506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.264506Z digest=sha256:9017f82591f37495d01d7bd11b6b39c25d93885dc531d84e3f8880ab0862298f

Observation a902841d-c1ec-4fe4-a833-ce3415268967 · outbound

This paper cites CASIA Online and Offline Chinese Handwriting Databases.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design CASIA Online and Offline Chinese Handwriting Databases

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.268002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.268002Z digest=sha256:4058879def5e65b5a9dc68606f5eb22c90177fcfec110e9568dab50b4ad82634

Observation 4d2f8b5a-081e-4fd6-bf1a-1ffa9ec29460 · outbound

This paper cites Synthetic Data and Ar- tificial Neural Networks for Natural Scene Text Recognition.NeurIPS Workshop on Deep Learning, 2014.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Synthetic Data and Ar- tificial Neural Networks for Natural Scene Text Recognition.NeurIPS Workshop on Deep Learning, 2014

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.272620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.272620Z digest=sha256:54378084d08a96b082fc3e7266d12e272b2f66972996404e6d6cb71168a4dc3f

Observation da5367cf-bb71-4807-b92f-93e9bb264af9 · outbound

This paper cites Industry Documents Library WebDataset (IDL-WDS).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Industry Documents Library WebDataset (IDL-WDS)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.275617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.275617Z digest=sha256:40d70fcb203e1b2dd657a6f24bde795c5df0e41349960791c533baf9f79e47da

Observation 879067f2-dcc1-4eb1-8501-360a96a7c8f2 · outbound

This paper cites Towards End-to-End Unified Scene Text Detection and Layout Analysis (HierText).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Towards End-to-End Unified Scene Text Detection and Layout Analysis (HierText)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.279624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.279624Z digest=sha256:f55a10b7099c74271769255f4bdb2795f928450d0a60ddbfc45ed1cf1e09fce1

Observation 11ae5cf8-52a0-4409-9b95-76b4bb4e1930 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.283181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.283181Z digest=sha256:81fd1815c81c4653a41ce89d74251b48e897f0dd4db4ad2237dd167c1993bcfb

Observation db8235ac-0579-45c3-84ea-dcf60d4b40ca · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.287336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.287336Z digest=sha256:ca7be4dada5637b50315be49c0015093d6a4acd7fdbd32b0b4b8f4f6842b2c17

Observation 9cc58260-158e-4a6a-8744-f9783e95a48d · outbound

This paper cites Wave-UI-25K: A Large-Scale UI Element Dataset.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Wave-UI-25K: A Large-Scale UI Element Dataset

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.291766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.291766Z digest=sha256:50423a1b709aa80ee5f7f8ff96df949525b02545ef98d51d2c8bdd97cb76e46b

Observation 7c0f49e6-6962-4d7e-bd80-d8cc05cf163b · outbound

This paper cites MobileViews: A Large-Scale Mobile GUI Dataset.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design MobileViews: A Large-Scale Mobile GUI Dataset

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.294985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.294985Z digest=sha256:08f34a749b9c50139ea3061163b2b85620c307559e2709fec26be4864a55b9fe

Observation 3a459686-be9a-4794-8612-420a1fd5b8ca · outbound

This paper cites Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.298496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.298496Z digest=sha256:f99a32eaf25fc869f6f3ea630904ab2268d3f46f1d60f6710ff84bd6d992619c

Observation caadc118-e7b5-46a9-a0da-ccd54efd1929 · outbound

This paper cites On the Effects of Data Scale on UI Control Agents (Android Control).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design On the Effects of Data Scale on UI Control Agents (Android Control)

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.302292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.302292Z digest=sha256:cd26a0ff9656638d1f23f2266c912560755148edcc158b364bc93f48bad2b148

Observation 31eca24e-9ca4-40ad-8189-24b81e0d0b25 · outbound

This paper cites Android in the Wild: A Large-Scale Dataset for Android Device Control.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Android in the Wild: A Large-Scale Dataset for Android Device Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.306246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.306246Z digest=sha256:4713ea2b03dda53b9021885b84dbc3d5165d29bd539cb7964c6602e66bd27fc4

Observation 0309373b-7ffe-41ae-81b0-9fafd6cccf61 · outbound

This paper cites AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.310475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.310475Z digest=sha256:59596b1ced3f9c7519d4520255da8efacc6463072a37e8a75e98c9a32bddd0a1

Observation 84927a1b-ee8b-4cec-b087-7535b26a95e1 · outbound

This paper cites GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.314255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.314255Z digest=sha256:3ae50eac1a521c4731f3828ef7badfb2f4f07f8eb944685b413d37a81e2eb1bf

Observation a055f88c-85dd-41e8-819e-a927e538c56a · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.318215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.318215Z digest=sha256:fdd72b6b490a04c9e92259220d7c4366937dc5b9aab208d6b4143e14fa7ed322

Observation 91d7a79e-bac6-4579-9f25-5592099b4f2e · outbound

This paper cites ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.323618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.323618Z digest=sha256:9f16a85c936858958328c9958698e366e23d114b6e4e4ca182581efb9129dc76

Observation 4a22167c-5dd6-4ea1-874c-24ad3cfcce60 · outbound

This paper cites an unresolved cited work.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.327302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.327302Z digest=sha256:327ed1b0ba70662b1e254a7f1a7949f5914d0bcdfb8204151cce102d14ab158e

Observation 87c17104-d164-45ad-90ac-83f3a15bdd4e · outbound

This paper cites an unresolved cited work.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.330625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.330625Z digest=sha256:5cf06306a5c89891b8d47d6f2db848df8832cbb6d6fd5c620fcbddd0b0de8f49

Observation 8b3485b4-b37a-4157-9271-d9063f3a375f · outbound

This paper cites ChineseDocVQA: A Chinese Document Question Answering Benchmark.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ChineseDocVQA: A Chinese Document Question Answering Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.334261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.334261Z digest=sha256:8ec33ce038ed6470f1e1f333f1915340b37d8e9124d9c01a876a3dbf6a3efad1

Observation 1388f081-3b65-42a1-b04a-a13eab3cebfb · outbound

This paper cites ChartQA: A Bench- mark for Question Answering about Charts with Visual and Logical Reasoning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ChartQA: A Bench- mark for Question Answering about Charts with Visual and Logical Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.337528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.337528Z digest=sha256:6826876cb0aa5a253be6e5272f586e15fc823efbb86f6b3ad576a246bdd3ea57

Observation 98814a26-da95-4d5e-9cad-b086e1462dfb · outbound

This paper cites Towards VQA Models That Can Read.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Towards VQA Models That Can Read

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.340921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.340921Z digest=sha256:e98cc33112e663da2f3df9068fe0d83ec35fb8ba90050a01c30242aa0ba34696

Observation 0f1fb5fe-2c2b-4196-83af-2313e2b49f2f · outbound

This paper cites Jawahar, and Dimosthenis Karatzas.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Jawahar, and Dimosthenis Karatzas

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.343909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.343909Z digest=sha256:e084c4b4098c0f8db7ff95c04edb52b2d93926ac2d5ca5f514ab04c915e71ca8

Observation 27d6c8cf-ba30-45ea-8d95-6ea11dd20c54 · outbound

This paper cites OCR-VQA: Visual Question Answering by Reading Text in Images.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OCR-VQA: Visual Question Answering by Reading Text in Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.347245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.347245Z digest=sha256:a9e98f5653685662127a417224600dbe46ce7eae554ac0c5dc3c68078ac6d13a

Observation 915c1431-d9cc-4bc1-aebb-41a0e707528d · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.350522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.350522Z digest=sha256:715fe5013d2e6b60be0f5d7484400b91f3e25ddd6ebebef80787d722dcfda99f

Observation 88f3e7ec-5781-4205-9046-69f9909e6932 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.354687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.354687Z digest=sha256:433ad3b5ae355b7809533cf3af59cb2443290f43359a569907e4f4f7d556f199

Observation 05dfab41-1e55-47cd-8cb4-7b8ea17466e6 · outbound

This paper cites LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models.Advances in Neural Information Processing Systems (NeurIPS), 2022.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design LAION-5B: An Open Large-Scale Dataset for Training Next Generation Image-Text Models.Advances in Neural Information Processing Systems (NeurIPS), 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.358331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.358331Z digest=sha256:1cee4fdab3e85d2da2ba48b5ac78d71b24508763a6010e6d7218ae83add92170

Observation 389c4a9a-f7e0-4ff2-9ee5-61702507223e · outbound

This paper cites Conceptual 12M: Pushing Web- Scale Image-Text Pre-Training to Recognize Long-Tail Visual Concepts.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Conceptual 12M: Pushing Web- Scale Image-Text Pre-Training to Recognize Long-Tail Visual Concepts

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.363826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.363826Z digest=sha256:a981d3606eea7109757c7539fc13e874decd3c6fd7556bbfb62f9449ecd19550

Observation f3a35fbd-f121-478c-91c2-91e5284e213e · outbound

This paper cites Wukong: A 100 Million Large-Scale Chinese Cross-Modal Pre-Training Benchmark.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Wukong: A 100 Million Large-Scale Chinese Cross-Modal Pre-Training Benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.370703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.370703Z digest=sha256:db10840c00223fa4b4b09defd5b21f7be958bae439b6f9ee2634e960fe18008a

Observation 4a0b70be-8b7d-4f4c-af5e-b43b294917e3 · outbound

This paper cites COCO-CN for Cross-Lingual Image Tagging, Captioning, and Retrieval.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design COCO-CN for Cross-Lingual Image Tagging, Captioning, and Retrieval

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.384859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.384859Z digest=sha256:a66bc5cf82f84a4296fa5f82bd041ac91f361eb81ecf4459de7baddd5683d201

Observation e3a060c6-b187-420d-8de6-72f705ad4bc9 · outbound

This paper cites Flickr30k-CN: Chinese Version of Flickr30k.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Flickr30k-CN: Chinese Version of Flickr30k

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.399767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.399767Z digest=sha256:c1af74fad2603cb4c80388377e9dcaf26a7df0c7de9b860336ac526c5a7e999a

Observation fc030710-7f3d-4c6a-9c50-c2a7651aeed0 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.411882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.411882Z digest=sha256:da7ef15de5be7f93889f15882c9d763232f7919b7fea05b9c368fd92d0546feb

Observation e6778c1e-0c14-4df5-9266-c6ded23775cc · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.424454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.424454Z digest=sha256:a59d00c3b84b0fa79e75e3d966ed8143e1638d6607ff7fc8a2fb34109179ec3a

Observation c751b3c5-70c5-455f-b00a-0cf0302719e4 · outbound

This paper cites ShareGPT-4o: Comprehensive Multimodal Annotations with GPT-4o.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ShareGPT-4o: Comprehensive Multimodal Annotations with GPT-4o

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.435425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.435425Z digest=sha256:af9f34db5324405408ab17c571f48ad804a0185024a03f7a75ce5337d7a917aa

Observation fe6631cc-985d-41ee-a96b-da517e43bb84 · outbound

This paper cites SVIT: Scaling up Visual Instruction Tuning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SVIT: Scaling up Visual Instruction Tuning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.447932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.447932Z digest=sha256:daeea7a96bc572a05ce6d9f25783c81bc38c46825208a6b36b702475c745d730

Observation 2b61bb21-d93c-4668-9318-9a3bec3ffd8a · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervision.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Learning Transferable Visual Models from Natural Language Supervision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.460150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.460150Z digest=sha256:da6c125f469e8d8486e1d5ab6805deb180a27bb4db50800efa09648f3efe652c

Observation 6865b9ac-9130-4815-8cf5-aa5db9f1075f · outbound

This paper cites Implementation and Benchmarking of Perceptual Image Hash Functions (pHash).

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Implementation and Benchmarking of Perceptual Image Hash Functions (pHash)

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.472649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.472649Z digest=sha256:3b5ee606bfbc07bb50ef522da06cadd92152a2b0a87f2d1f4cf2ec9aad5a9596

Observation acf07bab-cb36-4b96-b0da-8e49b912cf2c · outbound

This paper cites an unresolved cited work.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.483109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.483109Z digest=sha256:a32b9e0a7b4078c9efc0ca8ea881dbc166fcb17000ac9c207137687e1cc8c667

Observation fc91d887-15ab-426c-84d0-ec2999ffc4ec · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.499913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.499913Z digest=sha256:8e6d3b726a98c1f1b83122d10a45b020db471cdbb1345f13a7ddca7e71a58dac

Observation 19b876e7-b1ae-4b02-9659-5cc140b6f77a · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.517595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.517595Z digest=sha256:a12886cbe908fe980dde08e7fd63b475f799aacba3bb496d6b771c7e231b0546

Observation da69dde2-7265-4a01-898b-666d18034e74 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.613077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.613077Z digest=sha256:1d8257a7f1a70b4233832fb19906c416631c97fd82a55689bf396f956a21ea8b

Observation 4403b786-7262-4849-adda-33fb400c4398 · outbound

This paper cites RLPR: Extrapolating RLVR to General Domains without Verifiers.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLPR: Extrapolating RLVR to General Domains without Verifiers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.704541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.704541Z digest=sha256:345e453e707486d6b8838199b5134a82a9935d405aa8961038e2359b8b3e6297

Observation d275308a-bed2-4e73-82fa-d76d418983ad · outbound

This paper cites RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.836000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.836000Z digest=sha256:7cf1a6639b608299dd2d75b440573927ef069c138970dd27387e6ca7b82d12e1

Observation efa52f14-ce52-46e0-98db-3aa2b84f35a5 · outbound

This paper cites Manning, and Chelsea Finn.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Manning, and Chelsea Finn

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:11.946959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:11.946959Z digest=sha256:ccb523ae27e538064541578d53290a41da519b07dbb59ff7520bbb6e05b47a21

Observation ddbb1c71-25a5-4db1-a048-2ef1251b8e29 · outbound

This paper cites Zico Kolter, and Zhuang Liu.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Zico Kolter, and Zhuang Liu

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.001958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.001958Z digest=sha256:a2fa343cea573061802b732d18e89e3f8b4b962cb2dafe8de7c9d80cb7b75bbc

Observation 10b596ba-bdef-4d85-acb3-17fa3d25dca0 · outbound

This paper cites Qualcomm AI Runtime (QAIRT) SDK.https://www.qualcomm.com/ developer/software/qualcomm-ai-engine-direct-sdk , 2024.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qualcomm AI Runtime (QAIRT) SDK.https://www.qualcomm.com/ developer/software/qualcomm-ai-engine-direct-sdk , 2024

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.052313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.052313Z digest=sha256:e9c86e758aaa748a1eb898ea94de0ce4da8f0a9649982554878a65718e8cc90a

Observation 4888ccbf-44c0-4fb9-8df2-0f2e858a6659 · outbound

This paper cites Hexagon Tensor Processor (HTP): Qualcomm On-Device AI Accelerator.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Hexagon Tensor Processor (HTP): Qualcomm On-Device AI Accelerator

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.139676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.139676Z digest=sha256:94241252a9d88dea5a0871a7853fca8dd52e99c9f23bb4cc3ffba80854c169c4

Observation 4dfb00cc-24ce-4a7d-98ac-67cb90c3ea9c · outbound

This paper cites Qualcomm Genie: AI Runtime for On-Device Generative Models.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Qualcomm Genie: AI Runtime for On-Device Generative Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.235795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.235795Z digest=sha256:3062752149802fe393deb5fc2fa63b0bfd6abf9481ded39e93a3996e2f315fc5

Observation 3bf75391-926e-4235-8cdc-dab1a3c4915c · outbound

This paper cites ONNX: Open Neural Network Exchange.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ONNX: Open Neural Network Exchange

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.318022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.318022Z digest=sha256:9a147dc07b5b60484fe0dc43f99120fb7f12623a02d0b807586eefb9b0f9af59

Observation 57296692-7d1d-4526-9fb4-4b8dc2b784d4 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.377437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.377437Z digest=sha256:a6bb487f8d57d1da6cc4d767143af19d4a1463767e3fba4a44d3cb9bd36e47c7

Observation 68b888a8-fd2e-4524-8105-30d343422d36 · outbound

This paper cites Paech, Paul Pak, Rom N.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Paech, Paul Pak, Rom N

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.430191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.430191Z digest=sha256:4265e48e15d3a76467839d83023a4cf625b2dbd1b7c41417448486d4eddd9b0a

Observation e6c76062-526e-49d9-a39e-f2b276ac32ac · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.485226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.485226Z digest=sha256:104a35b5f0f7efd84d3c2340ec91186704524679515e4c63f8af7c1819c9984d

Observation 9f6cffe7-2dc7-48dc-89c1-eda5a011adbb · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.518516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.518516Z digest=sha256:eef285a1bddb32ffce1d21342ce99940779bedd772f45b13ce3b6bf72169a2e0

Observation 62460a42-b29e-4695-83c3-e89de74c20b4 · outbound

This paper cites Berg, and Tamara L.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design Berg, and Tamara L

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.549423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.549423Z digest=sha256:ea69ce863a4c770ab86819ab537487b293e7a3dc448ac434956687c8535c7aed

Observation bb25b327-c346-478d-8158-2dc600fbf812 · outbound

This paper cites ࠞ”ཟb ശđ໡ᄝ “ࠞת.

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design ࠞ”ཟb ശđ໡ᄝ “ࠞת

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T16:31:12.584059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:31:12.584059Z digest=sha256:e995dee205854be0d531e822b5767caeff05b01935c5742bb7f07924c86a10aa

Pith citing papers

No inbound Pith citation observations are available.