Pith. sign in

Paper Citation Record · LEDGER

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2607.14675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14675 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:28:01.619301Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved61
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2ce58ce-6303-49e0-b331-15fa5d0c434c · outbound

This paper cites You Only Look Once: Unified, Real- Time Object Detection,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots You Only Look Once: Unified, Real- Time Object Detection,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:54.965088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:54.965088Z digest=sha256:6aace52e63927856ccdc21ef73b001fc0ca3f1948c58e35074e2f6a80152eec0

Observation 0206d837-1c8c-4fa5-b818-a84b87fb1d35 · outbound

This paper cites YOLO9000: Better, Faster, Stronger,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLO9000: Better, Faster, Stronger,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.106710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.106710Z digest=sha256:7ec575bd571162fccdbb93fd2024cbdba4b1596378552bd8e5ed51d4c4800603

Observation a487f2e2-b5d7-48bb-a547-648a6e9f1e2d · outbound

This paper cites YOLOv3: An Incremental Improvement.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv3: An Incremental Improvement

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.249830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.249830Z digest=sha256:3cbffc0539cca5537b717d8ece4e27174690bc2796c018db9734fbc7c6e2f949

Observation c89d9af4-6478-493d-9f79-e71424803a66 · outbound

This paper cites YOLOv4: Optimal Speed and Accuracy of Object Detection.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv4: Optimal Speed and Accuracy of Object Detection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.366748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.366748Z digest=sha256:374e7df06647a64c0d731676ff59e48c76d5fa8c972c8ef2c2cbf60bcd6e16bf

Observation 2b00bf5b-0d3b-409d-b137-cd77950ab7c4 · outbound

This paper cites YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv6: A Single-Stage Object Detection Framework for Industrial Applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.468740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.468740Z digest=sha256:df88890ceea8ae4a10321f033b97f8ead10a813d4d58e82a78a43ccf0eced6d1

Observation 407bc2a0-34e8-4795-bef0-bfa14a1643c7 · outbound

This paper cites YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.595785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.595785Z digest=sha256:b2c5c3f3da2903a8d3f138d4a9146d6cc0983d4cbacc259860b080ccd0c6b18c

Observation 6a9d5fba-216c-41b1-8a52-62069ad65662 · outbound

This paper cites YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.698686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.698686Z digest=sha256:af108fa10b81cf3a0161e39d294a1e8848c771724c3af5cfb8873a7d2056205b

Observation cd4ba6c9-1851-4e57-8835-62513c752426 · outbound

This paper cites YOLOv10: Real-Time End-to-End Object Detection.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots YOLOv10: Real-Time End-to-End Object Detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.781407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.781407Z digest=sha256:1a4c10180bd243845221b9e18ee261368e29f5ebca77328692124ece3ee14cd2

Observation a77f38f0-843c-4166-98c5-342bf9827cda · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:55.929803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:55.929803Z digest=sha256:2a82de40c1be50d4e04b7ffc385a4361c9f2d0d77bce5a31e38c5f6239bf15f6

Observation 76d9d287-dd28-4a99-b21d-680bdcffcd2b · outbound

This paper cites SSD: Single Shot MultiBox Detector,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots SSD: Single Shot MultiBox Detector,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.089146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.089146Z digest=sha256:47dfa8a54736fc7d980282166b27a131c38f5e0ac8250fbeeb3688eac17e0d16

Observation d260c4ba-b770-41a1-bcb5-d4f08cda10b4 · outbound

This paper cites Focal Loss for Dense Object Detection,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Focal Loss for Dense Object Detection,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.263469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.263469Z digest=sha256:d57983c97d6a4de890c4659af9f47344c5904fd2421e5c279b84594e3dc1d00d

Observation 2dac9ae9-56de-4b5a-a5e8-5b89281a07c9 · outbound

This paper cites End-to-End Object Detection with Transformers,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots End-to-End Object Detection with Transformers,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.402511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.402511Z digest=sha256:c60ebbc6bf50cf339d8406937ff72acda4da3e935df92504622ee93196a22ccf

Observation 68422a9d-4ffa-4f37-9279-f61b9b32d2cf · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detec- tion,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Deformable DETR: Deformable Transformers for End-to-End Object Detec- tion,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.550576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.550576Z digest=sha256:1701153fcb6e6f55436f2508a492a870b52eab0c398c02735bef1db782b10dda

Observation 8cfe396d-30ac-4cf7-af10-16442898df59 · outbound

This paper cites Mask R-CNN,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Mask R-CNN,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.671701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.671701Z digest=sha256:8cde20b711195330e3e50330435053383028f211a958369f81339ae31f564b63

Observation b8c01bb8-a2b2-4453-8fc9-d3554e66e97b · outbound

This paper cites CBAM: Convolutional Block Attention Module,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots CBAM: Convolutional Block Attention Module,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.826527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.826527Z digest=sha256:3d3a67844d7789ab6d0a0bf2f5fe6fe8951025b0ae5f79bbdfca6466afa36e63

Observation 98139c05-7f86-4599-b1c6-a5b9fa4a6700 · outbound

This paper cites Squeeze-and-Excitation Networks,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Squeeze-and-Excitation Networks,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:56.942471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:56.942471Z digest=sha256:4452da95cce6da5435e37433734bae41bcb777c9a3fb7e953cbe7acd854b1324

Observation a1165607-0ed8-4580-817a-495fdec6f434 · outbound

This paper cites Attention Is All You Need,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Attention Is All You Need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.040193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.040193Z digest=sha256:c91f210b762eb0ea8d88183ebf51e08c0fa543ff52010c9e63538b48c58c147c

Observation 55301acf-0213-451a-ac3f-9d29dc05aaa4 · outbound

This paper cites Non-Local Neural Networks,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Non-Local Neural Networks,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.157890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.157890Z digest=sha256:17da5dd00dd23a77eba27e18156aba395c0ea7cfd0f1159605f6cbc709a916d0

Observation 6a4c3653-b680-40f5-bfe9-02fdb960c4c4 · outbound

This paper cites Coordinate Attention for Efficient Mobile Network Design,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Coordinate Attention for Efficient Mobile Network Design,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.235872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.235872Z digest=sha256:11c481e4fa5d84122c08e3f9a972213ae80d99e6c4dfdb4825da37e0a0a69d44

Observation 7ea68dd5-6f07-404e-99d2-98b9f96d7af7 · outbound

This paper cites Generalized Intersection over Union: A Metric and a Loss for Bounding Box Regression,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Generalized Intersection over Union: A Metric and a Loss for Bounding Box Regression,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.316228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.316228Z digest=sha256:3e40a0b2eda43e517251ddc743f0de311026e3b4854276f7e7fb47a05134b32d

Observation 0ab16a1d-43af-4054-ac52-f0bf57dd9940 · outbound

This paper cites Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.382839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.382839Z digest=sha256:7a5b60a87110c0326d152969e12c063e225b30c3221e0a3f10ffed60d4cfde3a

Observation a9f037dc-066a-4ed6-8f47-40540d7f5c77 · outbound

This paper cites SIoU Loss: More Powerful Learning for Bounding Box Regression.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots SIoU Loss: More Powerful Learning for Bounding Box Regression

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.440376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.440376Z digest=sha256:2874e99687bdc2b997e06a72ab9b12850c48837fe13ffc2e23b97bcd96c5639a

Observation dbda39a1-d9b3-42f8-ba47-9e2cc9d94e94 · outbound

This paper cites Focaler-IoU: More Focused Intersection over Union Loss,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Focaler-IoU: More Focused Intersection over Union Loss,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.505845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.505845Z digest=sha256:49ea72e621485b1dcf1e807ccf5bd418029efedb3b11a2830ffa2a0acf919a62

Observation 793f9ed5-c70d-4ad6-9b2f-77b75a5bbe64 · outbound

This paper cites UnitBox: An Advanced Object Detection Network,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots UnitBox: An Advanced Object Detection Network,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.575970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.575970Z digest=sha256:205592307467a83e6191fcf06fa61fd2d83c22021eeb32bf6c825f13dec8642e

Observation d3599dbc-274c-4dd4-8d32-b9c8091d3d91 · outbound

This paper cites Language Models Are Few-Shot Learners,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Language Models Are Few-Shot Learners,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.649886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.649886Z digest=sha256:5dcd11bfb82e11402df5ce4594d0b2d52216cdba96af1de44942291117d29752

Observation 12d0812a-1545-452a-ad37-2fb9689e2c93 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.731195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.731195Z digest=sha256:40fa6a7363d86c5ff00d66339aab9dbf6efd8f73d0c3aaaefc7fed3992e363db

Observation 2cb89733-8871-4351-b369-3d2eec495057 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.813669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.813669Z digest=sha256:9627434abf300c5e6ca088ed0fbedbc6832eef0ee693c9aacf14e580a2c4ef8e

Observation f540a740-09be-4290-aa15-15b567122c12 · outbound

This paper cites GPT-4 Technical Report.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.898391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.898391Z digest=sha256:e5d74fe45a2de36b5752d5d1647dc56d00f997c8a736e3367a1b767518635463

Observation d4659746-3ce0-44cc-8266-f5ccbefbdade · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Chain-of-Thought Prompting Elicits Reasoning in Large Language Models,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:57.978487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:57.978487Z digest=sha256:cf6922f08257aeebcf2de3ee39c759b64b9717201e3c78af1df8274cadce412a

Observation 2acdb142-bac2-4f57-b7c9-b4dad7c1115d · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots PaLM: Scaling Language Modeling with Pathways

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.065543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.065543Z digest=sha256:61e6f0b414fd99bc371048683579abbeb468f9e09bdbc03cc3dc874e79d73258

Observation cf795011-e53c-40fc-adbf-7c1e059732ec · outbound

This paper cites LearningTransferableVisualModelsfromNaturalLanguageSupervision,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots LearningTransferableVisualModelsfromNaturalLanguageSupervision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.150328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.150328Z digest=sha256:ef002ca8fe73a318b3df69a5a27c19931db34ba9e1b1554255fd1b3f48697370

Observation c72cb270-fe74-4077-84f8-c42a3af9c2ad · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.161162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.161162Z digest=sha256:6779672cc1c56e3a2acf283999d9355e8b7a788dfe12ae9174eee024882a93ec

Observation ad04f198-185a-4e8a-8bdd-43d308cc2fb8 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.299457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.299457Z digest=sha256:f110878b57ca1941a62c49e8ddfb2d178dfee51047e740d3cc0fe8a0a035a3cf

Observation f6c4cb4b-82d9-45f3-bd8c-53863d05ebf1 · outbound

This paper cites Visual Instruction Tuning,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Visual Instruction Tuning,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.336581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.336581Z digest=sha256:80613731461fa3a379602b3712320a20fe58af04eee55c83047b19702fda20c9

Observation af88eb50-f6c2-4e5f-b046-ce4594eab5ca · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Improved Baselines with Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.446531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.446531Z digest=sha256:2773149ca2049dd7131d8d7acc0eef93f1d307003666a8a90adbc84e896fdba7

Observation ddd7f4f1-911f-4956-ac4b-906c8af9e190 · outbound

This paper cites Flamingo: A Visual Language Model for Few-Shot Learning,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Flamingo: A Visual Language Model for Few-Shot Learning,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.552364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.552364Z digest=sha256:73c0d764020ac4f3839593914903d8abe4c5146b05fed8351b96e0789737499d

Observation 0b6953cf-7bac-4e9a-a63c-84bfe5177377 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.680234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.680234Z digest=sha256:c0ee99995e35810bed455e1d136160505c03929a54fe1b55829801eb2ed3f027

Observation 12d98d6d-6fea-4685-8478-e7d6cab107b1 · outbound

This paper cites AnImageIsWorth16 ×16Words: TransformersforImageRecognition at Scale,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots AnImageIsWorth16 ×16Words: TransformersforImageRecognition at Scale,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.794645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.794645Z digest=sha256:9a31cd6c22aa006b5bdf1869e41811395ab3b4ef410e14ed54452919a86323a2

Observation f32b98a4-57e7-41ca-a205-f4e25e6c45d2 · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Masked Autoencoders Are Scalable Vision Learners,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:58.904736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:58.904736Z digest=sha256:8a27fe7ec37c2e821c1325cc8e00c720c3ebd35eed80ec7989609f40f8c327db

Observation 6ad2960d-2457-449c-bcb1-05c0c9c928cc · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Do As I Can, Not As I Say: Grounding Language in Robotic Affordances,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.041232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.041232Z digest=sha256:aa62df3128172a9e4f98737f13ff9e3c59220c63d87919d67e149024e9e2b77e

Observation 97fe3b86-7c0b-41be-9b84-54018a558a9b · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.207004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.207004Z digest=sha256:35baefc649c26b2782d2044f1a2e1731347d6be8defd562a09f9194869a38fcc

Observation d42dd1ed-68d4-4df6-8588-42e44d6d2b8f · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.373066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.373066Z digest=sha256:e66ef6b815d9883aa065f2dede3123037e9b42f861680c061a091ce8a6693c8d

Observation fa5bc69e-1a46-41d5-881d-effeae353d2a · outbound

This paper cites CLIPort: What and Where Pathways for Robotic Manipulation,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots CLIPort: What and Where Pathways for Robotic Manipulation,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.492385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.492385Z digest=sha256:9a49544aee4f64570f48e02a1c0807cfa2767e47f2425d7ea3258507d881d401

Observation dd8244b5-e65f-4d3b-b105-97a1f3eb9d82 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots ReAct: Synergizing Reasoning and Acting in Language Models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.638626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.638626Z digest=sha256:268756763ce2919c2a86ed2e8bbdf099a3521b66eeb17c188c1a9e63f167eb65

Observation d871c382-15c8-42c0-bcd2-c9f7e2f615d2 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Generative Agents: Interactive Simulacra of Human Behavior,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.804962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.804962Z digest=sha256:5e942e84c92bb351d2f32fd5a7cfbbc52c65a6f9b9f5f81abab48ff2d7dfab29

Observation 41ae76cf-4d39-4d0e-8a41-4fe04d1ea2e7 · outbound

This paper cites Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:59.970912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:59.970912Z digest=sha256:ed37aadd16a0a74ab965237f09aff5ee8393ab36adbf994a14237c167eba70b8

Observation 2c7cf377-d878-4898-9ca5-f2ecd97c45e0 · outbound

This paper cites A ConvNet for the 2020s,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots A ConvNet for the 2020s,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.131308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.131308Z digest=sha256:7b48e65bd2d6b6add48fb052eded34a7b18bf6da85c89c26c93027275a3c94f0

Observation e15cf791-f9dc-499e-9dc7-3755f6664649 · outbound

This paper cites Deep Residual Learning for Image Recognition,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Deep Residual Learning for Image Recognition,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.239227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.239227Z digest=sha256:7e66cf2653b4dafe8284aa2b8a932e1ecf5c9392325ad9d30a4d64005da22e90

Observation 83d1710c-361b-4c41-8bf2-eb3b87186448 · outbound

This paper cites MobileNetV2: Inverted Residuals and Linear Bottlenecks,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots MobileNetV2: Inverted Residuals and Linear Bottlenecks,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.397659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.397659Z digest=sha256:2abd0b508e69e7c07e13e3b0363df721cabd1b9cd36f6b97e6e8d9c8a77b2cb2

Observation 0cbca189-8953-4ecb-8706-e7719ed4ca05 · outbound

This paper cites Searching for MobileNetV3,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Searching for MobileNetV3,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.564351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.564351Z digest=sha256:37a4c45fc7ae9dc033347b5ce8a3cf8f8ff7ced79a89f32195a72fbf805e7ec3

Observation 80647a6e-89a9-4ad6-a465-ba8ed57febbb · outbound

This paper cites EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.731294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.731294Z digest=sha256:94351820a7a4394510760c360310424f4bb120bda63b72968962be894dd8c15b

Observation 8b0b7568-45b0-47da-b0fc-96b32c6b1cda · outbound

This paper cites Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Edge AI: On-Demand Accelerating Deep Neural Network Inference via Edge Computing,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:00.892238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:00.892238Z digest=sha256:66745c6fe4fdeca54169bb46ccbae9450efe0247b7932291bc0f550b005d0359

Observation e50a6a68-e2df-4860-b409-3727791d5e46 · outbound

This paper cites MLPerf Tiny Benchmark,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots MLPerf Tiny Benchmark,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.004357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.004357Z digest=sha256:eb2ba8f5ac256d32883978a2af450d4ff72c6022044710af9038a5ca53fbda99

Observation 1920287e-6f0b-45cf-898c-e0dce011808d · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Very Deep Convolutional Networks for Large-Scale Image Recognition,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.083960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.083960Z digest=sha256:896b1dc29dfd41f54d35f8d6c9292ccc413166df6260a56f7a964f03bd26bd10

Observation 8be47a5c-48d2-4fc2-b1f5-eb5735cdb833 · outbound

This paper cites Rethinking the Inception Architecture for Computer Vision,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Rethinking the Inception Architecture for Computer Vision,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.176513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.176513Z digest=sha256:13fc126d645a07e18dc5a932725f9be854945e02c194f6f5f9b79ad6e4fbad70

Observation 055f7727-9abc-4f90-b64f-43f36f599023 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots DINOv2: Learning Robust Visual Features without Supervision,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.256440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.256440Z digest=sha256:559ed54ef6c1692a16bf379d4d545fb67eddb84103688c1269329b9a6bfcd8ee

Observation ec84c30b-501a-48fd-bf88-8f4599bdfc9a · outbound

This paper cites Penetration Testing for System Security: Methods and Practical Approaches,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Penetration Testing for System Security: Methods and Practical Approaches,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.340115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.340115Z digest=sha256:4d21018cb7c995062ff61c3e6b7564d52079f3e627e414bb4df4120f9148c1ba

Observation 87472f62-df8d-44c9-bfb8-ab1b9a1ae142 · outbound

This paper cites SmartBugBert: BERT-Enhanced Vulnerability Detection for Smart Contract Bytecode,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots SmartBugBert: BERT-Enhanced Vulnerability Detection for Smart Contract Bytecode,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.423115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.423115Z digest=sha256:3eb0fc1426ce94688a06a3c8312287cab76d197a94121444492e22b88db48c6b

Observation 6902c036-1c7a-48f7-8ff1-8517ddd6e91a · outbound

This paper cites Exploring Vulnerabilities and Concerns in Solana Smart Contracts.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Exploring Vulnerabilities and Concerns in Solana Smart Contracts

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.429733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.429733Z digest=sha256:45e77e846c2cd3ac9d62b2caef58d28c9663b157d6af17e7b6f25acd365e0ad8

Observation 90fae803-4798-4927-b9f4-6fc83e2ef225 · outbound

This paper cites Interaction-aware vulnerability detection in smart contract bytecodes,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Interaction-aware vulnerability detection in smart contract bytecodes,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.498143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.498143Z digest=sha256:6f67dd4cadec6d67dbf824eefbebb0c4bcf1f54a1f16d61d150dece7e2d93d51

Observation 2bb1dd54-0ee0-4931-ba20-987e97299127 · outbound

This paper cites Penetrating the hostile: Detecting DeFi protocol exploits through cross-contract analysis,.

An Intelligent-Cloud Edge Multimodal Interaction System for Robots Penetrating the hostile: Detecting DeFi protocol exploits through cross-contract analysis,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T01:28:01.619301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:28:01.619301Z digest=sha256:ee0d0ee89660dc99379507ff3084153a06a96ef6f808292906d6e2c37a6b5a1b

Pith citing papers

No inbound Pith citation observations are available.