Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T05:26:04.960844Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 100 of 208 outbound references and 100 inbound Pith citation observations for arXiv:2505.07062.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T05:26:04.960844Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T21:37:07.589391Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T16:15:06.198778Z
100 of 208 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation afcfd011-6c93-41fd-a90d-904ea8bc7c11 · outbound
Seed1.5-VL Technical Report Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2fac0f76-f434-4805-9089-0cddead8d2fa · outbound
Seed1.5-VL Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e43bdde6-fc62-4d79-83a5-a18769811fe9 · outbound
Seed1.5-VL Technical Report Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a1c198b8-abb0-44c3-b207-214fe00f3088 · outbound
Seed1.5-VL Technical Report Understanding Alignment in Multimodal LLMs: A Comprehensive Study
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1bf4ed3-812b-4a0f-b4e2-a11f97f052f0 · outbound
Seed1.5-VL Technical Report Claude 3.7 sonnet system card
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5d7fb1ab-4250-4767-bd01-d397437cc2cb · outbound
Seed1.5-VL Technical Report Claude’s extended thinking
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d46de2ad-e269-45ff-a58b-31e1eeb1baac · outbound
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1ca5c67f-9af3-438f-8585-b5e9c1637ef1 · outbound
Seed1.5-VL Technical Report Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ae0a26b0-86c9-44c1-bcf6-1e8a0a4070b9 · outbound
Seed1.5-VL Technical Report ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d0cb9a49-493f-49f8-a2a3-5ea434b3113c · outbound
Seed1.5-VL Technical Report PaliGemma: A versatile 3B VLM for transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e7ed0343-1903-4b5c-b731-67e105619faf · outbound
Seed1.5-VL Technical Report Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · outbound
Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b78cbcf5-abcd-4ff4-88d0-d514868036a1 · outbound
Seed1.5-VL Technical Report FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ee4d3bb-01e7-41cf-86b4-87ff10d51410 · outbound
Seed1.5-VL Technical Report MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2caeec6f-6c80-40cc-9668-6d42f76dcdf6 · outbound
Seed1.5-VL Technical Report Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6dc2d73c-eed9-495b-9d1f-6eabae4a22a8 · outbound
Seed1.5-VL Technical Report Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 52f4aac0-f2bf-4fc2-bb29-f3dc51fa2321 · outbound
Seed1.5-VL Technical Report Yolo-world: Real-time open- vocabulary object detection
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 44c2e2ee-1414-4644-a691-54fb50e9a9d3 · outbound
Seed1.5-VL Technical Report PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 769e071d-a57c-4a42-a863-a23d22964c86 · outbound
Seed1.5-VL Technical Report Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e299bb67-faa1-4f78-a3cf-f2fe4ceb9ec0 · outbound
Seed1.5-VL Technical Report Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e4e88b71-6c17-4e55-9f22-d64811bd6c3d · outbound
Seed1.5-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 68302095-fe8b-43b9-aa82-07fa136a25dd · outbound
Seed1.5-VL Technical Report Imagenet: A large-scale hierarchical image database
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · outbound
Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c32d8719-ac79-4347-9124-6935b7d63034 · outbound
Seed1.5-VL Technical Report Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 033ed50e-56d8-41f9-89f3-a6252a2938af · outbound
Seed1.5-VL Technical Report An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0744f927-0bff-469a-bd64-35c6682fd51c · outbound
Seed1.5-VL Technical Report Counting out time: Class agnostic video repetition counting in the wild
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4c5f184d-9b44-4e23-a71a-19977cb6d7c2 · outbound
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4a417498-d993-4e04-8995-825af5143ba1 · outbound
Seed1.5-VL Technical Report Eva: Exploring the limits of masked visual representation learning at scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bfdebabc-c4ea-45fd-ad31-39e70d3dfba5 · outbound
Seed1.5-VL Technical Report Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1ef4e04e-0b5b-4f35-b2db-c16925e41dce · outbound
Seed1.5-VL Technical Report Helix: A vision-language-action model for generalist humanoid control.https://www.figure.ai/ news/helix
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 53ca0743-daac-4203-b767-93aff87a3a4a · outbound
Seed1.5-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6516d9d2-cdca-414d-a6cc-52f0cb4445ce · outbound
Seed1.5-VL Technical Report Blink: Multimodal large language models can see but not perceive
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a70836d8-b460-4680-a901-79595d8f3d12 · outbound
Seed1.5-VL Technical Report Tall: Temporal activity localization via language query
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 05887b6e-1f85-4531-b72f-50449df61832 · outbound
Seed1.5-VL Technical Report Experiment with gemini 2.0 flash native image generation.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc8c93b0-3905-4e7a-85c6-46a92376c5ed · outbound
Seed1.5-VL Technical Report Saliency-guided detr for moment retrieval and highlight detection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 30db478d-cbaf-4879-aa3b-5b22634ce670 · outbound
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 49f84c85-9b74-4a42-87e5-1e7d4ceceb8a · outbound
Seed1.5-VL Technical Report Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · outbound
Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 203ec044-cc71-4792-877f-e90610d595e9 · outbound
Seed1.5-VL Technical Report Lvis: A dataset for large vocabulary instance segmentation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d5ebee7c-323c-46ff-a9db-0e45abb7a7d2 · outbound
Seed1.5-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 37746a3e-98f4-4307-b361-d4434033f04f · outbound
Seed1.5-VL Technical Report WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1b50b982-6127-4f1c-a8e5-3f58838874f7 · outbound
Seed1.5-VL Technical Report The many faces of robustness: A critical analysis of out-of-distribution generalization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ff96f097-bebd-4349-a2aa-1db2fa5e4b6b · outbound
Seed1.5-VL Technical Report Natural adversarial examples
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0be140a3-ab89-4a5b-b008-ddf7dce85e9c · outbound
Seed1.5-VL Technical Report Scaling Laws for Autoregressive Generative Modeling
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6b7e3d2b-dca4-42e4-88c0-aad33465ab40 · outbound
Seed1.5-VL Technical Report Training Compute-Optimal Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 48407459-62b0-4714-a90d-1bdd31dc1aaa · outbound
Seed1.5-VL Technical Report The Curious Case of Neural Text Degeneration
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c389374c-5eab-4d3e-9c2a-c10924b44ed9 · outbound
Seed1.5-VL Technical Report MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8b90fa82-393e-438c-9a3a-e2da6583651c · outbound
Seed1.5-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d872c490-ca22-425b-928f-5311fd07972b · outbound
Seed1.5-VL Technical Report Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f1005a7c-390d-42be-bba1-17b0d6d91301 · outbound
Seed1.5-VL Technical Report Online Video Understanding: OVBench and VideoChat-Online
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6182c567-ca84-41f4-a445-04c7b292a1db · outbound
Seed1.5-VL Technical Report Classification done right for vision-language pre-training
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 602425a6-015e-4768-9605-dabf9c43e2e4 · outbound
Seed1.5-VL Technical Report D., 2007, @doi [Computing in Science and Engineering] 10.1109/MCSE.2007.55 , 9, 90
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4b45629c-a7b9-4085-a3e9-431dd431efec · outbound
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b67af7ca-89d7-4017-8dcd-a9aed805bb63 · outbound
Seed1.5-VL Technical Report $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b05d473d-0521-41c9-a644-524f69447186 · outbound
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 309dea63-518c-4c51-af93-3783303ff936 · outbound
Seed1.5-VL Technical Report In 21st USENIX Symposium on NetworkedSystems Design and Implementation (NSDI 24), pages 745–760
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3f3c3296-3673-4a32-a746-7c6ae81b6f56 · outbound
Seed1.5-VL Technical Report FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7f19c731-9841-4029-9af4-962cfd0fea24 · outbound
Seed1.5-VL Technical Report Scaling Laws for Neural Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ff76b86-eaa7-46f9-9c10-6a321a390258 · outbound
Seed1.5-VL Technical Report Referitgame: Referring to objects in photographs of natural scenes
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2ba3d0f2-ebc1-4318-91ba-aa660858ecc6 · outbound
Seed1.5-VL Technical Report A diagram is worth a dozen images
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 890d3d42-54dc-4fa6-8e10-aebe70c29ca4 · outbound
Seed1.5-VL Technical Report Ocr-free document understanding transformer
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1c6a1699-9822-4bb0-adc9-e031f6a7bbfe · outbound
Seed1.5-VL Technical Report Openvla: An open-source vision-language-action model
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 12256451-28ef-4dea-9daa-58621d2d0fbd · outbound
Seed1.5-VL Technical Report Adam: A Method for Stochastic Optimization
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 87a025c5-d2a8-4348-9836-ddf276b9acf2 · outbound
Seed1.5-VL Technical Report Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c012fa6c-48c5-4d5d-b938-ebd8d6ed12d8 · outbound
Seed1.5-VL Technical Report The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 02d3c154-8b2e-4bc1-b46c-4f627f3fe759 · outbound
Seed1.5-VL Technical Report Gonzalez, Hao Zhang, and Ion Stoica
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd045068-0ad5-4337-b55a-975fd7ff9702 · outbound
Seed1.5-VL Technical Report Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fe7e3a20-3c4b-4e94-8988-52459e76e5dc · outbound
Seed1.5-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7cd21718-a876-4dc6-abd9-b77cb7dd2a94 · outbound
Seed1.5-VL Technical Report Echarts: a declarative framework for rapid construction of web-based visualization.Visual Informatics, 2(2):136–146
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e66a46c0-3c91-4517-ac37-8e63025eff47 · outbound
Seed1.5-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 33b9e398-5f45-423e-92dd-6c3272960c36 · outbound
Seed1.5-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 675eeab5-8790-4c16-bf3d-5c39d5151d88 · outbound
Seed1.5-VL Technical Report Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6cccf783-ee74-4d1e-b9d0-eac6fdb1ab84 · outbound
Seed1.5-VL Technical Report OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6d210c9c-db51-4ece-b83e-e8e59d775169 · outbound
Seed1.5-VL Technical Report The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a7739048-e734-4807-b2d6-c769a53a1e94 · outbound
Seed1.5-VL Technical Report StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c963f759-fa33-4d40-926c-2d7b7c60b7df · outbound
Seed1.5-VL Technical Report Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e125210c-4fab-4ab3-9c5a-50a03211a77f · outbound
Seed1.5-VL Technical Report Visual instruction tuning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 45680f7d-05e7-4bf0-8c52-88abf138915e · outbound
Seed1.5-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 40f7a881-bd38-4256-b550-fd9872191a45 · outbound
Seed1.5-VL Technical Report Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ddb64011-8cbc-4818-b680-416ee764cb66 · outbound
Seed1.5-VL Technical Report Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation dee4ab59-fd2c-489e-8f14-73b3a8735974 · outbound
Seed1.5-VL Technical Report TempCompass: Do Video LLMs Really Understand Videos?
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7cd549ae-5d9d-4e8e-af05-5d325c973b62 · outbound
Seed1.5-VL Technical Report Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a34b8a8b-4669-4658-85b2-d5557e30d5d1 · outbound
Seed1.5-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5b7c0424-0ad2-48dc-ae2b-ae49ede7034b · outbound
Seed1.5-VL Technical Report Ursa: Under- standing and verifying chain-of-thought reasoning in multi- modal mathematics
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4d6b2d37-f2b3-4221-b7d1-cfc6bc8689b0 · outbound
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fdb0da2f-6ebf-452b-a2f8-a44d1a594cf2 · outbound
Seed1.5-VL Technical Report Egoschema: A diagnostic benchmark for very long-form video language understanding.Advancesin Neural Information Processing Systems, 36:46212–46244
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ddd73e2-591e-4fce-99f0-8b25006ba738 · outbound
Seed1.5-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 89a6b06a-1091-4261-a221-92a35397ba27 · outbound
Seed1.5-VL Technical Report Docvqa: A dataset for vqa on document images
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ba0a29d2-3a8c-4cfd-886f-cf1f7e7539d8 · outbound
Seed1.5-VL Technical Report Info- graphicvqa
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fe0f7ed6-1c69-4b57-8df3-d082f58edab1 · outbound
Seed1.5-VL Technical Report The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a476d431-18ca-4b33-b46d-eb513dae0cf3 · outbound
Seed1.5-VL Technical Report Modeling context between objects for referring expression understanding
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 127ec65b-80c7-4015-9ea1-2867c375c79c · outbound
Seed1.5-VL Technical Report Memory-efficient pipeline- parallel dnn training
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4b8b888f-1bae-438a-8669-c29a28dd0630 · outbound
Seed1.5-VL Technical Report Efficient large-scale language model training on gpu clusters using megatron-lm
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59e7d375-3180-4bdb-a701-d888ef651391 · outbound
Seed1.5-VL Technical Report Indoor segmentation and support inference from rgbd images
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9b4338fa-51ad-456e-b352-b7e501089b62 · outbound
Seed1.5-VL Technical Report Gpt-4v(ision) system card
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1d8d762f-4c4a-457e-8604-ae5f53943323 · outbound
Seed1.5-VL Technical Report Addendum to gpt-4o system card: 4o image generation
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b569b03-50fc-4213-9727-1736b58ec368 · outbound
Seed1.5-VL Technical Report Operator
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6f2c3e87-a354-4abf-bb94-4c044ab34a4e · outbound
Seed1.5-VL Technical Report DINOv2: Learning Robust Visual Features without Supervision
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cccf8905-622b-4a0b-8a49-6fe756b8475e · outbound
Seed1.5-VL Technical Report Training language models to follow instructions with human feedback
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f4f514b0-5bae-4d7b-bb64-ab29ea276fb6 · outbound
Seed1.5-VL Technical Report How predictable is language model benchmark performance?
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b01406b5-88b8-4a52-989b-02fc4a1d11d9 · inbound
LVBench: An Extreme Long Video Understanding Benchmark Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e29e291a-9035-460e-a735-6373514e41b5 · inbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 31960e2f-eab4-4db4-a130-b6b293d6ee4e · inbound
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Seed1.5-VL Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 93cdf003-0ea1-4adb-ae41-92f3d9eff97f · inbound
GTA1: GUI Test-time Scaling Agent Seed1.5-VL Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a61bccae-b27c-4896-903c-de4f5ce04ff3 · inbound
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 00ceeeda-b9c8-4dc3-82ac-78f64201205f · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Seed1.5-VL Technical Report
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9f6d2274-fe55-41f4-8389-1e9fe22a5061 · inbound
Seedream 4.0: Toward Next-generation Multimodal Image Generation Seed1.5-VL Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4e95f10a-35a2-499a-9591-3c343fe030d8 · inbound
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Seed1.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0ccb065e-fc53-4243-8d35-1912f4ee181b · inbound
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c6b6c83a-6dca-4ec9-aed3-b060ce150775 · inbound
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents Seed1.5-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9f752b9d-7723-4de2-aa7b-efa6ad70be72 · inbound
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e80dff-649b-4887-be08-6a2e5abb1b77 · inbound
MiMo-Embodied: X-Embodied Foundation Model Technical Report Seed1.5-VL Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 33745277-755c-4e19-8869-9eae62425e40 · inbound
Boosting Reasoning in Large Multimodal Models via Activation Replay Seed1.5-VL Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 15010b91-8e2e-4448-afe6-bf85f8207e74 · inbound
OneThinker: All-in-one Reasoning Model for Image and Video Seed1.5-VL Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f9e5bdac-24fc-4247-8288-958bce8c4530 · inbound
Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning Seed1.5-VL Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d30e0b-b72c-432e-bfeb-24d3a70d46b4 · inbound
Grounding Everything in Tokens for Multimodal Large Language Models Seed1.5-VL Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ef4a5d2-12b8-40ac-9ecd-cd22d91fc50c · inbound
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning Seed1.5-VL Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 00b8b6ef-d0be-4677-88e4-d814d681dd06 · inbound
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Seed1.5-VL Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b153076-3b0f-4b79-9444-d69987996b3c · inbound
CountGD++: Generalized Prompting for Open-World Counting Seed1.5-VL Technical Report
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bed1ebcf-69f7-4e96-9e5a-9f763a375257 · inbound
LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data Seed1.5-VL Technical Report
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0952cbb7-12ab-4624-af42-810838bc210f · inbound
Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method Seed1.5-VL Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e15eb89-a3f8-48b1-aeec-ab53b8a4d1ae · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Seed1.5-VL Technical Report
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2ac7a027-5a9b-492d-9e3f-e67bf1daad3e · inbound
Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction Seed1.5-VL Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f142a9-63b0-4cb1-a101-7438cfd479a6 · inbound
Kimi K2.5: Visual Agentic Intelligence Seed1.5-VL Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c9d91b91-63b2-4212-ba47-7455d8f22b9d · inbound
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f597c6e8-5eef-49d7-b298-43cfc185f295 · inbound
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Seed1.5-VL Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d46b1046-30c1-4b3a-9694-0a43c78c9f30 · inbound
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Seed1.5-VL Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 70193c6f-cc14-436f-8edf-b32e7c284204 · inbound
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Seed1.5-VL Technical Report
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25d4cdf8-7156-44a0-8371-ae0638bea619 · inbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9158f500-ae05-4265-8f33-ca70f58be589 · inbound
PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335b71fd-9a91-4702-bd6d-6032358880bd · inbound
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f9eef8-ed59-4106-b25e-0304f4fb151c · inbound
CodePercept: Code-Grounded Visual STEM Perception for MLLMs Seed1.5-VL Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1cbe2ac-e98a-42db-8740-f94b9316b201 · inbound
AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison Seed1.5-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 705d5836-de75-4a73-a082-ccc5d2e8c2e7 · inbound
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Seed1.5-VL Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d7687b2-011d-40ca-84c1-4979940d7d04 · inbound
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a792e4-84d9-46c4-892a-1ca8771ec2b4 · inbound
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Seed1.5-VL Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684d642b-a880-4338-a93b-9e3415a180a9 · inbound
InstructTable: Improving Table Structure Recognition Through Instructions Seed1.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ca1359cd-6476-4fd8-96b4-b15426c2db73 · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9f7f9bfe-e4d0-46cf-86c9-40be3fad9562 · inbound
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51bd245e-7b26-474a-af68-f51e37231b67 · inbound
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Seed1.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4729d643-ea08-4e50-874d-f574b59a4bc1 · inbound
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence Seed1.5-VL Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c15ddaba-36c1-4d49-9990-edd8358fce2e · inbound
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Seed1.5-VL Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1313a7be-ad69-4774-a5ac-00d1a6eb1146 · inbound
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Seed1.5-VL Technical Report
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f6083b10-d61c-459a-a5fa-761cc5435c5b · inbound
LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation Seed1.5-VL Technical Report
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac4bdaa8-d253-4395-9cd6-0c806ee9d2ce · inbound
EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9911c8e4-65b5-4c44-888e-b94494bbab64 · inbound
EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8eea66-cc36-4ae1-9e8e-f1e993ef632a · inbound
Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation Seed1.5-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c343d3b9-ca6a-4f4e-acf0-5081ed2d6f6e · inbound
CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping Seed1.5-VL Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bb5762b0-f8a7-4249-8fdd-0af5a468e403 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Seed1.5-VL Technical Report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cd114e61-0326-430a-b985-10dfaf86ebbb · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Seed1.5-VL Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1ad42240-2d80-4bc4-919a-06069a60de67 · inbound
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8b8c1bf1-2dd6-48d3-b435-21b9c24ecb7c · inbound
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99aedbb3-c5f6-44aa-9dfd-c2bc4e0c655c · inbound
UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding Seed1.5-VL Technical Report
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 30c6fd44-077b-4ec3-a935-a855b3328712 · inbound
Seedance 2.0: Advancing Video Generation for World Complexity Seed1.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9fb25da2-9171-4cab-a7b5-87b2b715a864 · inbound
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow Seed1.5-VL Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b4feee8a-4cc1-4113-89e0-628590d048e3 · inbound
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior Seed1.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 07575c65-6c00-4c56-98c5-3146de1aebdb · inbound
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning Seed1.5-VL Technical Report
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5f3ec7c3-edf9-44e2-9199-03e7af082812 · inbound
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1298f9e2-02e3-458a-af07-8c4ef8c99a56 · inbound
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3e3c3bf8-286b-42b4-a6d5-be387d7f8822 · inbound
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7deaca6e-06fe-4c94-bbc3-8133b4d2ba07 · inbound
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Seed1.5-VL Technical Report
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 706b69ac-33c7-4e0e-86ba-b603988810b7 · inbound
dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a9231b2d-39e1-4c77-9296-b322c8ebba8f · inbound
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Seed1.5-VL Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 22df99a6-c8d7-4bcb-8dbc-02cb5f53f4ec · inbound
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection Seed1.5-VL Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7fb67bc8-972f-4fa8-a465-cab0cb1f9ec4 · inbound
Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7d8bdbba-e9e2-4c21-af96-42e6b75f2e53 · inbound
Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2a0858e7-282c-4077-b08b-cc03e9aa955e · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Seed1.5-VL Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b4d31d0f-9da0-4311-96b4-c06ac22dba24 · inbound
Leveraging Verifier-Based Reinforcement Learning in Image Editing Seed1.5-VL Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ddb424df-4a1e-4320-9cb2-b5338fa84e2b · inbound
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants Seed1.5-VL Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c1618e6e-388a-433d-a17f-8c03e4b1e196 · inbound
AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding Seed1.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3853f9ff-e339-4bf5-a7cd-9cf5f15b5235 · inbound
Perceptual Flow Network for Visually Grounded Reasoning Seed1.5-VL Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 92b39476-037d-4d4a-8793-1011ece6264f · inbound
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e566e246-7453-46c6-8e9c-2239f39a2067 · inbound
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Seed1.5-VL Technical Report
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e64c2ca0-8620-4994-9fab-e199fdd2e8b6 · inbound
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7b5eb592-f1b0-4888-8e83-cb3e03d2d195 · inbound
Text-Guided Multi-Scale Frequency Representation Adaptation Seed1.5-VL Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e1ae2a60-51fa-4193-aa76-45e2944494ce · inbound
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a47c2dd-a9f8-47d2-8d14-1c8c02ccb7e7 · inbound
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8857fdb6-37ee-4f42-a42b-422fd953fb38 · inbound
Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9399df48-cdc2-41cf-aa76-70e174d1d9c0 · inbound
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Seed1.5-VL Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8307dabb-8164-4771-8412-811d9cd009d1 · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ba3a781-2891-43c1-995c-1ce13eceffac · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3d997b2c-a990-49d6-82c8-960b564d1e8b · inbound
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs Seed1.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 95d74c2c-fd9b-462f-891e-7549a93c223b · inbound
Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning Seed1.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f568a41a-3ffc-451d-92ca-f0c9eb2cf10d · inbound
How Mobile World Model Guides GUI Agents? Seed1.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 590f1c15-5a8a-4c90-b186-a1a93ad60818 · inbound
How Mobile World Model Guides GUI Agents? Seed1.5-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 77c60c3f-f062-406c-9f78-4abe05f0d8f9 · inbound
AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation Seed1.5-VL Technical Report
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a915e8f3-18a6-447e-9509-ea2dccc0f5aa · inbound
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seed1.5-VL Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 89c1cc12-55ee-4f13-9398-29448fa0a91d · inbound
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Seed1.5-VL Technical Report
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4d94be9a-8397-4b60-bc2d-952613344424 · inbound
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Seed1.5-VL Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6ab74e-6ef3-42eb-81a2-fcff4d039c5a · inbound
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Seed1.5-VL Technical Report
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 327c5b0a-0aab-443c-8c45-d424af987953 · inbound
ViMU: Benchmarking Video Metaphorical Understanding Seed1.5-VL Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation af230fbb-e5b3-4c96-bd0c-6cf9a9e2e689 · inbound
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Seed1.5-VL Technical Report
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 958e1bf9-61e2-4070-a824-c181ecec9f4e · inbound
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs Seed1.5-VL Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d5db45d8-d9cf-4209-a957-23ef81ba112c · inbound
SEED: Targeted Data Selection by Weighted Independent Set Seed1.5-VL Technical Report
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac9dabdc-9edc-420d-ba36-614fb1a950a6 · inbound
Unlocking Dense Metric Depth Estimation in VLMs Seed1.5-VL Technical Report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 77e0a633-260b-45ba-9df4-a5dfe759a236 · inbound
Unlocking Dense Metric Depth Estimation in VLMs Seed1.5-VL Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ec6f2d9-b656-4687-ba11-211bc949893f · inbound
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making Seed1.5-VL Technical Report
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 717659f3-0fcd-44f0-99e8-71690ae8bfa0 · inbound
FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation Seed1.5-VL Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6006882c-cdfb-41bf-a927-eae5cae77f02 · inbound
Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors Seed1.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation aff8dc3e-e3cb-4b7c-89c4-9e2da8a5af59 · inbound
RISE: Reliable Improvement in Self-Evolving Vision-Language Models Seed1.5-VL Technical Report
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.