Pith. sign in

Paper Citation Record · LEDGER

Seed1.5-VL Technical Report

As of 3 August 2026, this Paper Citation Record lists 100 of 208 outbound references and 100 inbound Pith citation observations for arXiv:2505.07062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07062 v1

Coverage vector

measured 100 of 208 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T05:26:04.960844Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 100 of 156 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:37:07.589391Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T16:15:06.198778Z

Reference resolution

100 of 208 outbound references displayed

  • verified exact51
  • verified fuzzy45
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afcfd011-6c93-41fd-a90d-904ea8bc7c11 · outbound

This paper cites an unresolved cited work.

Seed1.5-VL Technical Report Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:26:06.278271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4ca044aa930980db6dc9203ed597cf6f15a02f083c88eb2bf1b4dba906958056

Observation 2fac0f76-f434-4805-9089-0cddead8d2fa · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Seed1.5-VL Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.205044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:94553ef84a9bf2434ed0bf7a9c850c8a4e03e92db5e37aad260b5fcd6bf3246e

Observation e43bdde6-fc62-4d79-83a5-a18769811fe9 · outbound

This paper cites Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837.

Seed1.5-VL Technical Report Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.288023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:47359e19052ef98082753e3f49abf4f5a8514ef8d9328da9c678578a9879900d

Observation a1c198b8-abb0-44c3-b207-214fe00f3088 · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Seed1.5-VL Technical Report Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.032367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1070a03c8e7fcbfc58578303d4949292af9aa349b50fe6ebe67db111a25ff08e

Observation f1bf4ed3-812b-4a0f-b4e2-a11f97f052f0 · outbound

This paper cites Claude 3.7 sonnet system card.

Seed1.5-VL Technical Report Claude 3.7 sonnet system card

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.294589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:66758bc24e5d3f6606f0b86c5bb9bd599ddcd207d41963c9945cce1af5502ae8

Observation 5d7fb1ab-4250-4767-bd01-d397437cc2cb · outbound

This paper cites Claude’s extended thinking.

Seed1.5-VL Technical Report Claude’s extended thinking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.299128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:20207a6ce63588ef6a3d695b5b040f52c96e4bf671e97e34973758a1799fd1fd

Observation d46de2ad-e269-45ff-a58b-31e1eeb1baac · outbound

This paper cites Qwen2.5-VL Technical Report.

Seed1.5-VL Technical Report Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.047883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9659124a14d05cd813a26a87f45106ebabb6a0aab028d27cde309a4fd71c838f

Observation 1ca5c67f-9af3-438f-8585-b5e9c1637ef1 · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

Seed1.5-VL Technical Report Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.303629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:abec832e97661484d5b76d836fedc41e3f9e76ccc5cb5c3bf30bac782d447cb8

Observation ae0a26b0-86c9-44c1-bcf6-1e8a0a4070b9 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Seed1.5-VL Technical Report ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:47:09.105767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a1eaa69550ed5867ec3d4db30f58cf2529beb9fd8ed61002769da4e28fe5b4ee

Observation d0cb9a49-493f-49f8-a2a3-5ea434b3113c · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Seed1.5-VL Technical Report PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:58940d5d6e377b45f1e64f26150fb0ba48ba8ada201aa57e768abe649e2b14be

Observation e7ed0343-1903-4b5c-b731-67e105619faf · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

Seed1.5-VL Technical Report Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.586427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6f86164a99eab4d2b7be39b9d483d2c92554f6df04df32c6acf9d972a51e2b84

Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.681583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:370433835bd5ba040bcf7ba39fa60ca027d8c8ee2fe9f35ce15612b273513276

Observation b78cbcf5-abcd-4ff4-88d0-d514868036a1 · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

Seed1.5-VL Technical Report FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.905725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5e78c6b760c7dec42861ab4ddbb474512e81cce1173f282706459be4e8723778

Observation 7ee4d3bb-01e7-41cf-86b4-87ff10d51410 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Seed1.5-VL Technical Report MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.955218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6b5b66d801e012f10aa728f162dbe05d1780e79186e7318910fbed36a8a14512

Observation 2caeec6f-6c80-40cc-9668-6d42f76dcdf6 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Seed1.5-VL Technical Report Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:033fcb423c31b0f13dc5b326289fba00cb4f952fa7a62911ab9eac456eea25fc

Observation 6dc2d73c-eed9-495b-9d1f-6eabae4a22a8 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Seed1.5-VL Technical Report Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.311222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c99b2fd2543ac3919ece4e2108415dc8e8fd2ffe97ed8475f3bfd4ab2a1893ec

Observation 52f4aac0-f2bf-4fc2-bb29-f3dc51fa2321 · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

Seed1.5-VL Technical Report Yolo-world: Real-time open- vocabulary object detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.318911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6469011f622a8830e7a9fec10f94f63bd4e3b0eaed8f31eea9d94be579de508c

Observation 44c2e2ee-1414-4644-a691-54fb50e9a9d3 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Seed1.5-VL Technical Report PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.081961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:32c64d267dc97b412f7135b2053463ef0e1be41ae4130032601fdd62a4663a36

Observation 769e071d-a57c-4a42-a863-a23d22964c86 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Seed1.5-VL Technical Report Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:06.098537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2b03880b42cbf5d7eb52b66dad5640e3386c7d8e90e2c7b587b672cdd852db75

Observation e299bb67-faa1-4f78-a3cf-f2fe4ceb9ec0 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Seed1.5-VL Technical Report Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.323904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:885de0863d12c0884bd813b65472eb6ee7b00554fa02c898bb704e511795dab7

Observation e4e88b71-6c17-4e55-9f22-d64811bd6c3d · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Seed1.5-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:dbf508d7dc587754751cc8b7cd3fa4437d00d220eda96ea98741041abaee9d98

Observation 68302095-fe8b-43b9-aa82-07fa136a25dd · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Seed1.5-VL Technical Report Imagenet: A large-scale hierarchical image database

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.328500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b7ddbd85dc772f4b7db868afc8d27af640444c0c5e0ca66a7f0b71c30f05a539

Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a4ad942d3202c9efc7e4a0582652de897cadc4e5daa4dbcfb8b5498950e6e032

Observation c32d8719-ac79-4347-9124-6935b7d63034 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Seed1.5-VL Technical Report Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.485132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f6a8417255df87a9988cc2cce04d85407c226d918d4c7157f5a0089a76c27428

Observation 033ed50e-56d8-41f9-89f3-a6252a2938af · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Seed1.5-VL Technical Report An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.491357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c9e287d41ca7c1e73ae180529cb773bacbacbd46089237e069174ba450a3612a

Observation 0744f927-0bff-469a-bd64-35c6682fd51c · outbound

This paper cites Counting out time: Class agnostic video repetition counting in the wild.

Seed1.5-VL Technical Report Counting out time: Class agnostic video repetition counting in the wild

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.336018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:69ae095a8bf02327cb1a074f1f9196455ec8c1d56e16737f989e2ec39f20007f

Observation 4c5f184d-9b44-4e23-a71a-19977cb6d7c2 · outbound

This paper cites Data Filtering Networks.

Seed1.5-VL Technical Report Data Filtering Networks

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.559184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f94bce07743976cfcc87e1555135296c915ea78889a9f1ff01a969afe53dbab6

Observation 4a417498-d993-4e04-8995-825af5143ba1 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Seed1.5-VL Technical Report Eva: Exploring the limits of masked visual representation learning at scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.341778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a338f236182e5ebfd5371e3f73ecfee28406d2da7dd7890cfaba7495bc90d1dc

Observation bfdebabc-c4ea-45fd-ad31-39e70d3dfba5 · outbound

This paper cites Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation.

Seed1.5-VL Technical Report Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.608208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1b9cefa12a7e28602d13a75a7db03a1d86fdf62f8177f42177eabb4d0a15d663

Observation 1ef4e04e-0b5b-4f35-b2db-c16925e41dce · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control.https://www.figure.ai/ news/helix.

Seed1.5-VL Technical Report Helix: A vision-language-action model for generalist humanoid control.https://www.figure.ai/ news/helix

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.348494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c0bb6425f91ba6fbae4fc7821f11ee4549a53b8702f0b9d65cf7dca2f97e6275

Observation 53ca0743-daac-4203-b767-93aff87a3a4a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Seed1.5-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.774438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:38aa19cdc09d5da4c72ad33029496ba853d0e1b5c06fd3ec8d8304985d7ba092

Observation 6516d9d2-cdca-414d-a6cc-52f0cb4445ce · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Seed1.5-VL Technical Report Blink: Multimodal large language models can see but not perceive

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.353238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7bd4d9d8a79c640587d8de0988ae47e838dd2a84e164ca5005b44c991e798d90

Observation a70836d8-b460-4680-a901-79595d8f3d12 · outbound

This paper cites Tall: Temporal activity localization via language query.

Seed1.5-VL Technical Report Tall: Temporal activity localization via language query

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.357984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b912d380273f4e7a07079a3e8f3e5f19eff668d530ed6154766881c73f429f62

Observation 05887b6e-1f85-4531-b72f-50449df61832 · outbound

This paper cites Experiment with gemini 2.0 flash native image generation.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation.

Seed1.5-VL Technical Report Experiment with gemini 2.0 flash native image generation.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.367407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:8029c5e41a45ff144f3bec6ed8bfabe5da9cbb363b7ac4f69546027f3cdcd4b4

Observation fc8c93b0-3905-4e7a-85c6-46a92376c5ed · outbound

This paper cites Saliency-guided detr for moment retrieval and highlight detection.

Seed1.5-VL Technical Report Saliency-guided detr for moment retrieval and highlight detection

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.981905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f00ee868bcf89b89113be3cc9f9ca1d2ec0c6fbaf7c76fb48017f5816355baf6

Observation 30db478d-cbaf-4879-aa3b-5b22634ce670 · outbound

This paper cites The Llama 3 Herd of Models.

Seed1.5-VL Technical Report The Llama 3 Herd of Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.015925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:00a164c5f6c38b194d5b72b13c53f06f97fbe0e00414b921ed88bba702dad1bb

Observation 49f84c85-9b74-4a42-87e5-1e7d4ceceb8a · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Seed1.5-VL Technical Report Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.372253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b8b0d8b4367dbf794058528c56da3de070b6da5ddf7a777452d8962efc1be47e

Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.042133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2a5804396296d5a8587ecd66aaf439a327fb112b8194d6982586757ab4bd6b53

Observation 203ec044-cc71-4792-877f-e90610d595e9 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Seed1.5-VL Technical Report Lvis: A dataset for large vocabulary instance segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.376984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9187bbfb6061c669230f622da31a5d173deca9de5f5093a22c89a871dcceefcf

Observation d5ebee7c-323c-46ff-a9db-0e45abb7a7d2 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Seed1.5-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:843c71e7bfc343a15173653577db0124da5d7eb7eaded410f724bb8431f9324d

Observation 37746a3e-98f4-4307-b361-d4434033f04f · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Seed1.5-VL Technical Report WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:43:35.479641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e969285682783b13ee4c3aac03d8a3c360f1e2c43166a5d7add27bf4e968f3a2

Observation 1b50b982-6127-4f1c-a8e5-3f58838874f7 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Seed1.5-VL Technical Report The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.381706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b40ae659d423f6930b15aea7f0dbd432122fe96b3e3c3a13cc8e88da6f1a2732

Observation ff96f097-bebd-4349-a2aa-1db2fa5e4b6b · outbound

This paper cites Natural adversarial examples.

Seed1.5-VL Technical Report Natural adversarial examples

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.391576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e4373b071d8a0e2cea68151fec8ebde410d36b59bbea07b8d2c376743ae83ece

Observation 0be140a3-ab89-4a5b-b008-ddf7dce85e9c · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Seed1.5-VL Technical Report Scaling Laws for Autoregressive Generative Modeling

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:49:43.923372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9493027f76d2437e187eed116007c2dde76c4834c7dca3f6cef4c961d6ee41d9

Observation 6b7e3d2b-dca4-42e4-88c0-aad33465ab40 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Seed1.5-VL Technical Report Training Compute-Optimal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.128471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5f2f142cabeb6ec84a08ef72f16a139fe09b8522f4a8dc3f2fbcae4384168aed

Observation 48407459-62b0-4714-a90d-1bdd31dc1aaa · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Seed1.5-VL Technical Report The Curious Case of Neural Text Degeneration

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:18:24.042352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:77e3f16bdd79f521bd60465fb104fb4aaffd5a82812733a659c82c05086908dc

Observation c389374c-5eab-4d3e-9c2a-c10924b44ed9 · outbound

This paper cites MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models.

Seed1.5-VL Technical Report MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:48:11.661227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c75556182e3c39fddcd313302edf4ffc1d1f30dc8b00cc558ef8bd616cdee6f0

Observation 8b90fa82-393e-438c-9a3a-e2da6583651c · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Seed1.5-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4c6956584db42c03dc70afb527f7b236cfcd456632f9a59cc05ff642b9b73ac0

Observation d872c490-ca22-425b-928f-5311fd07972b · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Seed1.5-VL Technical Report Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.396470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:71f1f57436a338875c04b49e4c27133e9d085e1b5f98faec48d5c239967ba56d

Observation f1005a7c-390d-42be-bba1-17b0d6d91301 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

Seed1.5-VL Technical Report Online Video Understanding: OVBench and VideoChat-Online

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.215670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c2373d7e2522c43054961cea55f8ddcb1ac7ed6e8118871f55d8e5995d20eec6

Observation 6182c567-ca84-41f4-a445-04c7b292a1db · outbound

This paper cites Classification done right for vision-language pre-training.

Seed1.5-VL Technical Report Classification done right for vision-language pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.400636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:636917ff6c1d0976e144d07ea6f0717c8d5402e2a77a881aafe72a582f48ef1d

Observation 602425a6-015e-4768-9605-dabf9c43e2e4 · outbound

This paper cites D., 2007, @doi [Computing in Science and Engineering] 10.1109/MCSE.2007.55 , 9, 90.

Seed1.5-VL Technical Report D., 2007, @doi [Computing in Science and Engineering] 10.1109/MCSE.2007.55 , 9, 90

Reference 53

Resolution
verified exact
doi, observed 2026-05-11T05:26:05.347076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f5ef3d0895220d4e6520e0cf0ad03e09d23d13b13ee7cb48ebaa63452ecb6098

Observation 4b45629c-a7b9-4085-a3e9-431dd431efec · outbound

This paper cites GPT-4o System Card.

Seed1.5-VL Technical Report GPT-4o System Card

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.379349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4f95dd7dcbc2ba72f1bc9dfb5fd6d5c0d8fc5fc723a97407c8865537a19304cc

Observation b67af7ca-89d7-4017-8dcd-a9aed805bb63 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Seed1.5-VL Technical Report $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.406172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:cad8cf5c85fcce18fe222e8c6b61f7adbf1a882dec9d2fa60dffc6877bcdf6a4

Observation b05d473d-0521-41c9-a644-524f69447186 · outbound

This paper cites OpenAI o1 System Card.

Seed1.5-VL Technical Report OpenAI o1 System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.415512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0000accc906e7b60207763661bbbe2eca9e764b67a42374eb67a1cd1777146e1

Observation 309dea63-518c-4c51-af93-3783303ff936 · outbound

This paper cites In 21st USENIX Symposium on NetworkedSystems Design and Implementation (NSDI 24), pages 745–760.

Seed1.5-VL Technical Report In 21st USENIX Symposium on NetworkedSystems Design and Implementation (NSDI 24), pages 745–760

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.404958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:64cca92606a02a02e0634f62539c1480aed66f7ce5891c87142303b79363872b

Observation 3f3c3296-3673-4a32-a746-7c6ae81b6f56 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Seed1.5-VL Technical Report FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.448721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f9f113677500f05915c31a0ce06e4835bbf6a32ce184f3df96a3022328b958f1

Observation 7f19c731-9841-4029-9af4-962cfd0fea24 · outbound

This paper cites Scaling Laws for Neural Language Models.

Seed1.5-VL Technical Report Scaling Laws for Neural Language Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.456643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:458e3de6d3f5f52e1f55a65ea5384da3d47f66d398a415a974b1270e508539df

Observation 8ff76b86-eaa7-46f9-9c10-6a321a390258 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Seed1.5-VL Technical Report Referitgame: Referring to objects in photographs of natural scenes

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.415359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:19462c49ecefacb4960dc2175ee1e9e6d6e1888c1d3118a3b6ece377a5b07bdd

Observation 2ba3d0f2-ebc1-4318-91ba-aa660858ecc6 · outbound

This paper cites A diagram is worth a dozen images.

Seed1.5-VL Technical Report A diagram is worth a dozen images

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.423351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a560de1ea2257fac3f8e9f51135eb3c566d55988ee2db311c88e519edd4d50b5

Observation 890d3d42-54dc-4fa6-8e10-aebe70c29ca4 · outbound

This paper cites Ocr-free document understanding transformer.

Seed1.5-VL Technical Report Ocr-free document understanding transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.429262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:22a11080f66319a9cfb23241bedc4499e3f4632797df856af53d9f19a7d339da

Observation 1c6a1699-9822-4bb0-adc9-e031f6a7bbfe · outbound

This paper cites Openvla: An open-source vision-language-action model.

Seed1.5-VL Technical Report Openvla: An open-source vision-language-action model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.439687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c3babe644e3ab814ce9d6705c5bd43433f4857a815dd8134deca7d4bac1c69ed

Observation 12256451-28ef-4dea-9daa-58621d2d0fbd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Seed1.5-VL Technical Report Adam: A Method for Stochastic Optimization

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.552726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3790084e2dd6f38a10ff50df89aea8fd9c67ddd386e91ce175d866710146ae04

Observation 87a025c5-d2a8-4348-9836-ddf276b9acf2 · outbound

This paper cites Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353.

Seed1.5-VL Technical Report Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.445032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a2e930d903186f21c8d1ae20bb745182331310f3c554f89e2521f7d869e25a14

Observation c012fa6c-48c5-4d5d-b938-ebd8d6ed12d8 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV.

Seed1.5-VL Technical Report The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.451376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:98da1db17c64f0a6ff29cc3c0e32f36801e669f880087c6928da1447b073fffd

Observation 02d3c154-8b2e-4bc1-b46c-4f627f3fe759 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Seed1.5-VL Technical Report Gonzalez, Hao Zhang, and Ion Stoica

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.460726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d781826b9bb0df39322f840af9773895615c32b6cdd563db96a2167fb16f2ae2

Observation fd045068-0ad5-4337-b55a-975fd7ff9702 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Seed1.5-VL Technical Report Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0f54a2111b26a8a18e8bbe9e70bf27e2eb0ba869110a25942723252d1596a62e

Observation fe7e3a20-3c4b-4e94-8988-52459e76e5dc · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Seed1.5-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:26:05.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c21c47f5e903062b80fb2e536f41ecb9374dfe819d4db82efbc8c0749b76743d

Observation 7cd21718-a876-4dc6-abd9-b77cb7dd2a94 · outbound

This paper cites Echarts: a declarative framework for rapid construction of web-based visualization.Visual Informatics, 2(2):136–146.

Seed1.5-VL Technical Report Echarts: a declarative framework for rapid construction of web-based visualization.Visual Informatics, 2(2):136–146

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.465664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3158fd7ca603c741682d422e35a3ff3f60621533f77b5c61506bb3aab4fa5419

Observation e66a46c0-3c91-4517-ac37-8e63025eff47 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Seed1.5-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.585341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d670b0efa1a3592109d54f4979e3d5e5a996bc64b2c2247daeb82d6b76def65d

Observation 33b9e398-5f45-423e-92dd-6c3272960c36 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

Seed1.5-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.723122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:af71c824bc8bbc962bc6edea958b7b18fcde6562783a8bb9f5890561b246efc1

Observation 675eeab5-8790-4c16-bf3d-5c39d5151d88 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Seed1.5-VL Technical Report Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.470167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:77d6cdd811382516969fa3407dda173e3b585327f4a84d7b69a0b635d8f127a6

Observation 6cccf783-ee74-4d1e-b9d0-eac6fdb1ab84 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

Seed1.5-VL Technical Report OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.803350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:996518d5dd182fb3ed5ad9e607694b6abd3feb235d98fb10dd6fadb935da7df7

Observation 6d210c9c-db51-4ece-b83e-e8e59d775169 · outbound

This paper cites The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models.

Seed1.5-VL Technical Report The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.816865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d5a3936bbf90d92babd65c2a704598d3b5eb9b2734b0609d56129f7ddb1c99b6

Observation a7739048-e734-4807-b2d6-c769a53a1e94 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Seed1.5-VL Technical Report StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.849502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ce9a1abdb4b73707585ab2ee1b1153b719e7e889cc3b5f9ed24f1d2d0a42f864

Observation c963f759-fa33-4d40-926c-2d7b7c60b7df · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Seed1.5-VL Technical Report Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:28:28.418803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f54a359ba81f45709913a41370bae61f166c51ecbc9f6c4a4325be261c65bab2

Observation e125210c-4fab-4ab3-9c5a-50a03211a77f · outbound

This paper cites Visual instruction tuning.

Seed1.5-VL Technical Report Visual instruction tuning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.476449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6998d2880b925efb8e4fc2975e4046e97826b745baa3023e64c0b79e8cb319e4

Observation 45680f7d-05e7-4bf0-8c52-88abf138915e · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

Seed1.5-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.946519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a1a64501767adaac6d247fe739e2305db95068d9db7de4e7867728dc28319a2e

Observation 40f7a881-bd38-4256-b550-fd9872191a45 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Seed1.5-VL Technical Report Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.485651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5b428bbae0f6c3384a86fc4564f3e26fb06fab68b04fcdb57d40e823f94736ad

Observation ddb64011-8cbc-4818-b680-416ee764cb66 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Seed1.5-VL Technical Report Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.493387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d7d496fc89b9e3a0941df8b5887bc9aa5168ae373ff711a4c939928200a648a0

Observation dee4ab59-fd2c-489e-8f14-73b3a8735974 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Seed1.5-VL Technical Report TempCompass: Do Video LLMs Really Understand Videos?

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f66af5692c2db2bb76287b894cac3abf7c971c119481002f0fddeeb4dff09dd9

Observation 7cd549ae-5d9d-4e8e-af05-5d325c973b62 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102.

Seed1.5-VL Technical Report Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.498097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4ebdf14576b5e994c3f1c3be950ec4640adfe286ee444538cfa0a44eab46d30d

Observation a34b8a8b-4669-4658-85b2-d5557e30d5d1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Seed1.5-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.989375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:123fcbfa19140a7bab3d1ffd83a33a170de800ae87662594ea36b1e9962610ab

Observation 5b7c0424-0ad2-48dc-ae2b-ae49ede7034b · outbound

This paper cites Ursa: Under- standing and verifying chain-of-thought reasoning in multi- modal mathematics.

Seed1.5-VL Technical Report Ursa: Under- standing and verifying chain-of-thought reasoning in multi- modal mathematics

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.001472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5417c53da1a7f7b68b19144795d88a77552f14911b534258bf9868445a74db19

Observation 4d6b2d37-f2b3-4221-b7d1-cfc6bc8689b0 · outbound

This paper cites Generative Reward Models.

Seed1.5-VL Technical Report Generative Reward Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.009302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:11fee7bae95f46f051ad03dc3a142cd72923a92581ab3091b21ce252108bad7e

Observation fdb0da2f-6ebf-452b-a2f8-a44d1a594cf2 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advancesin Neural Information Processing Systems, 36:46212–46244.

Seed1.5-VL Technical Report Egoschema: A diagnostic benchmark for very long-form video language understanding.Advancesin Neural Information Processing Systems, 36:46212–46244

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.502713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0fd93cbc65276c9a5d3c9670e79dd56b672f416ef0f0fd6bc4436e9911aa2ecc

Observation 7ddd73e2-591e-4fce-99f0-8b25006ba738 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Seed1.5-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:13:07.396442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7705019d3b995fa3127d4d2ce4dd734172b0e2ba5c71dcfb6e75c8bbe2416cb2

Observation 89a6b06a-1091-4261-a221-92a35397ba27 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Seed1.5-VL Technical Report Docvqa: A dataset for vqa on document images

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.507185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:050ac8510d8d9dac9846903807ebefe93342e317c8b4b8fb40cddfe1f6b74417

Observation ba0a29d2-3a8c-4cfd-886f-cf1f7e7539d8 · outbound

This paper cites Info- graphicvqa.

Seed1.5-VL Technical Report Info- graphicvqa

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.511470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9c86e6a19315aa83ad8b4025142e97869051e57907f5c3e9c8b2ff244fcd8813

Observation fe0f7ed6-1c69-4b57-8df3-d082f58edab1 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Seed1.5-VL Technical Report The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.515776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c170c05e17aa6e318e63a583520810b604f29792ed9076c040c8515d5117c9fe

Observation a476d431-18ca-4b33-b46d-eb513dae0cf3 · outbound

This paper cites Modeling context between objects for referring expression understanding.

Seed1.5-VL Technical Report Modeling context between objects for referring expression understanding

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.519936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:21db99879e8093986731018d436dcd8d1e3eaf88544a0c8d30b0868d81078b03

Observation 127ec65b-80c7-4015-9ea1-2867c375c79c · outbound

This paper cites Memory-efficient pipeline- parallel dnn training.

Seed1.5-VL Technical Report Memory-efficient pipeline- parallel dnn training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.524415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:fbccaacccc371df776ee4e51bbd6b0d878b6d0db14c6879f701d170ffaed97ad

Observation 4b8b888f-1bae-438a-8669-c29a28dd0630 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Seed1.5-VL Technical Report Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.528739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:eee080587122e5bcb2696fa4886291f10714e1499f58a596be25382d4f712681

Observation 59e7d375-3180-4bdb-a701-d888ef651391 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Seed1.5-VL Technical Report Indoor segmentation and support inference from rgbd images

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.533141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d76d7adba1ce71ca2989ab4a8540238f344a45e7a0e3f0a08f78949dd9a31725

Observation 9b4338fa-51ad-456e-b352-b7e501089b62 · outbound

This paper cites Gpt-4v(ision) system card.

Seed1.5-VL Technical Report Gpt-4v(ision) system card

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.537573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e2c1ee2533b1c971b242810b8324cbb4cbd2b64d65172af18f93180f391a8b90

Observation 1d8d762f-4c4a-457e-8604-ae5f53943323 · outbound

This paper cites Addendum to gpt-4o system card: 4o image generation.

Seed1.5-VL Technical Report Addendum to gpt-4o system card: 4o image generation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.541908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a3a5d102ad0d690f49f8943e285e6e1bb7920a50cdbf7c496aefbd8fe6e0c022

Observation 7b569b03-50fc-4213-9727-1736b58ec368 · outbound

This paper cites Operator.

Seed1.5-VL Technical Report Operator

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.545987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:85fa33845ae945cff75b0fa9ab1d379d5f8d4bd8a771a288b047fc9a77a750c2

Observation 6f2c3e87-a354-4abf-bb94-4c044ab34a4e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Seed1.5-VL Technical Report DINOv2: Learning Robust Visual Features without Supervision

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.146024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b2769546eb4c19a91a683dbd14762d1b1a53ff15bd578bfcd1335f7a99ae5753

Observation cccf8905-622b-4a0b-8a49-6fe756b8475e · outbound

This paper cites Training language models to follow instructions with human feedback.

Seed1.5-VL Technical Report Training language models to follow instructions with human feedback

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.156357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:313d7fbeddb8730e553d2abf8ae821c60048ff9259f75d8c6b2f268d9221591c

Observation f4f514b0-5bae-4d7b-bb64-ab29ea276fb6 · outbound

This paper cites How predictable is language model benchmark performance?.

Seed1.5-VL Technical Report How predictable is language model benchmark performance?

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.167444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:cd18a8a5a8355c9111f637b827f906274e8b0bc56b9026a99cf09b8440140ad4

Pith citing papers

Observation b01406b5-88b8-4a52-989b-02fc4a1d11d9 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.168739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:0f75e05d8d6b4f89292567d73775bb1d6ca446af629d86d5d170c2b9d943f871

Observation e29e291a-9035-460e-a735-6373514e41b5 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.418167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:640a0c63a9a3084a0d47c1a7c885070d9074a40ffef3437821c176abf2d8daee

Observation 31960e2f-eab4-4db4-a130-b6b293d6ee4e · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:05a9cae8a969ef7ee8c6f21770e4175684e04df95e5235084dc3ce561b01ec19

Observation 93cdf003-0ea1-4adb-ae41-92f3d9eff97f · inbound

GTA1: GUI Test-time Scaling Agent cites this paper.

GTA1: GUI Test-time Scaling Agent Seed1.5-VL Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:55:00.062838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T13:54:59.938216Z digest=sha256:b34904f5c388eb95143b88b3d2ce65ef2381be2e663ce1a4b9de78b01ec79076

Observation a61bccae-b27c-4896-903c-de4f5ce04ff3 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Seed1.5-VL Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.591241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:35d8a0d9319571b81d6a0d0818887ba2d7b4f1719982585c63966ec2fc384548

Observation 00ceeeda-b9c8-4dc3-82ac-78f64201205f · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Seed1.5-VL Technical Report

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:2b8da460c9727b6312ce3e5e2549587d9cd4dd2ca0f1d1d2b5145c17138b1656

Observation 9f6d2274-fe55-41f4-8389-1e9fe22a5061 · inbound

Seedream 4.0: Toward Next-generation Multimodal Image Generation cites this paper.

Seedream 4.0: Toward Next-generation Multimodal Image Generation Seed1.5-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:39:00.920258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T16:39:00.828442Z digest=sha256:4335852690e0cbb22614ed08de4fd7ee64c969b6f7eb0fb347adc8d6190359a4

Observation 4e95f10a-35a2-499a-9591-3c343fe030d8 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.987859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:21dd472d2861d466fa97f98a34d3f089ff74a5d79b9ebc511eae188106f042e5

Observation 0ccb065e-fc53-4243-8d35-1912f4ee181b · inbound

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training cites this paper.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.633881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:71969f5d265172fffc2ee23df8ecbe4ec7f60ee934e0ac1d3f086ec0cf36f347

Observation c6b6c83a-6dca-4ec9-aed3-b060ce150775 · inbound

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents cites this paper.

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents Seed1.5-VL Technical Report

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T08:16:06.590003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-18T08:14:51.102085Z digest=sha256:f41d6b5dc8816688d817f2fba992dbc9fad356e670830d4933becef1598e6c21

Observation 9f752b9d-7723-4de2-aa7b-efa6ad70be72 · inbound

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks cites this paper.

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T21:37:07.589391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:37:07.589391Z digest=sha256:7a17de3b78906dbb9c2b0064d557d6b77685cd9c43b71668633259cd09c807e6

Observation 63e80dff-649b-4887-be08-6a2e5abb1b77 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Seed1.5-VL Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:42:05.738526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:5392889faaf3eb0fce427bc25d6e56427a9bdb6cf6397af42b5f9ea986e96b6a

Observation 33745277-755c-4e19-8869-9eae62425e40 · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay Seed1.5-VL Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:09:04.026356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:223008bad02134c81b70c27e71813f3e8c3ef428905e4d294080b0d0b8b84a00

Observation 15010b91-8e2e-4448-afe6-bf85f8207e74 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.581673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:d166ccf1b3c865888d8b259b6d3f63e72b68ec40bc34070e6213fba071513d3b

Observation f9e5bdac-24fc-4247-8288-958bce8c4530 · inbound

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning cites this paper.

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning Seed1.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:30:50.912992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:30:50.912992Z digest=sha256:eb90d6ce4d42377c90887d26207f8341b26557a70f6b8b752f935e04399b8b1a

Observation 49d30e0b-b72c-432e-bfeb-24d3a70d46b4 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Seed1.5-VL Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:21.904826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:a9355c569d656cb2b7f01625eada4e111c2bd576020cf881545d483e13fdda0a

Observation 8ef4a5d2-12b8-40ac-9ecd-cd22d91fc50c · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:44:15.952662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:da2deecb00955e82d9195e71f4a62e16267d770703858d27cb2297510437f06d

Observation 00b8b6ef-d0be-4677-88e4-d814d681dd06 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Seed1.5-VL Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.766906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.766906Z digest=sha256:373501e770d1caddc82a9bbd8b7555e1f0db56347779bcb44d2f467d0ebd1be8

Observation 3b153076-3b0f-4b79-9444-d69987996b3c · inbound

CountGD++: Generalized Prompting for Open-World Counting cites this paper.

CountGD++: Generalized Prompting for Open-World Counting Seed1.5-VL Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T13:45:45.266385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:45:45.266385Z digest=sha256:511607141ce52c136fe23f1bd4a9f185c57cfbdaa32e25ce54f07d33459b4814

Observation bed1ebcf-69f7-4e96-9e5a-9f763a375257 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data Seed1.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.244864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.244864Z digest=sha256:54713e017ff07bf3d8c6212950449d6d47fc8e34f6002410a7c45a620554c009

Observation 0952cbb7-12ab-4624-af42-810838bc210f · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method Seed1.5-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.145651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.145651Z digest=sha256:3ae4a34e19b81ae8ddb4dd8c4a55e72614c295fc2779cb45a0c2e9dcaaa39f89

Observation 8e15eb89-a3f8-48b1-aeec-ab53b8a4d1ae · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Seed1.5-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:17:51.869714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:bee17da36ef3a970a94c8bf4f5cf8b8f355e6d616b4fe78cccf42a180c279e91

Observation 2ac7a027-5a9b-492d-9e3f-e67bf1daad3e · inbound

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction cites this paper.

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:38.768084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:38.768084Z digest=sha256:0b162df52efdda9f7eac705e263cdbc98df3991b1cc1e309d8c5118bdd3f6af9

Observation 77f142a9-63b0-4cb1-a101-7438cfd479a6 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Seed1.5-VL Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:d4d020fe1d78be31c6b996791dc856367666840b610b8d42a051592ccdadf910

Observation c9d91b91-63b2-4212-ba47-7455d8f22b9d · inbound

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning cites this paper.

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:40:42.201090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T06:39:29.010937Z digest=sha256:6ba332f25a33afd1af600200047998c2c229dd684108fb20007c9e6d72bdc78e

Observation f597c6e8-5eef-49d7-b298-43cfc185f295 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:10:43.152968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:36eb07ceefbdc57f1486241f97563be184bddcf0c5d60f1aabd043fe84d6964a

Observation d46b1046-30c1-4b3a-9694-0a43c78c9f30 · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Seed1.5-VL Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:56:50.375022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:2383542e0a67dc347644d33f323bc8eb7080fdda0defd611eec644481fa678a6

Observation 70193c6f-cc14-436f-8edf-b32e7c284204 · inbound

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments cites this paper.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.921774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.921774Z digest=sha256:f88cfde1f11103c555a9e873a2a87c3c33c1af7486772060e40e5b965576766e

Observation 25d4cdf8-7156-44a0-8371-ae0638bea619 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:05:09.941488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:ecf2ce60a49b9ac2d2be58e184792c690da47743ae80982cd9958388e70f3fba

Observation 9158f500-ae05-4265-8f33-ca70f58be589 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.438697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.438697Z digest=sha256:2f013aac531b9ea0165af8e5555ac6285dd9cead08a892516238c262f0a1f521

Observation 335b71fd-9a91-4702-bd6d-6032358880bd · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.590628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.590628Z digest=sha256:e850ccf476a89a906cce57d69450356d6d3322e0c2b719b00c5833e406d62fb4

Observation e4f9eef8-ed59-4106-b25e-0304f4fb151c · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:8139a0ce0fc5e76de9767f748210d55cb9e8d4be66e0d4a1a4de0455c036a733

Observation c1cbe2ac-e98a-42db-8740-f94b9316b201 · inbound

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison cites this paper.

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison Seed1.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:55:33.279954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T11:54:18.587529Z digest=sha256:85dff68f73d532fc05ed3a72f6e15a0a2c7504ccf82398828f0010aaaf1f5e1d

Observation 705d5836-de75-4a73-a082-ccc5d2e8c2e7 · inbound

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP cites this paper.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:2321ef90b70ac7f7366faa65d6c5268a9a3305bbfc18121e03499bc0d1d55dda

Observation 5d7687b2-011d-40ca-84c1-4979940d7d04 · inbound

Peel neighborhoods cites this paper.

Peel neighborhoods Seed1.5-VL Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T20:08:04.785209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:08:04.785209Z digest=sha256:24f47991ad0ad16bf4d2b6d2bdae7199ba0b8778125e664c5578d45b18251356

Observation 20a792e4-84d9-46c4-892a-1ca8771ec2b4 · inbound

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios cites this paper.

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:13.857975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:05:13.857975Z digest=sha256:33b6b903361f6846c3dcb83238a472139f64f8ed73da6ecb4fd518b974ae1b09

Observation 684d642b-a880-4338-a93b-9e3415a180a9 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:58:16.051935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:0341d9b5bddd550df6463c938f2d5808d282434e91b1a15519b94f2d5994c52d

Observation ca1359cd-6476-4fd8-96b4-b15426c2db73 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:1fa522518185224d3918fe7502c5f11254e6870b33b56a140b358d08d3721b9d

Observation 9f7f9bfe-e4d0-46cf-86c9-40be3fad9562 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:b8daaa12550efc24dc528a6a7ae201110607b8edc8f5dd7291f59366d9fd9aec

Observation 51bd245e-7b26-474a-af68-f51e37231b67 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:a3ddf7f926a48499ed0cb8ac6d51f077fee746579b34e7a9528b756852724322

Observation 4729d643-ea08-4e50-874d-f574b59a4bc1 · inbound

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence cites this paper.

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence Seed1.5-VL Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:20:56.579502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:41:31.544716Z digest=sha256:73be95366c3dd4d244d6ee8391fec5bd76f27b9ca065db5df15e22f4f7f37e05

Observation c15ddaba-36c1-4d49-9990-edd8358fce2e · inbound

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents cites this paper.

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Seed1.5-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:46:07.941683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:57:36.038091Z digest=sha256:c8883fd074c84a5e6447a972c66313caac179114fa9c35b1402343a2ed727fe2

Observation 1313a7be-ad69-4774-a5ac-00d1a6eb1146 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:19464b992bc713863b1bbbaf9d371953fd12df0a7f9cc0ea36c679d33e1f91ba

Observation f6083b10-d61c-459a-a5fa-761cc5435c5b · inbound

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation cites this paper.

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation Seed1.5-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:50:57.521289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:58:06.892531Z digest=sha256:94c284819232810068a58131f12325b91d30a96a70c3bf986c291c52c13cc5da

Observation ac4bdaa8-d253-4395-9cd6-0c806ee9d2ce · inbound

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration cites this paper.

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:41:01.158961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T17:02:05.265029Z digest=sha256:d7315f435f8f195e920a6a1e3e8679ca61f094a7b93bda4e1ec44ca04bb28a47

Observation 9911c8e4-65b5-4c44-888e-b94494bbab64 · inbound

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration cites this paper.

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T23:17:40.794140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:17:40.794140Z digest=sha256:7d6773814342e65cd9a5428a2031d0072cc2a2b617000611435aca780d227b24

Observation dc8eea66-cc36-4ae1-9e8e-f1e993ef632a · inbound

Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation cites this paper.

Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation Seed1.5-VL Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:01:05.280745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:12:42.995562Z digest=sha256:93b4365e8eefe33a254c7f61d226a2a4fc2399f6e05f1e1d84566e4a9151d166

Observation c343d3b9-ca6a-4f4e-acf0-5081ed2d6f6e · inbound

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping cites this paper.

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping Seed1.5-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:59.855359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T16:42:02.271796Z digest=sha256:df0e5304d6d5db9f0f91472a34402d0da9c08302c4765144127fb69e84b6897e

Observation bb5762b0-f8a7-4249-8fdd-0af5a468e403 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Seed1.5-VL Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:41:04.089096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:ac7f18bbff78faa72719c7d23345aba7d34afbe9bcfb2484361cf10398f6ff3b

Observation cd114e61-0326-430a-b985-10dfaf86ebbb · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Seed1.5-VL Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:04.630482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:e8bc6813b9294e1a5838530a88175c78b733c15796f680a704f1d2ffdaa3ccf9

Observation 1ad42240-2d80-4bc4-919a-06069a60de67 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:b3d361d2cf26f78a44378e43529186d18c757825700bb823cbed5cffaa081f14

Observation 8b8c1bf1-2dd6-48d3-b435-21b9c24ecb7c · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:20.921262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:20.921262Z digest=sha256:181ab5150a3ba3c6e82703f01a7c8cba328f428259d410adf47b81824af46b31

Observation 99aedbb3-c5f6-44aa-9dfd-c2bc4e0c655c · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:aa5992ab7c8f679879f4315cb4b1aa91b65a627efb4380cd0811e1f1b38ec452

Observation 30c6fd44-077b-4ec3-a935-a855b3328712 · inbound

Seedance 2.0: Advancing Video Generation for World Complexity cites this paper.

Seedance 2.0: Advancing Video Generation for World Complexity Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T13:34:36.248186Z digest=sha256:56749be3f2bef6e96871c161222a768d1a8e206211650ba63bddc7388ea4c13a

Observation 9fb25da2-9171-4cab-a7b5-87b2b715a864 · inbound

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow cites this paper.

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow Seed1.5-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T09:02:09.075097Z digest=sha256:ecb23adb8ad99b7b0bc6708d3cdc65479f6a72a8ac1c1f664932b103bd5fdde3

Observation b4feee8a-4cc1-4113-89e0-628590d048e3 · inbound

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior cites this paper.

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T07:21:45.633310Z digest=sha256:132fee0ffd89dcd647e745ed8cdbcdec67825a3f3f1dcbd65aea90c361d63fdb

Observation 07575c65-6c00-4c56-98c5-3146de1aebdb · inbound

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning cites this paper.

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning Seed1.5-VL Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T06:31:30.778309Z digest=sha256:0def042f288a4ea8528a7e026e954b4f4c30e0df00e93f4e30cbaca44ebe6d30

Observation 5f3ec7c3-edf9-44e2-9199-03e7af082812 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T05:42:41.112158Z digest=sha256:5e6688017ffb23be951b33d667109cea90f6988ab32399be0bc6395f5c4793c0

Observation 1298f9e2-02e3-458a-af07-8c4ef8c99a56 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.937007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:9c16c7a98cd2992deaea8f9bc31b7b5bb5e252802acc57bd09e647b644ecd30f

Observation 3e3c3bf8-286b-42b4-a6d5-be387d7f8822 · inbound

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence cites this paper.

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:11:04.111222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T02:16:03.854650Z digest=sha256:d2469b472e3b375878737a4f0bc6b5e5b425652cb3e6e099c1277c55e23be9e3

Observation 7deaca6e-06fe-4c94-bbc3-8133b4d2ba07 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Seed1.5-VL Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:05.940196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:f6586c794b73bc8152757541920e120a7de0f6f9a10b5a7f3d7fb1efec7cfb36

Observation 706b69ac-33c7-4e0e-86ba-b603988810b7 · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:31:09.545993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:2cfc3396ff8fa161c251c0aed33b887247804de80bf7fafc1e146ef81a16881e

Observation a9231b2d-39e1-4c77-9296-b322c8ebba8f · inbound

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs cites this paper.

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Seed1.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:13.678645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:41:52.098355Z digest=sha256:3c0abd3d90b4e840f321ec094652c4616054b9c4fe09c0f4e213fd289e543858

Observation 22df99a6-c8d7-4bcb-8dbc-02cb5f53f4ec · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection Seed1.5-VL Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:36:17.378085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:0937720cd101ab252936c6ee40ff788caf36a0ed764876372c3fd2ac32d8dc9b

Observation 7fb67bc8-972f-4fa8-a465-cab0cb1f9ec4 · inbound

Benchmarking and Improving GUI Agents in High-Dynamic Environments cites this paper.

Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:13.993200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T16:55:29.332273Z digest=sha256:b9a03c4abd0d7f130a3652a037d1c8f85449346b19d49ab2c5dc069563909c8c

Observation 7d8bdbba-e9e2-4c21-af96-42e6b75f2e53 · inbound

Benchmarking and Improving GUI Agents in High-Dynamic Environments cites this paper.

Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T00:54:49.351703Z digest=sha256:90f8d4b251211ccd391a3eeeba08f8f7e12f0f05bba82f0355c5a81bb5d390b8

Observation 2a0858e7-282c-4077-b08b-cc03e9aa955e · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Seed1.5-VL Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:06:27.472749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:a11d044e77677216dadc7c5e1a4d7fdec8675fabd7bf5446fcc8d04fe5aabf57

Observation b4d31d0f-9da0-4311-96b4-c06ac22dba24 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Seed1.5-VL Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T09:14:05.986706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:f2f9e8c2a1c19fbcd22debd5c469cd89f71aac5bf49166c3b9ee1c90102eb37b

Observation ddb424df-4a1e-4320-9cb2-b5338fa84e2b · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants Seed1.5-VL Technical Report

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:31:29.805994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:e38950aca0c9a49232a9df22f3591cf55dd742656666b00e11292fba389b014a

Observation c1618e6e-388a-433d-a17f-8c03e4b1e196 · inbound

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding cites this paper.

AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T18:37:48.993080Z digest=sha256:096a879f002acbfa3cd1df9be9fefa63a33a9a977988582b2fce9a806316068a

Observation 3853f9ff-e339-4bf5-a7cd-9cf5f15b5235 · inbound

Perceptual Flow Network for Visually Grounded Reasoning cites this paper.

Perceptual Flow Network for Visually Grounded Reasoning Seed1.5-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:8282eaff86236ece1a1074d9edcea1567f1aadd839e5682022a4b44babd1fb5d

Observation 92b39476-037d-4d4a-8793-1011ece6264f · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:bab23cee6f476751438d74b3af63bd1e5dadc51f7987840230e1940f6a26255d

Observation e566e246-7453-46c6-8e9c-2239f39a2067 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Seed1.5-VL Technical Report

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:d8163345962424a71625ac0f71fb8b579779144b553d4023fb9e41b35a89603b

Observation e64c2ca0-8620-4994-9fab-e199fdd2e8b6 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:337ac47b36ee793c669f58530b09108105d712b54769bc48365fae96b8a508d1

Observation 7b5eb592-f1b0-4888-8e83-cb3e03d2d195 · inbound

Text-Guided Multi-Scale Frequency Representation Adaptation cites this paper.

Text-Guided Multi-Scale Frequency Representation Adaptation Seed1.5-VL Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:27.841458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T01:25:10.261289Z digest=sha256:a9e3a59a7d78eaa770292c3d50edb89c56adcecc734f5641b31a852279ba19a9

Observation e1ae2a60-51fa-4193-aa76-45e2944494ce · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:29.250911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T01:23:37.488121Z digest=sha256:dafebb52089f4cdda3f258ff4c4e3cfcad8685e1682e0413b77a0ec59e33a33c

Observation 1a47c2dd-a9f8-47d2-8d14-1c8c02ccb7e7 · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:07:00.433105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T01:03:30.995808Z digest=sha256:5a370e908822c81fcd55ef4c0948a1fb531a7a8cc445de67a689538623c77b67

Observation 8857fdb6-37ee-4f42-a42b-422fd953fb38 · inbound

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents cites this paper.

Done, But Not Sure: Disentangling World Completion from Self-Termination in Embodied Agents Seed1.5-VL Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:29:49.849167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T06:25:15.510083Z digest=sha256:7210b2244c365e5bd81f8ef8b87ec877346b54cbb8e52a73669f00c1b5c712f3

Observation 9399df48-cdc2-41cf-aa76-70e174d1d9c0 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Seed1.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T02:06:14.932541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:10044b6d9fb343b89ff14f78533172605edd5ba8f0bef752929562fd3aebfc01

Observation 8307dabb-8164-4771-8412-811d9cd009d1 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:01:17.892978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:4bf0b0bf0d942d7a8837c35530dca5601723f71b586955d60cfb55c5a94ee8f4

Observation 8ba3a781-2891-43c1-995c-1ce13eceffac · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:17:07.692946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:00dccd8286c7561a8c2e6f60a5b5201e9a0b54e7b85a630809944a1fca59b1d7

Observation 3d997b2c-a990-49d6-82c8-960b564d1e8b · inbound

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs cites this paper.

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T05:56:43.746483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:43:51.900721Z digest=sha256:40f4599627313f0a7f79279418bbaa67755918a23901dcceb0cff7dcd695ade2

Observation 95d74c2c-fd9b-462f-891e-7549a93c223b · inbound

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning cites this paper.

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning Seed1.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:25.709342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:39:55.109863Z digest=sha256:ce49b2ed75775c7749c8883a1572a11966ee499bd73705ae120a31115c3995dc

Observation f568a41a-3ffc-451d-92ca-f0c9eb2cf10d · inbound

How Mobile World Model Guides GUI Agents? cites this paper.

How Mobile World Model Guides GUI Agents? Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:16:24.612745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:28:34.562344Z digest=sha256:9e017c753a72a0ee2cb78d106f6e38c482f8445fdb5071b79840cb756378c8cc

Observation 590f1c15-5a8a-4c90-b186-a1a93ad60818 · inbound

How Mobile World Model Guides GUI Agents? cites this paper.

How Mobile World Model Guides GUI Agents? Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:05:26.669169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-25T06:00:34.560405Z digest=sha256:bf1671261140d9ceae82b8a4c90dfa0bc41572f4217e17c2790947ea629bfc1a

Observation 77c60c3f-f062-406c-9f78-4abe05f0d8f9 · inbound

AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation cites this paper.

AnomalyClaw: A Universal Visual Anomaly Detection Agent via Tool-Grounded Refutation Seed1.5-VL Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:26.245645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-12T04:08:30.902374Z digest=sha256:d5205415a0dd1ee22fe627dda8c06ae319412bef730b0bdad200a93a70b10b9a

Observation a915e8f3-18a6-447e-9509-ea2dccc0f5aa · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.697607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:553103004ffa727e907657f0be2d4ebd1dec40f74cf1564867c8e193084ebf65

Observation 89c1cc12-55ee-4f13-9398-29448fa0a91d · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Seed1.5-VL Technical Report

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T18:32:34.422468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T18:31:09.883545Z digest=sha256:02cfbba45284ed514efd2c9d02d803efab90de1dc09fd965e3f043499551ff82

Observation 4d94be9a-8397-4b60-bc2d-952613344424 · inbound

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models cites this paper.

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models Seed1.5-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T14:10:06.753843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:10:06.753843Z digest=sha256:3f81497ea0df58a55e261601ebde8f22d8304c7b296101e5c2972b6478dd2a1a

Observation 2c6ab74e-6ef3-42eb-81a2-fcff4d039c5a · inbound

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data cites this paper.

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data Seed1.5-VL Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:57:33.093036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T17:54:50.325820Z digest=sha256:07e40a0ac79bbae5538631217862c500930e8e0d0d6c073006b34cb6b844dd2a

Observation 327c5b0a-0aab-443c-8c45-d424af987953 · inbound

ViMU: Benchmarking Video Metaphorical Understanding cites this paper.

ViMU: Benchmarking Video Metaphorical Understanding Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:55:04.084678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T04:51:28.288476Z digest=sha256:b15168b22e2a23afeedb88d5d7cad452661c3a52f11544b347e2d7aa1e820f5b

Observation af230fbb-e5b3-4c96-bd0c-6cf9a9e2e689 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Seed1.5-VL Technical Report

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:55:04.195768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:84f54fb260f8bdb6b924b3b498664870b69f4c6f4d28d95f0380fa7df7c700d0

Observation 958e1bf9-61e2-4070-a824-c181ecec9f4e · inbound

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs cites this paper.

LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs Seed1.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:08:54.505767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T19:04:15.273932Z digest=sha256:b668c9b51402616d944721463bb8918046d4513e51f14c896b8fc721a05ee81e

Observation d5db45d8-d9cf-4209-a957-23ef81ba112c · inbound

SEED: Targeted Data Selection by Weighted Independent Set cites this paper.

SEED: Targeted Data Selection by Weighted Independent Set Seed1.5-VL Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T20:33:43.443630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T20:30:26.760966Z digest=sha256:ac7a1cbfcd2f6bc24fa8d7c8347cf7324024cd7597f4d3fa80db0dbdf040cd9a

Observation ac9dabdc-9edc-420d-ba36-614fb1a950a6 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs Seed1.5-VL Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:23:40.911345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:974d1cac05bd3f1c0970ca5339d1ce7a5f0b7904319c6b2bc4a38adafec8f199

Observation 77e0a633-260b-45ba-9df4-a5dfe759a236 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs Seed1.5-VL Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:51.125482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:0cc36bca57a538b99e1be7858192f54ed5f61cd812f1500abb1e6b52ddebaf47

Observation 8ec6f2d9-b656-4687-ba11-211bc949893f · inbound

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making cites this paper.

Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making Seed1.5-VL Technical Report

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:53:23.371221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-20T14:48:55.203993Z digest=sha256:5462245bcabd7f1a3dbae3d44a18f8ce5aa9e6b56c5e833bb6fa186eaf2adc71

Observation 717659f3-0fcd-44f0-99e8-71690ae8bfa0 · inbound

FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation cites this paper.

FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation Seed1.5-VL Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T00:32:54.629399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T00:28:48.871454Z digest=sha256:d078eb90b8b859c46ca63930f363a878a880be595c2666a2430fedbf814bbcc0

Observation 6006882c-cdfb-41bf-a927-eae5cae77f02 · inbound

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors cites this paper.

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:53:58.060141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T04:51:32.121026Z digest=sha256:cb8fac35a8bac0ea4d440f4c45fe5c0f456b653be8ec4914cebaf2da5f6e14fa

Observation aff8dc3e-e3cb-4b7c-89c4-9e2da8a5af59 · inbound

RISE: Reliable Improvement in Self-Evolving Vision-Language Models cites this paper.

RISE: Reliable Improvement in Self-Evolving Vision-Language Models Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:39:40.601375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-21T05:38:26.590720Z digest=sha256:5aba24bef18b44b8edf2a69e28161f3a9ae142c58105531b120428e33daa87f5