Pith. sign in

Paper Citation Record · LEDGER

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention

As of 6 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2606.07639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07639 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:22:31.310003Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact44
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 534d687e-f1f2-4106-8aff-bed773e4c5be · outbound

This paper cites Visual Instruction Tuning.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Visual Instruction Tuning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.939273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:e175434699e1cda7b9038c7d1a99be0f26fdbe1c2797a61034fabc3f7784c472

Observation 85416010-3373-4ff7-89de-6aa7fdbc2272 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.941516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:96774b39c140e4fd85ac4c09410cee38cd4dee4c18f21b8e2b8e362ca2b9b0b0

Observation b4edbaed-a809-41bb-8800-3283032b198b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.943893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:81166d3d61caa08ebafaaa7688c4c547e70f770774d03d192eb8aff90ad6ecec

Observation 60fafe70-7195-453f-91c4-e1289c0ed508 · outbound

This paper cites VideoLLM-online: Online Video Large Language Model for Streaming Video.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention VideoLLM-online: Online Video Large Language Model for Streaming Video

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.922925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:ca5d5d0837edf7bafaea06a515a64ee61a27fae6c9a349518d1641fe0124cc73

Observation cb9648f3-4db2-45db-9ff8-f7030cfec6b2 · outbound

This paper cites Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Videollm knows when to speak: Enhancing time-sensitive video comprehension with video-text duet interaction format

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.976228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:81d43fa6aab929175be7682b1cf4911faddaacf09c42fb56b50893e356fc1148

Observation 16ce67ef-091d-4468-a7a9-0547ad079b39 · outbound

This paper cites Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.927733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:0e9a09d2e308b5a792f5c5e71d00b8baa0162609655e032873448391f6c3c2fe

Observation 23e42558-8c68-4240-b219-8f04add94e01 · outbound

This paper cites Qwen3-VL Technical Report.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Qwen3-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.934788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:e91a1d527796a9a18d54f2349eea29688e5dd4ad02682d9caabee7906e2601ad

Observation a584e931-e40a-462b-bbde-022ef2d05f03 · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.892593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:35f16bfe5e302e204f1b200543a76307273c0a879076014f7a36e342c385090a

Observation 5aa55b11-e67e-4d3f-a6c2-e7d06f5661d2 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Flamingo: a Visual Language Model for Few-Shot Learning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.897598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:840fac1c26995e760e0e476bd4bfca7ef814b809d6c95f76bf96a46aeaefd49a

Observation e5e0311c-cb62-4fb8-a2ce-2eda2d3962da · outbound

This paper cites The Llama 3 Herd of Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention The Llama 3 Herd of Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.983411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4e0c390918f366f76c9eb9f8000abe9f82bc59d0f7221cc5b78821e4060e623d

Observation 447d702c-84d8-44c7-8173-805dae8e90c5 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.981195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:a5dac79e197033dc02ccfa50eb7ac0f1674fe4556fe87b1d6dae76e1c81eb737

Observation 0f9ca443-8094-4aa5-aabb-c23d93a03b40 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.968523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:a1ba3a57d153d43796ea5fba1a56e0a40500479cfcb4177bf862ea8ad5023313

Observation 9d6e6827-dc1e-47e4-b382-c73b69027ce5 · outbound

This paper cites Qwen2.5-VL Technical Report.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Qwen2.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.913108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:fc48de28d3d5596439566ebb1dc82837b9d9751e8de64df8011ea137de06d56a

Observation d505dc38-4621-4f4f-9ce5-87459b303448 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.910697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:4792a1ac243dd6b4d11e20df0bfd2ede1a9081c202317eb236a1fd49301f5c7b

Observation 6453e69f-deaf-40f8-811f-7358abe13d2d · outbound

This paper cites Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.902355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:a7f2cf29ed35cbdedda8f7689acf6b1927841ebed82d7a562689bc0be930baae

Observation a1ac21c0-ec49-4733-85b9-06aeb3e62c5e · outbound

This paper cites OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:26:17.907664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:b2e5603fe936636e8f26aedad7ba22c5c9dc96284b43a17428a7cbec6b411547

Observation 0701350a-dcd4-404e-ad4c-7204e4177d82 · outbound

This paper cites UnifiedVisual: A framework for constructing unified vision- language datasets.arXiv preprint arXiv:2509.14738, 2025.https://arxiv.org/abs/2509.14738.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention UnifiedVisual: A framework for constructing unified vision- language datasets.arXiv preprint arXiv:2509.14738, 2025.https://arxiv.org/abs/2509.14738

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.917998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:68130b03d52a87a7d074678ccf01ddb5937a0b083b73f31824263bce9453396d

Observation 3c4d1470-6226-4ca6-b937-0bda961159d8 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.925212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:1f3fccfcccab9ad2e6dc4e1f5a92b72d1f410b64dd3e80a49b213207b61b1e3a

Observation d9cb0997-5ccb-4b5f-89c3-df9657029bc5 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.970978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:834fc80a71bec434bb6550104e348837587654807a594c5858b70c9441803083

Observation f77726b6-5cd0-4e28-82dd-7801f22a5c80 · outbound

This paper cites DecoupledProxyAlignment: Mitigatinglanguagepriorconflictfor multimodal alignment in MLLM.arXiv preprint arXiv:2509.14735, 2025.https://arxiv.org/abs/2509.14735.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention DecoupledProxyAlignment: Mitigatinglanguagepriorconflictfor multimodal alignment in MLLM.arXiv preprint arXiv:2509.14735, 2025.https://arxiv.org/abs/2509.14735

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.973465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:fd85e71a4827a8d87378de17390751271a6a971f7078c97a018cce2bcc6100c4

Observation 4cb2791f-4d7a-4090-bdd8-ee8f967339d4 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.930062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:08560b7ac469a693451bcc2a51bedc1bd54f4848e244f97fefd57132bed575ef

Observation 67bf9386-ea11-4b0f-99dd-e994e8c4f025 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.877757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:88473fb842ec309dd08fb59e6d163ea78fa0795356c28f03880fd858ba0eb29c

Observation 371c0666-998b-4f2c-a4cc-9596ea918a4a · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.946170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:3b20118be3100db4e30c289bffd9f6e67f49baf90efbf3111d29b6aa1754cff6

Observation c87cbb39-249b-4dbb-8d9e-978cbcacd063 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.884886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:2b1c486ec44709fcf30c33ea234a90d4299b681490efa9b0ed79b832db3e956e

Observation 3d7ff161-b207-4f19-91ff-0fc4fcd01313 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MMBench: Is Your Multi-modal Model an All-around Player?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.904946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:453bcca3f65a90b623a2bef413860269d7f7dbf3dbbfcca5d3dfdb2f52060238

Observation f3ff5ea9-bbf9-4409-89c5-16109cac203c · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.932464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:88e2d85d880c3793a562f83bcfce659b21231dadb7ed37d8ab42a354c89bd057

Observation 6bdc7b07-3731-431c-8f1e-af0354dab525 · outbound

This paper cites RealWorldQA: A new benchmark for real-world multimodal understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention RealWorldQA: A new benchmark for real-world multimodal understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T15:22:31.310003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:d1b7fb8a0644c927e0e3d123b0cd48b57355cc5b9dade824bbbf6002ad5afa5b

Observation 097fb2a1-fe27-49a9-b5c3-89cd04b0f600 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:26:17.937097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:228138f238e2550b4f790a9a8adbf6c316745fc180b6ff6c86119af9fe0295fb

Observation ff3a96f7-ff3e-48c4-86c5-e012a0edf410 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.915441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:386af63896399b3e91ad11e11696ee6d53a94afaba5aeac810fadad731b31de8

Observation 73fe40b3-7ea2-4124-a1b0-8a33a705c4f6 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.963272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:b22ab5196b8b46b34c4c90eaeb0ee2bd3a20ff25945700b33e2b64bd815593ff

Observation cedf9d78-4bc6-4fbc-b4b2-628871e7c466 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Evaluating Object Hallucination in Large Vision-Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.953615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:842481e8c53612746c6f361b04e7faaa170d5bca39a9b406afbcab64c7913685

Observation 84b003d3-ed6e-47e1-a363-22ad5f129c1f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.948743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:8281c1f2f001688fc1854e7dab52880a15280430942d709a67aca23e1de32d99

Observation 79c455c0-6e7c-4123-848c-947c636d6fd2 · outbound

This paper cites V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.951269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:fad9a68a74e358c7abc976d84aca88352db66df05178c6db474f218849b254cb

Observation ac6e24e2-42c4-40c5-b845-0546edbf288e · outbound

This paper cites A Diagram Is Worth A Dozen Images.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention A Diagram Is Worth A Dozen Images

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.955779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:9fe435db42eea9e0821d69e5f7c68bff04bfd7a857e7f0c808e9d6fd77c38306

Observation cd534004-9d1b-4be6-bba6-129580f50e80 · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.958425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:dcfdf788421b7a35e96aef085cb780ba430e9ebb4e7b333ddbd3bef1736ced86

Observation 8068ad65-ea6a-480c-9149-ba38d04dddc3 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Vision language models are blind: Failing to translate detailed visual features into words

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.978721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:c8ae819bf80b2548d1acb5050f9be53d7ff1fbfe509b5f311b99c42b11b795f0

Observation a0a912cf-793d-4413-8cd4-8a37758a026f · outbound

This paper cites ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T01:19:04.778549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:7ecaf33f05f5dff4a8f93e87d0da9840435eb4e34f2a350b73838c8a1ae4e1c3

Observation 105ca07b-9b98-47b6-b929-cba92bd9d553 · outbound

This paper cites Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T15:22:31.310003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:3ff7697ad2d4bda3189acd4eaa7320da9842198f52d5ce1ed1aa496d55858968

Observation fc4d7ccb-aa1d-498d-b469-0e2de9b20030 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.887459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:1cdf91b4341088a8efb6b846d84ea826d339ecc70bf1314d7b0e0b8233b8616e

Observation 2ceec0df-85e7-4ce4-8fab-9895482f6318 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.890023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:d688bc310fd128b8b20d625fb863e5d49dd4a01ed874deec73a1255487812a9c

Observation a6cfca76-03ff-4d9c-8ff4-4126eae2738e · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention MLVU: Benchmarking Multi-task Long Video Understanding

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.899862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:77e10bc4af2ee98fa7b524b8623e427b831afa77c2dd801ca4ef67c05533ef6d

Observation 63c5fe9e-b0b1-4fc7-ae31-c4533b20762f · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention LVBench: An Extreme Long Video Understanding Benchmark

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.875305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:b4f1e87baf52768a31ae2661dcf39ef8a6f6943eb3ef419b94b68049430b9887

Observation 90702f37-b6a1-43ac-b4c9-8a60aa5978b5 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention TempCompass: Do Video LLMs Really Understand Videos?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.880051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:561b0e2b91a8f7648465782f3cab2582093e20cf8f1a67635fbdd8141ffa9442

Observation 2a3e8e24-21ee-441d-843d-3676576839c0 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:26:17.882445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:24dfe0e1def086a9a0eb3ea7e168a7ad9e5ff517ab7d656d9ad97fc19cf2faed

Observation eb17668b-b0cc-4e76-afc5-6aa23abbb5d1 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.895174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:7fda63c171abe39f4f9bf34ad7fbf1fe72c3f37a7ae121d58064707ccfd06343

Observation 345cf5bf-79cf-4c81-a392-20a444a110de · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.867656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:108a13f004678250fc7e32e3beaf5d09798279d9aec044fc15c74b977785b731

Observation ed4aedf6-16e5-45f9-8473-46938f380488 · outbound

This paper cites Proximal Policy Optimization Algorithms.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Proximal Policy Optimization Algorithms

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.920371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:fa98cf5bf9fa726e25ffbe83825bdb6ae5659950aea8650afa05d2c7abb577a1

Observation 7a305ea6-81a4-47e7-8f6a-0496b716f718 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.960754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:f6d7959e19280353df81d1ba76d8457f6c0a3ef83eed885ca135b2bd519f26a4

Observation e1ab2ef0-4a6a-430f-bc6b-f17efd817e57 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:26:17.870244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:5be87aa2a49a659f30ede7b3ed036d15ac18a15174fa43bdbc0aa9380834da62

Observation 702f9e44-561d-495f-a5ef-ae30dc9bc281 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.872984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:121aeef42d8e3f994bf23635d975ad4e5f069b9dcdeb7f501c20eaa0c3513a5a

Pith citing papers

No inbound Pith citation observations are available.