Pith. sign in

Paper Citation Record · LEDGER

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2409.12961.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.12961 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:02:24.696608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:50.847114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e0925f8d-efec-4110-8251-c6b1daa0f532 · inbound

NVILA: Efficient Frontier Visual Language Models cites this paper.

NVILA: Efficient Frontier Visual Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.235839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ce676a09d036081fdfa21f80204353e96e283cdf6b257b0e55ff077eef8bef61

Observation a0f969e2-d8f5-486e-bdf4-ec96e7c62893 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.747383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:f0be0dcd3ad979d288a27b8e49329a37a4e11384391c192d9546e04b15868d41

Observation 25818a3a-9cf6-4172-babf-0805141c6282 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.369891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:e2821e76783cb46a14eddcd1fc51d9b82774fab9f0e3eecd5a7f297c8cab58fb

Observation 96ae2de7-6e16-41b4-8353-9851ffc6d606 · inbound

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding cites this paper.

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:02:37.421236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:01:00.775721Z digest=sha256:9e98298697c3835968683b1a9bb475fe3bde48f6638abdf4c79927111bb3ac5b

Observation 5c2840ee-d5b0-4714-8645-3ddf4a95a080 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.722872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:a0c89072855f3fa2c0b822f1f0afa75a246bce529aff4a9198a4f68cde875340

Observation 78546f28-51b5-41e9-b238-f48203b0cc7f · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.381147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:a81f787f806da64d96ed12c83eff46e6f41c072a39f15746d9a266d96a322dd1

Observation c77b11ea-bf67-4cfe-b28a-8fe3ed3e4a48 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.995688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:59b8a3bad0602489d00d3b733a6673073816b738729380554877e27f3cb354aa

Observation 5d95c883-b5b7-4803-9ca9-0b08cdbbd3b0 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.933917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:3932c1f9ef675032b6ff03e5c5d3aaff762b342c6a8547a78954e43d667adb67

Observation bdd27281-da2f-4e45-85c9-51cdef6f0107 · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.255586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:1646e39e8a411a31cac5f73f727be2d80b0b902b2ea6aa9ea8c77e92ed20a5cb

Observation f5616df8-bba1-4eb2-a19e-34372d632036 · inbound

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning cites this paper.

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:12:07.016238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:10:57.219445Z digest=sha256:d71a5768e031b95afd71da1cc64c014b7783ec37e07b5c5c6e651aaa3b865352

Observation b95fbbae-b9c3-43f3-8cf9-7f2c29168dbd · inbound

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting cites this paper.

Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:24.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:02:24.696608Z digest=sha256:56be25ad4c231d7aa902dc1120541bc53d5db8b262c0657d990741da900693ab

Observation 88aa53dc-00ca-49c3-b000-fd682ddfb9e3 · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.952463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.952463Z digest=sha256:bcefa427a7f5e01ee609c8071a9d1591dea27aa2a304edfc4f15d6ff900c3349

Observation 658ce853-647e-4c11-9558-ed2cef15393e · inbound

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark cites this paper.

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T17:13:07.966651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:13:07.966651Z digest=sha256:63903827976f282d75d9a609a63f59bee8a90f998f4d831e610cd4ea48c858c4

Observation 782a660f-92ca-44c3-bae5-d9aa665a2db4 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:59.187761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:17a3e2b4a6fad884558439a706b258c77aad79e599ea87f45afe81a71102ad74

Observation 8291684d-8999-4361-96c2-dda099eae3b5 · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.684255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.684255Z digest=sha256:fa1b99b15ede24fa8438cef09efc6b416dd9f1622bd847cb3415b8cea3ee3c76

Observation 79b3cbe0-4f3b-4d46-b8a0-4f9896202924 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:34.081170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:34.081170Z digest=sha256:74b900dd0a6922c076791f8c9b23d3b1bbc442e7505c7887dc56c82d91cffa14

Observation 470f744a-a41c-44e2-9102-190dfa7b92dd · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.612461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:e8effdf5ac6902aec7fea20966d59599cbd3414e36b9b0f60a1e69e49c51fd69

Observation 011af0e0-4d45-4789-8742-668caa3108ca · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.851540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:6fe45fae0ab5caf78ff344d0fb0e8eed5f8a4230b5ab3a60fe782bd415664969

Observation 30140cd9-f595-41b2-a384-a90470bebc5a · inbound

Social Caption: Evaluating Social Understanding in Multimodal Models cites this paper.

Social Caption: Evaluating Social Understanding in Multimodal Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T09:12:18.037151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:12:18.037151Z digest=sha256:610a3e6a131022f900e2d154c9903e549e593a1937e886f5e31bf10137de106c

Observation a45b1f45-aa83-4b36-b733-fce8c1a1d963 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.467842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:f4f1a5b39fd9e4c6496fc10ef1a24a5f7070b632daa64594a7730bd7344f6d5a

Observation 147c87bf-0038-4bb9-9f12-5d276f99f031 · inbound

Stateful Token Reduction for Long-Video Hybrid VLMs cites this paper.

Stateful Token Reduction for Long-Video Hybrid VLMs Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:05.151986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:05.151986Z digest=sha256:7e623eaac6fdc7ea2833500bf48c1a2a7a5812e74a12bef75d8bae6d8b84c62f

Observation d74697c3-69bb-4ccf-9ece-b3f48fee6c83 · inbound

3D-IDE: 3D Implicit Depth Emergent cites this paper.

3D-IDE: 3D Implicit Depth Emergent Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:38:11.440996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T22:34:04.833557Z digest=sha256:385ed41ee122ea70e7473e02d0b7cd7c4f4eda09d9e5438651203c66f52f89be

Observation 6df034c9-fb8f-4325-b306-44d5a82d595b · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.088181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:48:39.444933Z digest=sha256:eae36906fca515406c2df847bd9529b4619acb0776577f0bf7e6287f55fe2aad

Observation 6b37bb7d-5986-47b8-b93e-0126796083d6 · inbound

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding cites this paper.

VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:05:56.390960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:57:42.822121Z digest=sha256:54f7d921d3157ca75b77b38404ba788a5c6dede833447853265317962a332a15

Observation b90458eb-4969-41dc-b8a5-c7ce70b27228 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:08.785397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:06:27.953376Z digest=sha256:2dfdc12c30e6472ba56b0b81a7dbd727166a45a4ec57bd1f7bb8712a2a407056

Observation 343d0087-eb2b-4d90-97ed-fb4a66216af0 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:50:51.275095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:49:41.654207Z digest=sha256:7bded1f6a5ada915039cab206f094f397d1616a346607c55ee70f4804e16461e

Observation 06b676f5-4ff0-45b7-925e-51e09e45b1d5 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:28.052362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:35:59.553683Z digest=sha256:69ec607d072c50fbd85c3be845d94a78b748f19da93a432b1fcf0f1c5040c16e

Observation c7e94571-41a5-4ef3-9a69-a6b7ab33be83 · inbound

VISD: Enhancing Video Reasoning via Structured Self-Distillation cites this paper.

VISD: Enhancing Video Reasoning via Structured Self-Distillation Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:24.064572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T06:08:19.956833Z digest=sha256:7db8d727fb64e1fb1ffa609c1d80d472d5781b0c706bf15d1b5790b5f99bf3fc

Observation 45513e9d-92f6-4ba8-bc30-ada40750a6a2 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.244738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:92b12899beff075b235fe2e5274a00038ecec7dbc5b61e69a1f18b6356c02fe7

Observation 8e43fe16-f2a6-4541-ad1d-63cb1e5b1499 · inbound

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly cites this paper.

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:20.651921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T09:20:32.920925Z digest=sha256:849fd50ba8ce8d5412930b46f94fa23a5775e070fe3666958013e2cdee0c3ccf

Observation bf5a937f-eaf8-4524-924f-2ebead1d02e1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.738324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:8f50619ff8b2b0ba650ef8a277f10d5d1bbeea713c0cffa4ca8692e6f8f6bd5c

Observation 9fc1f3ba-b6a6-45d6-87c6-c59c5d4c378d · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.596700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:6224b167ffcf10b085ff1df22ffe1321dc7c4c14b658f015a492314bb3b3e5c8

Observation 687ae146-f056-42c1-bd57-d3b9df6a016b · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.842034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:abc8e8a5b51632c527dd98c99e525fc57f1ae6319ae124f430643c41d5cc42d1

Observation 89a1a066-f3e4-47fc-9d6e-9f3370537c85 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.170346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:fd1e218636e88d75eceb3e21ba9ae7cfe6affbebb80baea4c92376801e2169d9

Observation d7608799-4708-4045-b623-9ca07ce51d91 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:09:50.849452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T05:27:41.578243Z digest=sha256:dc48b7d2c30d158b33c098b47a00e856b3c923efdc2a7798f50f0e3046733178

Observation 1fa8c5de-9e72-4954-ac08-d538123a1901 · inbound

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution cites this paper.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.749751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:432715f1e993a291b3ee6feb700eb260efe7d5bbdbc1ef4212829187c6b646ad