Pith. sign in

Paper Citation Record · LEDGER

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

As of 5 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 4 inbound Pith citation observations for arXiv:2510.26113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.26113 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:23:09.304749Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:52:48.921328Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:18:13.630096Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c67ea8dd-45f5-495f-ab41-260ab92a9656 · outbound

This paper cites write newline.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.151058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.151058Z digest=sha256:8778800051c43d34b1d975e75435a0d3913e0b2e75c793d839da451ebfccafb6

Observation d6ba7251-ab3b-4e11-bf02-140b50c8fb7b · outbound

This paper cites GPT-4 Technical Report.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.206596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.206596Z digest=sha256:aea0b2fedfca6a400439672433d366bbdcc58bbe1cc43faa95f93bf6dc8b4c25

Observation 818ad9fd-dacf-47d6-baf8-31ac7beebc4c · outbound

This paper cites Qwen2.5-VL Technical Report.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.289185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.289185Z digest=sha256:63d27a16e7eacb740b8f06173db5e698a5a985f6a3b964e1b1cb2b4209c17062

Observation 4362ca1e-a179-4a0b-bbd6-3a949bbea7b0 · outbound

This paper cites Egothink: Evaluating first-person perspective thinking capability of vision-language models.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egothink: Evaluating first-person perspective thinking capability of vision-language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.338324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.338324Z digest=sha256:3ee2e131b6aa2ed92dac67c171fd3c57774b396aa3f928e53b2cabc1a6e5f474

Observation 924251dc-ef6d-422f-9cfc-fcf651a4d180 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.416590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.416590Z digest=sha256:e2bb3a96ff24ea63655dc4afa3afbb462aa533226528197cdfcc29020e7745fe

Observation 65aeb1ee-2fe7-48cd-a543-de5a1b93bcda · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.493511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.493511Z digest=sha256:a5c0a4fb2e55bd329989078859fbabbdc86cdab84ebac33c9dd6473bec8a5689

Observation 8bda65d5-9efd-44e8-89ea-24f8062f3b04 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.592902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.592902Z digest=sha256:086870356f61320265e08feb6d8a84dd9cd121071abbc19b9582dcd1e4c2e150

Observation e1cdb779-09db-4959-a7cf-87bfa3f5cdb8 · outbound

This paper cites Grounded question-answering in long egocentric videos.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Grounded question-answering in long egocentric videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.674055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.674055Z digest=sha256:39f14730cd5390d071cb7794461b6cc2ec8efbf1994039cb34b95e726a9918b8

Observation 00828e5b-4d62-4af0-8a61-3127e25064f2 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.858602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.858602Z digest=sha256:5804d7dffd3baea5ed79007e47d61a4380274eced91b3fd5afa09dda6361f4c2

Observation 24f6e97c-0632-4ae9-967e-a0d4b26fa599 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:03.904036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:03.904036Z digest=sha256:ac078818aed4726763de02b64403d10235e96cc7c4b5660f693b77581c610564

Observation e522de50-7231-4245-8e73-fed5048b2eec · outbound

This paper cites Tall: Temporal activity localization via language query.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Tall: Temporal activity localization via language query

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.050513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.050513Z digest=sha256:6d1ec27704cb9b088d8e5f3ee068b9c61583a3f3190d61d728d1772080c994bd

Observation 9b0f31df-a307-4b22-a7c5-068bec89bbaa · outbound

This paper cites The Llama 3 Herd of Models.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.132173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.132173Z digest=sha256:d3594e65bf9b1984a8d1e059046c4fcabf46032e08aa75d7874140bc1279f204

Observation 40da9d03-56e1-44bc-8e62-0621b546d254 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.198631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.198631Z digest=sha256:81d3cbe0c865c70ae35561de3e64bada639f39a8b7ce1d0757fc6ade108ab9db

Observation 6230194e-7128-41ae-9e16-2bf574635c34 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.272054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.272054Z digest=sha256:dfbb05b4203b63647edc01aec1b966f6e1c888ce46eab0c3a4146bfe3e3ba284

Observation 204b3c85-64a2-4538-ab67-1203535fc192 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.331779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.331779Z digest=sha256:5ae0a8ada140c3464004abc17dcd5ad6506c7162563af1a3a8c727566370f77d

Observation ded4086a-74cc-4ade-a563-ab2d0cdba6ca · outbound

This paper cites EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.393109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.393109Z digest=sha256:c815d9143270cfdcc46734ae8d568369393d86c20493592cadd378f38438872c

Observation f04e3edc-26d8-48cc-9088-56a511600f8b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lora: Low-rank adaptation of large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.467417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.467417Z digest=sha256:38e3817a06998b103e7ec7e48153bffd9510decf3a7552aea6712e406ef4ea2c

Observation a3bcb6c2-a76f-4a19-99c6-c2d2e8cfc221 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Vtimellm: Empower llm to grasp video moments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.548784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.548784Z digest=sha256:b6dd1473a5510d547e4727c9941eddee9bde1e8a9c64dc1d5847f5d4523eeb81

Observation a9995c41-a9e8-4e83-9bf5-67a22059a791 · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.625353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.625353Z digest=sha256:900f7768772c0eef5615106ef4bc60ea8ee61c94d9ab64863e9019305434468e

Observation efe5f9e7-559c-47b5-981e-e66e4ce67e81 · outbound

This paper cites Background-aware moment detection for video moment retrieval.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Background-aware moment detection for video moment retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.678137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.678137Z digest=sha256:79809e2df4c66c9030dcdbb15f369cc1742c4c3a49514fcf52098578a84e6534

Observation 72182d6e-0cb4-4c48-8c12-713a858c1fd5 · outbound

This paper cites On the consistency of video large language models in temporal comprehension.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding On the consistency of video large language models in temporal comprehension

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.783148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.783148Z digest=sha256:a8b5774249fd86fae8127fe12e0ad05b1dca3f2d7e050dc5896c21e76cf08ca9

Observation 7b956a3b-ee12-4fd9-9ec4-217354adda91 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.846166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.846166Z digest=sha256:07ccc053c1471680d967354bbdde03b9b393e3f022e39281d67ec6ae1176fa40

Observation b7eb983c-3338-409c-9783-3fcefa8a01d5 · outbound

This paper cites Ego-exo: Transferring visual representations from third-person to first-person videos.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo: Transferring visual representations from third-person to first-person videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.923689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.923689Z digest=sha256:61f67eb710269bbc8be8b84e2b8a6b7f0a9f18abc95c3d934290e9e3b4789748

Observation b9ea4eec-9311-461a-8e4c-a37c6bb2a7d6 · outbound

This paper cites Egoexo-fitness: Towards egocentric and exocentric full-body action understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoexo-fitness: Towards egocentric and exocentric full-body action understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.994341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.994341Z digest=sha256:162f3a2fbb782bcec6be8803c9caad8025d118aee8f41a0632ae9539b08c9791

Observation b220119e-658b-4c3a-b551-f46cb3136cad · outbound

This paper cites Universal video temporal grounding with generative multi-modal large language models.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Universal video temporal grounding with generative multi-modal large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.004849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.004849Z digest=sha256:f9c3a0d56768106326f8b0d0aa8c47f0bfef66e49dde5e73959113835cd33d01

Observation 85ebea19-bc89-4b8f-a73c-c97ba66a00b4 · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.079698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.079698Z digest=sha256:1de9582250607cfe6272a75edfe99bb20a8f1ce42ed36d1f11f03846b72afbbd

Observation 1e29ac5f-6fcf-45c7-a84d-1367cdd3281a · outbound

This paper cites Is Your Video Language Model a Reliable Judge?.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Is Your Video Language Model a Reliable Judge?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.208635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.208635Z digest=sha256:2fb970778aad3adac5fcc788adfb6550dee155f217f4483197beb2aa9d1053fa

Observation 576e35da-897f-4682-987b-2841d57b5bef · outbound

This paper cites Put myself in your shoes: Lifting the egocentric perspective from exocentric videos.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Put myself in your shoes: Lifting the egocentric perspective from exocentric videos

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.325842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.325842Z digest=sha256:24c13004db8f560fe16b45b4fefbf57eaccc857deeed7952b7fe0d0049a54168

Observation 28cdc54f-0505-495b-b2ab-f5e82d964198 · outbound

This paper cites Viewpoint rosetta stone: Unlocking unpaired ego-exo videos for view-invariant representation learning.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Viewpoint rosetta stone: Unlocking unpaired ego-exo videos for view-invariant representation learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.453908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.453908Z digest=sha256:9b93cd9e39cdbdef5e26d4e4abd6c54e6371d2c62569de9955a33de2232dfd22

Observation bc64ff55-0d78-41f6-ac22-f0bb340aafde · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.578638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.578638Z digest=sha256:040e44ab8be86fe28f7f879972eff4e5f04625d594d1c45a75ee9cac8535d69b

Observation afac3c60-b6db-49b0-b1f4-a3103d8b5ea1 · outbound

This paper cites Chrono: A simple blueprint for representing time in mllms.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Chrono: A simple blueprint for representing time in mllms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.660746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.660746Z digest=sha256:078c463742a38770139dac01e59a465256bd6962276c4e43864b7b076e0e2a6d

Observation 398e02e3-7de2-4b68-b19a-4efd16ad01ba · outbound

This paper cites Introducing gpt-5.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Introducing gpt-5

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.765100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.765100Z digest=sha256:65ec206c639003f6a5d2c23eee3ce155a464a2f72e1b08840056d24f36f3a103

Observation b0d565e6-f2a1-4b25-91ef-a253ef8806e0 · outbound

This paper cites EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.886731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.886731Z digest=sha256:a281004263ff3fc7850b84f5fe6b7184c9cb1346b9c096701cf05c9f0115b3b5

Observation 3a1bc3f7-50ca-4392-8f64-b594b993cd5a · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:05.997667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:05.997667Z digest=sha256:9723baf0f34dbff2056242611f4d77c6ec35bc0623793abbae8a9f086ec4e1bf

Observation a609b8e1-bbcd-48cf-955f-ee55acf01d50 · outbound

This paper cites an unresolved cited work.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.099820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.099820Z digest=sha256:483e1fcee036f42abc4a0e20025b99ec623ad8fec4fbd92ae4acfeb278100ccc

Observation d7998972-807d-446f-a68c-d266f04d7fc0 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.211860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.211860Z digest=sha256:dc6886006bf4feea07da1e2afd18491a6eb3ae4673c767e83fc6dc5e6d509436

Observation 62aa2089-f183-4bf9-adaa-b73f4c9b87ad · outbound

This paper cites Assembly101: A large-scale multi-view video dataset for understanding procedural activities.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Assembly101: A large-scale multi-view video dataset for understanding procedural activities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.284683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.284683Z digest=sha256:371a0810c83f1ab355b37eb08fd3aeaab832f54cc00034c90f57295f4213b5ba

Observation 07c9a36e-0566-4394-bdbe-94e24e9b92c1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.372618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.372618Z digest=sha256:c3c17e2e2e0ad51e905c48d8a2c9b9baea6a6bb2711ca3c90943d45ddcd4afb7

Observation 058739ef-4f61-4791-94b1-27af30bbb50d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.466398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.466398Z digest=sha256:571dcf51ffe39f5b0fcc9bfc47322efc07ec93d0164729b61df0e98dd882dae5

Observation ebf6d6a1-6ad0-4b0a-aa6d-c88e799f7210 · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Actor and observer: Joint modeling of first and third-person videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.616649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.616649Z digest=sha256:31834324b4dde7ad7059babb2e2da34b96c41f7acb544384c151e9a745a69d61

Observation 75fc005a-618e-4748-866b-52ccca80e918 · outbound

This paper cites Learning from semantic alignment between unpaired multiviews for egocentric video recognition.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Learning from semantic alignment between unpaired multiviews for egocentric video recognition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.803270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.803270Z digest=sha256:04bfbbcae8654a27b0b7e83a2016e61eb21c362c8db1c2ecef0a55a0ca43a057

Observation 451f6820-401e-48a2-80a4-a0c50b14720d · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:06.927026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:06.927026Z digest=sha256:592c79fa6c792c9883395ebfa0a9329b6f17a2fce66935c7870b90b9193c8f69

Observation c1154590-e902-4fc2-94e3-dfcfe8eeecbe · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.036700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.036700Z digest=sha256:5440451e33bcc79549b43b5b0547a87f2f6b49b4fd3dd1a45a309f3e12c41f5b

Observation 092dee13-de5c-4c06-89c5-20e794ea9026 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.108809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.108809Z digest=sha256:880a067c91f8688a0394712383e39a0c44022448b9525da698cbffaae6ac0349

Observation 5b064333-15cc-4f7f-9284-4c40a6317838 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.235364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.235364Z digest=sha256:d4b1293208fe660099150e53411c5a83ae722a424799a38e88d36f179bedd6fe

Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.344850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.344850Z digest=sha256:e753736561f2787938299c1cb548b9f4c002b35f94b672d2c7c3487e7be992b8

Observation 21ea69d5-bc46-41b4-9c8f-a27a96bdae55 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.484753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.484753Z digest=sha256:355109e1d6927355bc627ab1752b53971d482a42104cd876bf0ceaab99d77260

Observation c7089af4-5a67-4c21-b09b-25ae1ad9482c · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.609169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.609169Z digest=sha256:5bcb61389da2969829e5bc16a0dfb05d7041e44223dc4f4e580edd21afcb1ad5

Observation 538d22a1-53db-4d21-a70a-015b0e21c93e · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Can i trust your answer? visually grounded video question answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.649531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.649531Z digest=sha256:60a3582c32a9ac3ea5f3174fc89604ed418894cd6ed5be65a55c9a38f4119e3d

Observation d49cc283-871f-40cb-8554-8c45fbb62f66 · outbound

This paper cites Egoblind: Towards egocentric visual assistance for the blind people.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoblind: Towards egocentric visual assistance for the blind people

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.699810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.699810Z digest=sha256:af55f699646f658391623da42783ec97a44ab7a83963e7f90c56b2295925bc03

Observation 7b3121aa-73da-422d-9565-dc6e6dfe64e0 · outbound

This paper cites Text2Reward: Reward Shaping with Language Models for Reinforcement Learning.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Text2Reward: Reward Shaping with Language Models for Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.784752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.784752Z digest=sha256:8853360e366d72dcbeef495b7333de05cf1b527daa4ba77508c0803e3a08c7c8

Observation f3c79df9-f7ee-4172-89b7-79e38a469b74 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video question answering via gradually refined attention over appearance and motion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.882110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.882110Z digest=sha256:f34253e9e913d75f71e5d8fed4d2b09a4fa92936f07c091a8b773f99b207ee89

Observation ec8daf29-db2e-48ef-a8c7-c7a745bd8571 · outbound

This paper cites Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:07.961142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:07.961142Z digest=sha256:6712347b9b118d6f7e3b8e0723168797c7ae9f956a795b9b17341b70ef2bea71

Observation da52ac15-2194-4fb7-adb0-42283e78c112 · outbound

This paper cites Qwen3 Technical Report.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Qwen3 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.030046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.030046Z digest=sha256:79bca7c6065574ed3209c91332f06d42a684752989c6d934db4dc932f3cc29bb

Observation 7e57d479-bd9e-4d1f-9d1e-c94d17e07899 · outbound

This paper cites Mmego: Towards building egocentric multimodal llms for video qa.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mmego: Towards building egocentric multimodal llms for video qa

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.145926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.145926Z digest=sha256:c0957b4cc73749ad85825bc88215b2558c7b15bd32a2338cfb26ddf1bc910fe9

Observation d68df159-68ed-458a-89c6-b4686498c6a7 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.224753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.224753Z digest=sha256:150e992d06345703f119a4b1e1a62f46d1553e0266b67511496ac1d5fd9ef639

Observation a5502050-69e0-428c-8a3b-03e29de7de9a · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.354942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.354942Z digest=sha256:0d815dd27613802c268b977f5d97771b17a9d7d072578d98c0c5d25ac4c02501

Observation 8a4eb131-d608-4e6a-b021-23fcbd145c25 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.454785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.454785Z digest=sha256:23d110a907d29d3090bac645bcdb91d80fc83b9d7d4606be30a7f35f1794af92

Observation c6803488-262b-4bec-b68a-a90d1ae860ff · outbound

This paper cites Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.584753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.584753Z digest=sha256:502927e2fea878a3239581bac24db518695284e57c2877d3f688bf92461421a6

Observation 738943dd-7f6f-4f19-83b6-12b008fa9194 · outbound

This paper cites Long Context Transfer from Language to Vision.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Long Context Transfer from Language to Vision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.764864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.764864Z digest=sha256:619cc6b2efd90e2037d665b4671ac7dd57b9449e15f9afd3722768d7ad5742b2

Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:08.954750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:08.954750Z digest=sha256:0712332df6d280169ad9c1f197d4285b43722ee797fb8022724b55c1e540dd8b

Observation 6eaca711-a83a-4a5a-a29c-ad30b18f7463 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:09.070998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:09.070998Z digest=sha256:9bb5b635485047a41a93e9bc8c527f5d3072c4db84d02c1b44325837f7e3f0ce

Observation 510b3cf6-f4ea-4ced-8f42-9f164f620dfd · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:09.184170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:09.184170Z digest=sha256:aa4ee507620ad54311594a0ce0366effe247c5e73067696dc01d03227015fb78

Observation e63d3dd1-f44a-4a8c-b06f-f90c0545d318 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mlvu: Benchmarking multi-task long video understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:09.304749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:09.304749Z digest=sha256:9f6efd2a97f0570ce15c678fbb5c78dbdbbc72876359fd28df269b01756a2ceb

Pith citing papers

Observation 886e0750-ebdb-4523-9acd-ecf55f158a3e · inbound

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning cites this paper.

EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:52:48.921328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:52:48.921328Z digest=sha256:cf330f54b663e34bbd77d00449d5ef5fe8c86e52e047875a6906d04cfd80d36c

Observation 109af9ed-c9ae-4cbe-83e4-953f092c6891 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-23T02:11:59.295007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:28cc3791e92490cb93b32ad56e0b0fc6fd71d7d3b101735b031447816583c048

Observation 765d3134-f777-456a-9bb8-1b59c2d926c8 · inbound

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models cites this paper.

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-23T02:11:59.295007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:36:20.388169Z digest=sha256:43ca89b8c94aa8f8082dcf2808e1c3cf97b0d34f52c4056d94de815efda6ef70

Observation 883e2ed7-8b35-4b16-afe6-26bb39d7abb0 · inbound

EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos cites this paper.

EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-23T02:11:59.295007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:18:05.563078Z digest=sha256:48b3b84df2e3f1f37ce7687ed0192ad98b40edcc50043e8983eecf054425bd17