Pith. sign in

Paper Citation Record · LEDGER

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

As of 11 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2603.18558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.18558 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T22:32:19.740039Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6738284-e9e6-4e0c-98d9-987e1d095e19 · outbound

This paper cites Qwen3-VL Technical Report.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:0e614281a65c7340ce42bfacc528768a45e853db3154ee22c13a9e99c6a59132

Observation 9958e024-debf-4b83-a338-c1f63c36abe7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:02e40562ad5f243fbf292442490ca7033ac05b12c0b15d4ca650736bfec4fda9

Observation 4ae7a0ef-d2df-401f-a625-01de8a66f0c0 · outbound

This paper cites GPT-4o System Card.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering GPT-4o System Card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:c8b7defe33511f391e4c41000c98590fab54549a14ee3e824a687eb1d23aa5c9

Observation a73728c2-cab4-4408-90c1-b73503e49797 · outbound

This paper cites Learning transferable visual models from natural language supervision.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Learning transferable visual models from natural language supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:0982fb707d35b39a81cf059b60eea49c01c3cc80ef9568113b25d65604b5ff1f

Observation ec097bab-a728-45fe-9453-927093126e24 · outbound

This paper cites Sigmoid loss for lan- guage image pre-training.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Sigmoid loss for lan- guage image pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f4178aacacc7c36845f91fb3ed972f8484dbd6c5d2ad6a96e6fc6b261d507f8d

Observation 82487c28-2fca-440b-933c-74a95e60b8ea · outbound

This paper cites Bolt: Boost large vision-language model without training for long-form video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Bolt: Boost large vision-language model without training for long-form video understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:d6e581e58e9aae0151fbff4dd5791b37f4fdb084e3ff387d65b8e1d32ec722e0

Observation c968043f-7be3-4680-b77c-f8119f3fb747 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Adaptive keyframe sampling for long video understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:afeec2ddd8cb25525ad0abbf4f9c27fbd2316bb1b181468da9be2d4a19b2a739

Observation 8832c725-0f15-4777-a70a-0ff31e2dfc81 · outbound

This paper cites Mdp3: A training-free approach for list-wise frame selection in video-llms.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Mdp3: A training-free approach for list-wise frame selection in video-llms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:c11479b2d239432a0b2f7db08ff5db39800735051471a8e72da222fbea29427e

Observation 2142abf1-54ca-4439-8c8c-bc4c2c3bab69 · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videoagent: A memory-augmented multimodal agent for video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:29a58f0d0e4a66f3f1a524c8f270601dcde979d9953d39c390d2d11ba04a3aff

Observation d44b23e5-5622-4de7-a18b-873e22040413 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videoagent: Long-form video understanding with large language model as agent

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:4cf7cbf4082c03a976f6ad2099c92585558b4bce1982f39209b859cc8be6ac02

Observation fa88792c-7a4d-46b3-8c14-5fae0adfbd1a · outbound

This paper cites LVAgent: Long video understanding by multi-round dynamical collaboration of MLLM agents.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LVAgent: Long video understanding by multi-round dynamical collaboration of MLLM agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:e9a4fbe3d08171b35e02fbfabe43d5a8454dbbc1ce0c1bf0cbc77793cbc6af89

Observation 1a805de4-dfdf-49e6-a27a-88656248cfd5 · outbound

This paper cites Longvideoagent: Multi-agent reasoning with long videos.arXiv preprint arXiv:2512.20618, 2025.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Longvideoagent: Multi-agent reasoning with long videos.arXiv preprint arXiv:2512.20618, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:cf01b705645366d8699887a7d981d34888a9e670c6f63614c40c6c0f1e182fd6

Observation 2fb535fc-be98-4885-b1f6-98241c05d648 · outbound

This paper cites SeViLA: Self-chained video localization and answering via llm.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering SeViLA: Self-chained video localization and answering via llm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:056bd4329eabc58e060f06e8bd5be1f80106ccdd7c7871fdf8bce076cf37a3fd

Observation 8d591e8a-fccd-47ef-838b-c877caf6de4b · outbound

This paper cites A.i.r.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering A.i.r.: Enabling adaptive, iterative, and reasoning-based frame selection for video question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f0fd419ab6cdae7b7668eb87e32eeebe841bfa52a0d11ef83d6e6078bb2641a1

Observation 9ccf10fa-0efe-452b-a796-0381f79a7deb · outbound

This paper cites Video-MME: The first-ever compre- hensive evaluation benchmark of multi-modal LLMs in video analysis.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Video-MME: The first-ever compre- hensive evaluation benchmark of multi-modal LLMs in video analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:fc63f86d19750ba5b9d1d4457a16de4123edbc57df5ddaaa61a9df7bc63b3956

Observation b3cee4a0-0c8d-46b1-8d66-3921a4d4b07b · outbound

This paper cites LongVideoBench: A benchmark for long-context interleaved video-language understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LongVideoBench: A benchmark for long-context interleaved video-language understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:5966f742bab15d6b4a0a7684fa1c95c7925efc599bf5192ef5934b5dcfe54af3

Observation 8f3aa5d2-7c05-42fd-a2c7-ae2e4072ed79 · outbound

This paper cites HERBench: A benchmark for multi-evidence integration in video question answering.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering HERBench: A benchmark for multi-evidence integration in video question answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:9c08fb8e1262e49b668834f8c2f6ec58a46742e1c14afe89c5ac4c9bd0a24564

Observation 750a6999-5e16-44f1-81ce-fbf84f95c254 · outbound

This paper cites Frame-voyager: Learning to query frames for video large language models.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Frame-voyager: Learning to query frames for video large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:2a5e17d55139653b5c89aeeefd0454b9b003d42c36a879a0b363f382d56af9d3

Observation 1fcb63c7-1a55-439b-8732-800df64e37c1 · outbound

This paper cites Flexible frame selection for efficient video reasoning.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Flexible frame selection for efficient video reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:035770e49694792e53bdeea54087fc31ca0385d997aa066e88203ef49d9051a8

Observation 29a2740d-4d32-437a-b3cb-e982ca1d89de · outbound

This paper cites M-LLM based video frame selection for efficient video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering M-LLM based video frame selection for efficient video understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:a0e17338b60d23b73a5fee667d0d5a614e55e487f3f99d03161a9a2788a83e6e

Observation f949c76f-ee69-4056-8799-6da396f15b08 · outbound

This paper cites End- to-end videoqa with frame scoring mechanisms and adaptive sampling.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering End- to-end videoqa with frame scoring mechanisms and adaptive sampling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:7f5962df13897471b88f4dbd85a8cd506ca732f6538cbe3acbc5275ee3820ead

Observation d6b4837a-04c2-4c22-973f-12c158a2ccaf · outbound

This paper cites Re-thinking temporal search for long-form video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Re-thinking temporal search for long-form video understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f26c5ab2aea502e91a0e066f4777b413cc8314d90053094259346701db896f54

Observation c4b392d7-061c-422f-952c-cd1ba676b666 · outbound

This paper cites YOLO- World: Real-time open-vocabulary object detection.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering YOLO- World: Real-time open-vocabulary object detection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:97530adef969746618104c62a564d8349e6f3008b6b717e23f5255be3d044654

Observation 97377aef-fbd8-4043-a728-e47ec53a17af · outbound

This paper cites Logic-in-frames: Dynamic keyframe search via visual semantic-logical verification for long video understanding.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Logic-in-frames: Dynamic keyframe search via visual semantic-logical verification for long video understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:7ffc0f867d347fdeb4e79723774f4f552b55f91c6ed4541a8aa8ebcdfd25a81b

Observation 8ec8ec76-e740-47e3-aa28-e1bfb2158534 · outbound

This paper cites Neus-qa: Grounding long-form video understanding in temporal logic and neuro-symbolic reasoning.arXiv preprint arXiv:2509.18041, 2025.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Neus-qa: Grounding long-form video understanding in temporal logic and neuro-symbolic reasoning.arXiv preprint arXiv:2509.18041, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:920d73e1c694f5b0ea2ea393f8793148ffe159a9ac1865803f74cb83d45bef9e

Observation 39c31ab5-6778-49ac-9a9d-98c233cfaedf · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Large-scale contrastive language-audio pretraining with feature fusion and keyword- to-caption augmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:2d8232385b2ac874baab1fc08d9144842a1e02f287395649c0b9e2a32422c578

Observation 9d4120b4-b2af-43c9-a419-e20f7f227502 · outbound

This paper cites VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:13513bbbba88893af7ad1b447ba7327bb37bfe26b9b382e7488db19eeb2ecdb2

Observation cbdb37c8-e94b-4ecc-9760-a13f1f585eb2 · outbound

This paper cites Videozoomer: Reinforcement-learned temporal focusing for long video reasoning.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Videozoomer: Reinforcement-learned temporal focusing for long video reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f8f11e63d8eeffa120ab966b58a2e4ff50e2d5c750b2715266d3459bb3f9e82d

Observation 94d446aa-3605-4c2b-94cd-b1c4a5e75e73 · outbound

This paper cites Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f5fa930c1fc70a7df6606c2dd33f8df44814e3698040e1fa518690544af3a830

Observation c1e10ab2-780e-48ce-869b-371687cdf43c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:4c624338278ac336ae5d00c037605a87b59655e18fa54b679bc0d1800a24ef34

Observation 7b44f30f-fc35-4945-8818-6d42b7970cf9 · outbound

This paper cites docTR: Document text recognition, 2021.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering docTR: Document text recognition, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:1a14ea3bedbf207d934fbe029e6852ac863da18b31367073914e14c41cf1867f

Observation fc03578b-c0ae-44cb-9aee-0c1352a0eef1 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Robust speech recognition via large-scale weak supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:8af26052a3bab3d6eaa49c41947ce230b9e034269d8275b9dee99b2436a07130

Observation a23808f1-337e-46b7-b49b-431f687413ac · outbound

This paper cites Data filtering networks.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Data filtering networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:24bbc7056813c50cd654783357c1c2d34d101c44435a36ef009bb7e8b0e9648e

Observation 549a1198-f88d-47f5-b6f6-baf0c119bc92 · outbound

This paper cites VideoChat-A1: Thinking with long videos by chain-of-shot reasoning.arXiv preprint arXiv:2506.06097, 2025.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering VideoChat-A1: Thinking with long videos by chain-of-shot reasoning.arXiv preprint arXiv:2506.06097, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:19c23503112bf94eb547d5e1e7f4b6ea4ba5bd113e6375e61be18e7ae8746897

Observation 1a41949c-1634-4641-89f0-d4e9a82a8f00 · outbound

This paper cites Qwen2.5-VL Technical Report.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Qwen2.5-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:d99db61418303d5930dc67353ead12dd890468b60172543eba0e3d506e303f3a

Observation d99c68ec-9514-4525-983e-998dde4490f4 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:6b05e1a3a7d47b2942ada49cde28c256bf52d39c42b2c6e62e818d3bc8908d8a

Observation 3bb763ee-3ff5-4972-8ff9-7fcda555b71f · outbound

This paper cites EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:5751676276f12501e26fd781094a30a0aa00f7c751e5b84ab1222d3f846fc0b1

Observation 1947a29a-bdd4-4bc0-a9d5-acd628fd37d9 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:871fd10c3c52e7550ead9b17ad278a4ea2550be3abf69c1fb588a9bde1b7060c

Observation f0d36972-f3f5-4954-b87a-57bf0651a701 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:3345533e687f93ac8a2c5ce79c6a858fdcfc9ff7d8bbcf32eef8984dee8077cc

Observation ed63f80e-9304-436c-81bf-8eb3892c6e6a · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:24e42d0768d6e3f6a1c64e12eb7b0782baaf5e37d38cdc10b392674ed28849e6

Observation e8a163c3-2c02-47c7-aae1-136e7241d4cd · outbound

This paper cites person.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering person

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:dcce5e4cc2bae3bf3c8b975173aa77315dd13db66faa351b7063d382c6ad6ae0

Observation ea5ee306-92ae-473d-8226-735d3ac467fb · outbound

This paper cites Exit " ,.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Exit " ,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:04580ac517b9999e86d6b0580d74178e2d18de2194010cd44931259d85757029

Observation 1d494086-659f-4af9-a962-d9424631e46b · outbound

This paper cites person sp ea kin g.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering person sp ea kin g

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:ec7760b328bbdf7e8432e24019e8ccc974fce0e4052aa59c890e09c36b679f84

Observation d7802363-5cff-4fb9-ae35-de30c1b60986 · outbound

This paper cites Add an ASR leaf with related spoken ke yw ord s a l o n g s i d e visual leaves.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Add an ASR leaf with related spoken ke yw ord s a l o n g s i d e visual leaves

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:2efdfb04944fa8ecc2538f6200eeebe32e9dcd9e72d300c524237e20554b96ed

Observation 54e21d6f-ac7b-41cf-998a-175e972b21a1 · outbound

This paper cites do or bel l ringing.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering do or bel l ringing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:297650015cd4a6113845e289fa4c8b8cef60e5ca045a4938622f302edad71dd6

Observation 3c799bf1-e1b7-4d68-abad-8d430ca62ab1 · outbound

This paper cites Never make a tree with only one expert type .] [ ELSE : Use mu lt ipl e visual experts when po ssi bl e .] [ IF_ASR ].

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Never make a tree with only one expert type .] [ ELSE : Use mu lt ipl e visual experts when po ssi bl e .] [ IF_ASR ]

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:788b18fc09a65f679ce78d142b7493db7e0fb4db71ce3809d73b601fc974b8a0

Observation 41c5d913-dee5-4753-a11c-67553c717f42 · outbound

This paper cites [/ IF_ASR ].

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering [/ IF_ASR ]

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:db9d69a39b1a339b3564455b3c59ba76ddf6e6c362787b63640811978b1b0857

Observation 65aa8b31-bb53-4af4-b63d-f8bcc1f5eacc · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:cfabae2df2ff09f6df07b843b87e84765e8adb71310b0c600119a6594f377543

Observation 7e02f608-dea1-4cbf-9ae9-5b57503d0957 · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:c7a106062ea97211076638a8bc0a9421bcbd869d61a99b2cc41c35a83a0b4594

Observation 8d2ff39d-a7e1-4033-8e21-c130568ed575 · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 50

Resolution
parse uncertain
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:d37020c8f196d4a3fa8887d767a2f93b4d7db88c3173697595fab0a76bedb149

Observation 31a53722-8804-4dd9-aa85-22d0ae6f6f60 · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:6ae6c128ad353934e55e2f007688167cfc6d538830210ef06072cb8aa2d712d7

Observation 0df2777b-14ab-43aa-bfba-16145fb8d74c · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f995098671f1336966b313e51f813daf45c7391c01dcaaa1f2cc6e9ab5107489

Observation 5ee2a23d-d4f9-43e0-87b5-577f30c086b7 · outbound

This paper cites Same " ,.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Same " ,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:ab07bda46e970c894925b7df6be0b01c82d3bb335a8c6eb1cc4f896ae653d9ac

Observation 3713b8e7-a3be-43bf-b0b1-6e35b97742ee · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:0036dddaf3aef9ddd04ae88c5b19b8e5542e45b904c462243b816ee9fdd07adb

Observation 745758b3-123c-4406-9028-eb863821327b · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:f750f3fecc0434d6c5b57efe7344c700ebe4320fc48c5aad425773a611a97388

Observation 8b438ece-e30c-47d2-acc5-65551b996fc6 · outbound

This paper cites [ IF_ASR ].

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering [ IF_ASR ]

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:6d25b9b11b3d5cf2eb4d18e275765c41de14b2bd2b62d2912c3aa0468acac8a6

Observation df87a3ae-98a2-4c0a-aebe-53e14c050d6e · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:095a9670547ac84926ec326f9613fd36787f086be4ce1bbeb4370d8395b83c29

Observation 11798827-8f08-468e-9e10-9e2750afe87a · outbound

This paper cites an unresolved cited work.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:163d271fab2e1a8e0fadec9a949c0b7d6feb1bc3692084d60d982be17e771f41

Observation 18d0e814-6060-459f-a824-7f1670ead1b1 · outbound

This paper cites op ": " AND.

HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering op ": " AND

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T22:32:19.740039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:32:19.740039Z digest=sha256:81760df6104489b5be6c8791b9a073a8fd0c6a9ef2b9a57b4fc4adcd6cc8df62

Pith citing papers

No inbound Pith citation observations are available.