Pith. sign in

Paper Citation Record · LEDGER

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

As of 6 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2605.03716.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.03716 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T17:53:49.388949Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact10
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c93821f0-864e-433e-9b50-04a5da34a206 · outbound

This paper cites Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.285188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:f0d2e55acf18aead5c1cb33071abccc468310053889c57bda9885e468c80b6c7

Observation 62264369-fccd-43e5-be1a-67c73fb738bb · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:14.252534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:bcadd6879badab3030064f5874c0cc618dbe22e67fa3f4e11c6e3f8382182634

Observation d93218e6-6d02-432f-843f-0e7a18ab5ded · outbound

This paper cites CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.235460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:5e1175f7c6ef0496a74190599394242c6a89af2bf276019f91631ed68eb9417d

Observation 77631125-e5aa-4e0a-b75a-1af9c8523687 · outbound

This paper cites General Compression Framework for Efficient Transformer Object Tracking.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts General Compression Framework for Efficient Transformer Object Tracking

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.296555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:8c1501162cf7351272fa84887b739ada302a581099c110612e6b31627c89a232

Observation 6c6ac4e8-52b6-41d7-84dd-288da0f8b33a · outbound

This paper cites Decoupled Weight Decay Regularization.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts Decoupled Weight Decay Regularization

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:14.271638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:2bbb254eded31d1656629f21f68ccb16201b3527c4a6803d603d1cea86674c6a

Observation b492a5d5-21f9-4d17-92ac-005950b60abb · outbound

This paper cites SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:14.308567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:e838b6dc2fc06ea4350c2fa1c1f1e1a97953372f448a874efa0c99d3b00ad9b9

Observation 61716e06-76d8-4bbc-a511-3bf430a64ae8 · outbound

This paper cites XTrack: Multimodal Training Boosts RGB-X Video Object Trackers.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.322164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:b16b9a5ad4aa7f8d6a432ee3b7111cbdedbd9682f8dec712c21ee260b6d2a784

Observation a76fc86e-3210-41d9-a24c-a27ab7c7956e · outbound

This paper cites What You Have is What You Track: Adaptive and Robust Multimodal Tracking.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts What You Have is What You Track: Adaptive and Robust Multimodal Tracking

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.195229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:bfc7330b56ad21dc2502ecb00cf74476a3fa232cfe6434a8191d5004a9e820bf

Observation c2c7e7bb-4a36-4662-85c3-b1365f2b23ac · outbound

This paper cites Visevent: Reliable object tracking via collaboration of frame and event flows.IEEE Transactions on Cybernetics, 54(3):1997–2010.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts Visevent: Reliable object tracking via collaboration of frame and event flows.IEEE Transactions on Cybernetics, 54(3):1997–2010

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T00:38:35.016014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:10f16267a91d2770e209cfeb79237429d2979ccbd8ad5deb285d1d6d70581498

Observation cb63ef5b-b137-4b8a-a350-1ea06c2fc380 · outbound

This paper cites Qwen3 Technical Report.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts Qwen3 Technical Report

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:16:14.179350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:9917f669a99b53dd0067485d74542fa9551f499ae3e813011500c15bc1f909d1

Observation 170febd2-3627-485d-9308-9ef77a0bec2e · outbound

This paper cites HiViT: Hierarchical Vision Transformer Meets Masked Image Modeling.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts HiViT: Hierarchical Vision Transformer Meets Masked Image Modeling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.223414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:5239b9d4af7ca5717708d85b93e633b196688677e7a14ab5a4057c355d6b186d

Observation 4fb6d82d-bcd3-4a13-92df-b32f1d2d57f0 · outbound

This paper cites DeTrack: In-model Latent Denoising Learning for Visual Object Tracking.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts DeTrack: In-model Latent Denoising Learning for Visual Object Tracking

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:14.210308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:43ff748290c6b0abded161ead7b3a83d17e88c60151f706a755ec76bd8c43296

Observation b51b04f0-ffeb-47e5-9b99-119f979617a3 · outbound

This paper cites Generalization of Meta Merger We further validate the generalization potential of the Meta Merger by testing on anunseenmodality excluded from the training phase.

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts Generalization of Meta Merger We further validate the generalization potential of the Meta Merger by testing on anunseenmodality excluded from the training phase

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T00:38:35.019517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T17:53:49.388949Z digest=sha256:5afdf45744548e784f337e4a6f4b7735952c4668b2fd1ba36b123c36deee6630

Pith citing papers

No inbound Pith citation observations are available.