Pith. sign in

Paper Citation Record · LEDGER

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

As of 8 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 4 inbound Pith citation observations for arXiv:2506.05856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05856 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:59.503626Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T20:34:35.048469Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:08:59.885995Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f443e88-6f5d-4d08-ad89-862e47251202 · outbound

This paper cites Temporal-aware Hierarchical Mask Classification for Video Semantic Segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Temporal-aware Hierarchical Mask Classification for Video Semantic Segmentation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.249004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.249004Z digest=sha256:38b2435edf0a9ed5986740c8608d6c81ed9533b0eba780ae48b97aba12e6320d

Observation 725b2662-33f0-4a09-a40b-44a60a7ca21b · outbound

This paper cites Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.443359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:56.403870Z digest=sha256:8b3974fe31e18a5af3798d65ff213fc82916fda127fff0bc3fa7b20eee29d463

Observation 005689ea-7fa6-4dd9-a440-cf2077ddbf35 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Masked-attention mask transformer for universal image segmentation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.608701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.608701Z digest=sha256:830c2dee1cefea5e33386f9895540342990a3130a7233c8c903fe3e4e3ea5116

Observation 9a9bd9f0-db89-474f-9dbb-3059afafdf53 · outbound

This paper cites ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:56.887838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:56.887838Z digest=sha256:72588bd129b7fc9668b53d63150e42e3781ed4001e5f0e02bbc77a71fff34af3

Observation e9ede233-0518-4a9b-adb7-71326dc6b4df · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:04.082575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:57.082792Z digest=sha256:9ac7fc63680e5e9521074dc566ecb3a8781e189fb2ac1d672d39a7fced42a1c6

Observation dcc6bcdd-1821-4e1d-bdbc-b6f208832efe · outbound

This paper cites X-prompt: Multi-modal visual prompt for video object segmentation.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 X-prompt: Multi-modal visual prompt for video object segmentation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.750415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:57.211327Z digest=sha256:8dec953c45134d34d4c310db02f471a0f43582b5b6330ec9ae3f8b3626eec862

Observation 2248e8af-93c6-461e-8043-9a5845b92469 · outbound

This paper cites Mask r-cnn.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Mask r-cnn

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:57.352999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:57.352999Z digest=sha256:f8a3777740e6d2be96d197dfada400d984a385cce4b738daa0a5efc5d478a7e2

Observation 8c715af6-5be4-49c1-9adf-03527e671253 · outbound

This paper cites Segment any- thing.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Segment any- thing

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.388111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:57.574611Z digest=sha256:17a3a1a0dedb389d9c3ddc279c3ec6017dd1551c24c09604da1fc4cae261753f

Observation 1eec0d54-d72c-48c0-9cef-dd27f930a10e · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Lisa: Reasoning segmentation via large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:57.715692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:57.715692Z digest=sha256:33a0f01fc42737ae7e29643b2c4a4a454c131aac8ac62d5a16d27d4c5ab76118

Observation dffa37fb-3d66-4632-aa6e-62a3288a89c9 · outbound

This paper cites Onevos: unifying video object segmentation with all-in-one transformer framework.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Onevos: unifying video object segmentation with all-in-one transformer framework

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:03.076345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:57.915378Z digest=sha256:75005265eb85a263770d4f0256a583b56acf66fc3b6f4c641a0ac4a2647e818d

Observation bb2f6463-df76-41e6-a09a-21e7517ed54e · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27948–27959, 2024.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Omg-seg: Is one model good enough for all segmentation? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27948–27959, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.653871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.062171Z digest=sha256:d4069f3487fb560951ec650c60815f62ae17ecb92de5b10491da589d4969c10d

Observation b17b2a62-90ab-4414-8a7e-aeedd7231407 · outbound

This paper cites Visual instruction tuning, 2023.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Visual instruction tuning, 2023

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:02.276457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.185928Z digest=sha256:6da1a98fa2d6662b5a7c067433bf54eb089c76a7053e42099933cafbc6c87dd1

Observation 8f8f8368-3aaf-454a-9730-17e3c90ee25b · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Glamm: Pixel grounding large multimodal model

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.929591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.377127Z digest=sha256:40d913bf60ce9529aee7dc57c7cb1aa83124a4d8a9dc2998ae5e06a82232a76b

Observation c39d921c-cc6a-44fd-9681-6a09baecdfa2 · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Pixellm: Pixel reasoning with large multimodal model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:58.523241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:58.523241Z digest=sha256:97429523bd5e3aa1fd56f7ce9764625b102ea78f360077fd7a2eece050df7d3c

Observation dd7159c9-5b50-43aa-8975-22f8c1e711db · outbound

This paper cites Object segmentation by mining cross-modal se- mantics.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Object segmentation by mining cross-modal se- mantics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.625200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.644474Z digest=sha256:758908a0edaf4096e5bf3e0a3ad2ec24684333c8fd4c419a62d4dacd402713be

Observation c747ee1b-86f7-4769-86c8-34b97e4cbe48 · outbound

This paper cites Universal instance percep- tion as object discovery and retrieval.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Universal instance percep- tion as object discovery and retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:01.197110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.788169Z digest=sha256:467c3781b7f5e2e43cf6fbe21f695310718e0a30297b5711a0556ff97054f0a8

Observation 759a649d-20e8-4891-9aef-ac73f5c4a62c · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Psalm: Pixelwise segmentation with large multi-modal model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.878470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:58.981985Z digest=sha256:f49ad0cbc11902668258abce349894c972c9a56382e04b2c5bb0a6efc0773883

Observation 19814800-e546-41d5-9444-5c9be91d9830 · outbound

This paper cites Learning modality-agnostic representation for semantic segmentation from any modalities.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Learning modality-agnostic representation for semantic segmentation from any modalities

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.648495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:59.093530Z digest=sha256:c959f12726e7af2472b8ef4fdfded90ff32317d562b323f2bc92ce46bbb31f60

Observation 4410f5b4-1a5e-42e0-a987-d38127f4db51 · outbound

This paper cites Distilling efficient vision transformers from cnns for seman- tic segmentation.Pattern Recognition, 158:111029, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Distilling efficient vision transformers from cnns for seman- tic segmentation.Pattern Recognition, 158:111029, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.321486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:59.200838Z digest=sha256:db33a8a3e1bbf9fd43c55d16aa899d56559c740cdd2ea30dcc64e067c69a4aa0

Observation 78db28db-7858-4696-b071-c70b71236434 · outbound

This paper cites Camsam2: Segment any- thing accurately in camouflaged videos.arXiv preprint arXiv:2503.19730, 2025.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Camsam2: Segment any- thing accurately in camouflaged videos.arXiv preprint arXiv:2503.19730, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:59.275393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:59.275393Z digest=sha256:c71eeb49f84cb859bf712afd120bd43e248235f6fe5f6257c21f07e636d7cbc2

Observation c941f7fe-cc7c-4db7-95a0-8ab2d87d253f · outbound

This paper cites Segment everything everywhere all at once.Advances in Neural Information Processing Systems, 36, 2024.

Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025 Segment everything everywhere all at once.Advances in Neural Information Processing Systems, 36, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:00.032209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:59.503626Z digest=sha256:23444408029e8623c25701614bf6ad916a519ee4c6c1644428768533eff5c300

Pith citing papers

Observation 45f5f125-8168-42f0-a724-759311ab1d31 · inbound

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence cites this paper.

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:08:59.888437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T04:08:32.413555Z digest=sha256:9e80eb2e6e7e40c8838ae97fa7ac8224a3c33746c4676675224e47e50f8b0e46

Observation 378ac56e-6c33-4603-adf4-bd3f9b767014 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.453337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:5dbeb2f6194f0bb8a4b1e49bba180875ca56343f33576f9929a970969125eba9

Observation e1f06317-7c39-4021-8e94-02a06f54a6ca · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:25:26.135752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:35.738229Z digest=sha256:64a385acf8f57336b916bc39dfae1488b956de8bc972ef0f22dcd45356f3ed0f

Observation bead3897-5736-46bc-bd36-2ff7e272972b · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T20:34:35.048469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:34:35.048469Z digest=sha256:c56086eccaab0afd5e4b41bdeca69285deaa03a584a7d7d8bca5e8563bcc4257