Pith. sign in

Paper Citation Record · LEDGER

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation

As of 4 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2504.18662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18662 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T17:39:23.051951Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:29:06.207663Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact5
  • verified fuzzy19
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ac687d64-dcbc-4de2-af20-c55803a091ca · outbound

This paper cites Waypoint-based imitation learning for robotic manipulation.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Waypoint-based imitation learning for robotic manipulation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.277341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:9dc9481e9033a17e402bdd43486abb21f1e3c124328c82c70e8054bdc4cbbd94

Observation c3c215ac-acc4-4eb2-b740-cb685db73deb · outbound

This paper cites Conditionnet: Learning preconditions and effects for execution monitoring.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Conditionnet: Learning preconditions and effects for execution monitoring

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.273848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:4a601de54e1aadad283fc8c0f40f2d17a535c94728c72fac079d59306e501d88

Observation 57d4b70e-5cd8-4ccd-b61b-61628b5a4102 · outbound

This paper cites Temporal action segmentation: An analysis of modern techniques.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Temporal action segmentation: An analysis of modern techniques

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.270362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:2d95d567e7f1434e84274550abd969a828a3bee294dc52fb9dfd1c8f642723e1

Observation 15714c13-1d33-471e-b67c-ee90fbcb1ac2 · outbound

This paper cites Unsupervised human motion segmentation based on characteristic force signals of contact events.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Unsupervised human motion segmentation based on characteristic force signals of contact events

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.267060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:3523680da63f935879d723383a3a7ca68240b201734103e5dbf81ddb1037a74a

Observation b9bec5e0-00e2-47f4-a717-268520da3f24 · outbound

This paper cites Online task segmentation by merging symbolic and data-driven skill recognition during kinesthetic teaching.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Online task segmentation by merging symbolic and data-driven skill recognition during kinesthetic teaching

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.263638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:d585e3d2bd273f4f832ffc48c95a60da68108383b47386c4b9f8b267d08ef5f8

Observation 4ab1cfeb-dcc7-4413-b941-f7a80f8b88e3 · outbound

This paper cites Movement segmentation using a primitive library.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Movement segmentation using a primitive library

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.256881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:54b7c8717ef4c047891d7432efd658f3742d976f30c3fbfb7d7cccf19d398071

Observation ba088ce4-837a-4e84-a9bf-0a91afe9aff3 · outbound

This paper cites Gesture recognition in robotic surgery: A review.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Gesture recognition in robotic surgery: A review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.253338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:d7c8b87137ed6c432f14b0a0190acdfe077fd9c9b87e1f11bf87cfab4e5044ed

Observation 1d9d18e5-1892-46b2-a331-58dc22e02b19 · outbound

This paper cites Ms-tcn: Multi-stage temporal convolutional network for action segmentation.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Ms-tcn: Multi-stage temporal convolutional network for action segmentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.249631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:2e3539e3a4bb4b4fc91255a7934ad656b281d1b8d875a502ac9208e287f5004b

Observation 363b519c-fe79-4a76-bb41-819e5c8720a7 · outbound

This paper cites Aspnet: Action segmentation with shared-private representation of multiple data sources.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Aspnet: Action segmentation with shared-private representation of multiple data sources

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.246166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:f68bd48f049fd42987b4b760bf8e9dc822c6e8f4dfee46e9d7df13d56c8b65f0

Observation 5390e222-eed7-4a65-97bf-2d5a5e436702 · outbound

This paper cites Alleviating over- segmentation errors by detecting action boundaries.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Alleviating over- segmentation errors by detecting action boundaries

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.241997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:2df73a7d7c79ad7c646bcf850a97f87b5e308f48f9adca45100c5e5b6846ca6d

Observation f02716c0-dea3-48ed-b516-36a66cf8e4fa · outbound

This paper cites Diffusion action segmentation.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Diffusion action segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.238110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:5259f146e65f155c90cf9a70f35b0f16c811120a10af558bf6602d03b29b9b6e

Observation 705138af-57b6-41ae-8672-5b944b61672e · outbound

This paper cites Attention is all you need.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Attention is all you need

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.230778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:df49e627311150bf78dbb323e311b960fc8fee8dad75116c5b3ac5fef1ea99a5

Observation e2d1ecf0-fb9c-4e1f-b5a2-ffb5ac0542ec · outbound

This paper cites REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.092837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:1658ec2c8d2de895a7aa8ed89466843597b47ea554f72aa0cc88cf540791c86b

Observation 911c523c-d602-4153-af85-f3d3e988e765 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Quo vadis, action recognition? a new model and the kinetics dataset

Reference 14

Resolution
verified exact
doi, observed 2026-05-22T17:41:53.003804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:82886d05a0dc96479032175bd581a4deb67a762eee241769c89b20a64c1868df

Observation 64738afd-5232-4775-aa1e-ca4dfe9670b1 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation ActionCLIP: A New Paradigm for Video Action Recognition

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.079251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:3076a5da8667d44fc3f143ef859043d2c4de57b67a0e6df00d5a51cd6e6e765a

Observation 69142f46-84d5-448e-aa34-de5ed6ff2ef5 · outbound

This paper cites Bridge-prompt: Towards ordinal action understanding in instructional videos.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Bridge-prompt: Towards ordinal action understanding in instructional videos

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.260286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:0dae802d59ae5e70854200e72c1106a652f49f071de0c068d8d18c460bf54a99

Observation fd9cf97f-39bf-4d6a-8f9f-abcd2151ade1 · outbound

This paper cites Refining action segmentation with hierarchical video representations.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Refining action segmentation with hierarchical video representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.227328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:b335b92c11211d724a808a6c0c2b741659dc135efd6fe4d472bb81007c13b83c

Observation 8b222591-4a63-4f7a-a72d-06ccbfa997f9 · outbound

This paper cites Segmental spatiotemporal cnns for fine-grained ac- tion segmentation.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Segmental spatiotemporal cnns for fine-grained ac- tion segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.224001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:30ab7a4e04abc3f6b687dc27c46541151c8546415b48a174e00ba2db8ed83432

Observation bc26ed7f-e277-48ae-87d0-b086b5111700 · outbound

This paper cites Recognition and prediction of surgical gestures and trajectories using transformer models in robot- assisted surgery.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Recognition and prediction of surgical gestures and trajectories using transformer models in robot- assisted surgery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.220602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:be883108523ee2000bdc87455d124c9d05bd24efdcf9d02a7a427e8b79d9c46b

Observation dd8248a9-ae94-400a-872f-07474dc55f94 · outbound

This paper cites Multimodal transformers for real-time surgi- cal activity prediction.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Multimodal transformers for real-time surgi- cal activity prediction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.217716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:e70af36327af296e7f1b0498b00d91f001aa7753f123f924461eed0473a07f44

Observation d15fcc3d-a5ac-40e7-ad52-82e7c5442a09 · outbound

This paper cites AST: Audio Spectrogram Transformer.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation AST: Audio Spectrogram Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.072562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:d6ec2109f67e31d229d48af4e3b22ba072bbb7a7d3440b4aede17e59ceb4995b

Observation edd81d41-4f2d-4ff8-9f3c-e4940d130a5a · outbound

This paper cites Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.208715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:f2a1bd4702a9131bd2449ca7f40b97649cea2f514980c393614ccc1c99d4e1a9

Observation d1f2d5ae-f730-4ca7-9eca-fa651ed7230c · outbound

This paper cites Discovering action primitive granu- larity from human motion for human-robot collaboration.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation Discovering action primitive granu- larity from human motion for human-robot collaboration

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T17:41:54.214599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:33b225a9340e3f64566fdecfdf74e7c2c92d59ec2260de1b3e9e146caf9642f4

Observation f895161f-596a-4592-97e7-26e1fe02d573 · outbound

This paper cites See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.086128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:cccace99f88df68e241fcfb6d4370939dc1811316e0c30104802651544330510

Pith citing papers

Observation b9e930da-b1fa-4622-be2d-8a7b1353e845 · inbound

Learning Forward & Reverse Skills from a Single Unfinished Demonstration for Constrained Manipulation Tasks cites this paper.

Learning Forward & Reverse Skills from a Single Unfinished Demonstration for Constrained Manipulation Tasks M2R2: MultiModal Robotic Representation for Temporal Action Segmentation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T03:29:06.207663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:29:06.207663Z digest=sha256:c25c77f12dee98665ded75c9251c964c0724df536ff60f00014c0d265e05a254