Pith. sign in

Paper Citation Record · LEDGER

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2505.16594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16594 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:31.527586Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:06:33.139310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:30:58.338283Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact4
  • verified fuzzy13
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a669fc7-60d9-4fb6-934c-512dd9d7dfc8 · outbound

This paper cites MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.003967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.003967Z digest=sha256:ddfac1d122550658a597310bd9bcca40e77905f7f5cd04d81a2fea54d0011a00

Observation b6247258-8cda-4e7a-a7c6-2fc1cfed444c · outbound

This paper cites VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.118849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.118849Z digest=sha256:7df9d7571410a5cdd60c739a8709ba2e09fa08ab13f8b5f6bba0d19fbfc460fa

Observation 3b89894d-f481-4bf4-bdc6-2a0216f92b91 · outbound

This paper cites Valor: Vision-audio-language omni-perception pretraining model and dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Valor: Vision-audio-language omni-perception pretraining model and dataset,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.415561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.210380Z digest=sha256:66b62839d85962a850dda74444f402abfaa96f6cdddca23f2e0fb677cb5ff574

Observation e6e64717-ff25-49de-8e4f-328ddcff60d6 · outbound

This paper cites Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.305348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.384506Z digest=sha256:8e6f0bd1b47993e12796582e19aa6b6c8e4028ddd667755410ed17338dc2722d

Observation 4f7c5dd5-252b-4021-9d20-722220d3cc6c · outbound

This paper cites Unsupervised Pre-Training of Image Features on Non-Curated Data.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Unsupervised Pre-Training of Image Features on Non-Curated Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.192245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.503146Z digest=sha256:fb4fe7e4b07393607b72bb3d93301d3c0a55c949f7baef02e04083a76ad5f581

Observation b997e172-681e-4e16-924d-19d87ae7e4c3 · outbound

This paper cites Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Melanoma thick- ness prediction based on convolutional neural network with vgg-19 model transfer learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.280853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.606666Z digest=sha256:4653c47c57fcb3819a7bb4e9b2db62196dad941fc228b11bae564624539b8f67

Observation 9e3a94ff-32d1-4c6f-bc57-f825a6be1c7f · outbound

This paper cites Pre-training on grayscale imagenet improves medical image classification,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Pre-training on grayscale imagenet improves medical image classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:34.149539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.716523Z digest=sha256:1ce46c4f43c6bd369a2fcd99819891e8616f1c86a5b15b8ef684ec155c29b6a8

Observation 76007362-10d8-40c1-a4fc-dd3890e65c39 · outbound

This paper cites Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:32.020124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.782114Z digest=sha256:24c6e00062386f777ecafea2276756f379496c80ae3d6453a5abfafe6f3611ec

Observation 2f3fedb6-ae83-4424-bcc7-ae7b5f82b368 · outbound

This paper cites Sensor-augmented egocentric- video captioning with dynamic modal attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Sensor-augmented egocentric- video captioning with dynamic modal attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.986897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.879984Z digest=sha256:d040bed3ca8f1dffc6995c105a8a640e9c70e49057c902d86ea677b8116013f5

Observation b5cdd1d4-1dca-4046-ab49-b6c24b18fef1 · outbound

This paper cites Boosting Video Captioning with Dynamic Loss Network.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Boosting Video Captioning with Dynamic Loss Network

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:02:31.865243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:28.972803Z digest=sha256:56884424ddea20e93ef3952b2a3702e4debe7110ebd8174132af0261f2d19d23

Observation 5634ddb2-30e4-4d4d-8e6d-29b3333a1cc4 · outbound

This paper cites Panoptic segmentation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Panoptic segmentation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.036353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.036353Z digest=sha256:098691ff961573194b60ca747d9c0d035a55c8f17199913d7ee4ba16342b3061

Observation 949b1838-94dc-43d9-b8fa-5857cc82c437 · outbound

This paper cites An empirical study of context in object detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An empirical study of context in object detection,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.817392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:29.104499Z digest=sha256:cc71afd61946e1b8970d75a080c49c59a29754eb699df0b937ee2902d017b99f

Observation afa08809-6553-456f-99cb-ef07545707bb · outbound

This paper cites End-to-end object detection with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end object detection with transformers,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.171195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.171195Z digest=sha256:cc16a0238e1a010db1de856c0ebc0b80397d61eb7784fb388e8473ca555a175b

Observation aa1a084a-51b0-4e57-94fe-f26759cc1cfd · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.260811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.260811Z digest=sha256:f414e1bfbda46513d8015af8ef0a96c1052ea5d7ecbd12c939ebf6ec7be4ca83

Observation b85f535a-1e35-4405-ac04-20df1c5043b4 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.351108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.351108Z digest=sha256:872c9052b31d36d03802e8428566be57e2f8e74dc9a54f3027de042ff5f60e5c

Observation a9878bc4-bf6a-46ea-a9c2-b4267b312608 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.424508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.424508Z digest=sha256:bc759d3d5e9e2438c66a9eccc4fd0668224f302bcaf2b8715776685331d2f98f

Observation 9225c940-9e52-4dae-8ea7-c98fe7d9f569 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.688236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:29.503052Z digest=sha256:4868f29340c67c62d92e3112ae82d63417f73667db6b7f2532e83dc06819f522

Observation 2e89f73c-4b20-447d-bb17-e1b31044843a · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Lingoqa: Visual question answering for autonomous driving

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.524482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:29.555237Z digest=sha256:ceb12697151b842072426362189b2a4c4b3cf22595aa06698ee62be803341180

Observation a9ad56f4-133c-4ee6-bf79-03698936ac9c · outbound

This paper cites Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.616411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.616411Z digest=sha256:6fd31e04c8d28e8afa0ce245b479e22ea001fd4e7a0bba5cde27dd9aaff47bfb

Observation 925d6cbf-a016-4395-9a47-23a0f46b9d60 · outbound

This paper cites Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.264972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:29.651915Z digest=sha256:3d93c5cf4ea9c6d711205616d32ad65b11b5af9b0d0d8ff71779fde7e599d0c1

Observation 6b766470-5748-44f5-891c-013d71e87bed · outbound

This paper cites Language Prompt for Autonomous Driving.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Language Prompt for Autonomous Driving

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.686171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.686171Z digest=sha256:b72a69f0b39d33c63d0044a4418fbaf0aa5d62edae2c486c22e2096f6257e8e1

Observation 25fab8d1-523d-4e0e-9df9-39006acd7f1e · outbound

This paper cites Covla: Comprehensive vision-language-action dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Covla: Comprehensive vision-language-action dataset for autonomous driving,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.741400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.741400Z digest=sha256:bc8490c4e0d9bcbdf3a25a48e6fafdfd0c819e08833286a721ef7c47967b6a45

Observation 2b7733b6-b7d0-4bc5-83eb-ee539c7f94d8 · outbound

This paper cites nuscenes: A multimodal dataset for autonomous driving,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks nuscenes: A multimodal dataset for autonomous driving,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.802695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.802695Z digest=sha256:7db1698bdebe5284f560bee2ed01947c1bde623b6f903a4dbf84424e3ee6d517

Observation db60efbf-4e4e-445d-92e1-340dcad7f9fc · outbound

This paper cites Scalability in perception for autonomous driving: Waymo open dataset,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Scalability in perception for autonomous driving: Waymo open dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.858783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.858783Z digest=sha256:64bf280529f97ae00665d59b34cdea97c7bb137647c9c62c3ccad5d85931c618

Observation a54ae552-4df9-4c06-bf31-587fc7e0a5b6 · outbound

This paper cites Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.930690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.930690Z digest=sha256:c8220c9e7d344b337c8094c5c9680fdabfa7c59e3aaa4e67af349182721a2fd8

Observation afa8f4a0-e1a6-430d-9fb5-a6d3c098cadd · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:29.989811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:29.989811Z digest=sha256:944f7dcd1194fa9e6085032a18f370464b7899376b88c3602f69092048acfd28

Observation 2dee0491-bba6-46c1-a5e7-7ca30aeaa4df · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.026643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.026643Z digest=sha256:42458feedd4c3abb9243595f8cb04961fed8dfe640b9b5d01115525e4e7479ae

Observation 45e652a9-39a0-4b43-9bc1-48b7e076c746 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Training data-efficient image transformers & distillation through attention,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.122872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.122872Z digest=sha256:70a00bdaa03e3e7569463b7b8d0b865427e49ca8984fc303ac3f273305474d00

Observation 58587ce4-673f-4b34-bd10-5ab51d599d13 · outbound

This paper cites Is space-time attention all you need for video understanding?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Is space-time attention all you need for video understanding?

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:33.026550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:30.219968Z digest=sha256:cd3f3389075c6b5a90694ba3da68cb55f9ea009940bed608281937f3a47b95a2

Observation 07fedb53-97ba-4763-a910-ed65e8bc7cb7 · outbound

This paper cites Vivit: A video vision transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Vivit: A video vision transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.284228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.284228Z digest=sha256:efd415eef8523d1ddc8a34b1c0a57b473e8a8f50578cb7d8129cf8d0b78095dc

Observation dd8bcb5f-5428-46ce-b46f-026b0068a37b · outbound

This paper cites Tuber: Tubelet transformer for video action detection,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Tuber: Tubelet transformer for video action detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.884734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:30.350333Z digest=sha256:5c716eb336c548d7a1ac2267aa8bdc418c50ca8ab9f0ffa9041c8c3ec3728b5a

Observation 035c7db7-0b19-416e-906e-ce7a1a86a862 · outbound

This paper cites End-to-end video instance segmentation with transformers,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks End-to-end video instance segmentation with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.423962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.423962Z digest=sha256:02de4db9d7e530e38e12679ed437ca343d478f046fee1b29269bf06acc6de197

Observation b5ad08d6-56d4-4d43-95bf-2dd2207e1773 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.614682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.614682Z digest=sha256:213f331b15f66d34beef167661f0b8973c3117b1a5c4a87833adc33e01d9538f

Observation 27b94720-d27d-4e83-a7e7-a7abde96db36 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.732657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.732657Z digest=sha256:431f2be886d12e3664937787e601043a396663abc8560222a324b6ddb848efa7

Observation 8cb8e2b0-b50f-42e5-94a6-69dda43a7026 · outbound

This paper cites Adapt: Action-aware driving caption transformer,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Adapt: Action-aware driving caption transformer,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:30.839297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:30.839297Z digest=sha256:dca47f9201c8db6737cb9220a57d4ab2c69351bce754bb23f289f1ead45c9d5f

Observation 464de7ee-7eec-4f00-a9b4-cda0ef90f22f · outbound

This paper cites Textual explanations for self-driving vehicles,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Textual explanations for self-driving vehicles,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.005990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.005990Z digest=sha256:5034a761a5d620dddb9d5da071d818b99763cacb73d40ad5abdf392a99fca6c5

Observation 8a76d122-1a03-4998-9597-ac78e3d31d83 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.081961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.081961Z digest=sha256:361f774deea843ea1c0bbf0a3ba14787cca69b15886d28f3e00ea70d7e072f30

Observation 980049bf-5867-40d3-a61d-4ba25b3636a3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Learning transferable visual models from natural language supervision,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.163340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.163340Z digest=sha256:13408e1f6b2b231e39219cdabfc57cd3f6346e6d5ad09d80c1d96d542effee6d

Observation b9b456bb-b459-4b12-83a8-43367164c6c5 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:31.265644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:31.265644Z digest=sha256:26abd7cac818daa24bb78fd543c0db15bbaf317c9896a5fa8bc5ab63715439b2

Observation 1cb494fc-6a8a-4209-83c1-d2fbd34d1cd9 · outbound

This paper cites Can masking back- ground and object reduce static bias for zero-shot action recognition?.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Can masking back- ground and object reduce static bias for zero-shot action recognition?

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.723513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:31.322555Z digest=sha256:e1b404bd48ac04191761441bb9d3bdd6aac5c22c245d69f2145308ca17e164b9

Observation 9facd20a-5781-42ff-b0e5-8da5a3718d0e · outbound

This paper cites Another efficient algorithm for convex hulls in two dimensions,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Another efficient algorithm for convex hulls in two dimensions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.612712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:31.379805Z digest=sha256:bb9017ac1b5a1db47038bb5c915a0912f0f529d9d8095f14f9e31d3e81e971d0

Observation b2b4fe5d-c54d-45b9-9f43-7a3d6cf6d9dc · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks Bleu: a method for automatic evaluation of machine translation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:32.472030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:31.527586Z digest=sha256:c637a65239f5c66b5a810c862adb8db22a11be77a11eef5d189399bfe8609975

Observation a349a548-3736-43a6-b28e-36458a90196b · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:28.305093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:28.305093Z digest=sha256:9dffe7a44772ba222da1808f22220786dadbe04d56b234860039f185f10b33af

Pith citing papers

Observation 11b69413-3366-4aa4-99e1-a5a6081f42eb · inbound

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding cites this paper.

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:58.341882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:06:33.139310Z digest=sha256:29f2e4368b2dc0d4ec710322ffa9d460b8af7cc43e46253de84b1dfbb2910ec2