Pith. sign in

Paper Citation Record · LEDGER

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.13364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13364 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:24.303188Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact8
  • verified fuzzy35
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d59a1872-489d-4553-9de4-41191b6890f5 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.223191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.223191Z digest=sha256:5a2d32fe5c31e9bf7afec7ea9488038cd4f6aac512da9c6f9193331f176b9966

Observation 58c6bae3-fdc2-4237-8b13-d67425b5116f · outbound

This paper cites Objects that sound.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Objects that sound

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.301171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.301171Z digest=sha256:0321091310b03c15188710f52bdc67ca054f538af05a7cceb1edd16504051fa3

Observation af1c03c2-eed2-4c63-8061-3e9de9304f1e · outbound

This paper cites 3d seman- tic parsing of large-scale indoor spaces.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d seman- tic parsing of large-scale indoor spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.441879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.441879Z digest=sha256:20baaf9efab89b02e207707a43c97c3bff7570ffa892203d181a8c29cd00f77b

Observation 39d3b6b5-846b-49fd-87e5-30dc0c0f116b · outbound

This paper cites MAE-AST: Masked Autoencoding Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.572549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.572549Z digest=sha256:30c9cc668407de9c9b8f5afffa848356b9c82bba934ee4c60284782c2fa8cf97

Observation a21ca1f0-fa77-4a38-b1f9-6cd59471f3d0 · outbound

This paper cites Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.702418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.702418Z digest=sha256:b9aac1ad943060cda994ca35ead7b126e284e716c0fe8a60463b8c67356761f4

Observation 2bea425a-ac9b-44e4-8ad7-d70b29a3bf8d · outbound

This paper cites Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.854860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.854860Z digest=sha256:96d75a8374d1d8536ec8efb2848667467add599471b00ce74ef30209aa4d66e9

Observation 3c461085-a472-455b-901b-a4c5fd5b9a5b · outbound

This paper cites HiP: Hierarchical Perceiver.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning HiP: Hierarchical Perceiver

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.988580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.988580Z digest=sha256:32c38305a8bd8120984ab635f3f8c0dbc492c35994a3f495c4d6f868c6dd14ba

Observation 56e36f43-8424-40ea-a6a1-99d718592339 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.091237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.091237Z digest=sha256:29c988b2ba696b3ff538199de7502799520427e8fdf8fc0cce8a1b944de86878

Observation 54ea8e34-ffa0-4245-9962-ef434d3f3ae7 · outbound

This paper cites DialogSum: A Real-Life Scenario Dialogue Summarization Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning DialogSum: A Real-Life Scenario Dialogue Summarization Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.207070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.207070Z digest=sha256:98c158fc1d7adba24a24a38d1e3a1ab0f6c788fd1f95766e2de65385a7534cd9

Observation c5a18620-9c84-4893-8dc8-db6d1adb2641 · outbound

This paper cites Multi-Task Learning with Deep Neural Networks: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-Task Learning with Deep Neural Networks: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.356316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.356316Z digest=sha256:5979a203b632b514123d79ac381d75be66a1759958db1c9276802ebb1596f571

Observation 56770bc8-7652-4291-a236-385ac86ae989 · outbound

This paper cites One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.526503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:15.468308Z digest=sha256:3c2c31c38e79604cf0c39be587b361a3e6ca4f7379f62be05f2f2dd5e1018a16

Observation 9f3f72e0-b581-414b-8320-fc49d0f25b45 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Imagenet: A large-scale hierarchical im- age database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.576775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.576775Z digest=sha256:f0541f9656f1b42684ad2e6d2732dba437d1f9a9cd3f49d7d0e634e19c469f71

Observation fd8e2161-2dab-474d-aafc-305379b80e06 · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.745612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.745612Z digest=sha256:2f7cc4a7c2d0d72a18937a35c2c9770c83636071e2090eabcd8ba44e5249ab53

Observation 37f95f8a-075d-45ee-b344-173b91ad14b3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.911782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.911782Z digest=sha256:a23235ae15f1097e6bc42492ab8fbc211168f0e77b7a28e3d9dfd7c0af8b54f1

Observation 55f8eeca-e4a0-4882-b8a4-0aead2363c93 · outbound

This paper cites A generaliza- tion of transformer networks to graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A generaliza- tion of transformer networks to graphs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.035240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.035240Z digest=sha256:45bdbe95eeef185278fb8fd4414b1301d937288e66e26774d6c9f480e4140322

Observation dc85f918-01e8-42ac-9e37-65f9d313ef58 · outbound

This paper cites Efficiently identifying task group- ings for multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficiently identifying task group- ings for multi-task learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.197878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.197878Z digest=sha256:b23cafb13cc08819a477300b131855339108c4b694dccbe1d00cd7f9ecaac295

Observation dae79af5-b80a-4199-8116-20bd3ce75789 · outbound

This paper cites End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.324991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.324991Z digest=sha256:df3356210253144244f22e833a347c3ef4411fee9a215ee9423ed6a1bd7b32d5

Observation d9b07949-7982-4622-b307-34c03a5d24c0 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audio set: An ontology and human- labeled dataset for audio events

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.490217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.490217Z digest=sha256:8fbfe7f6e141cc3c3c0dadae3a743383faf8db82a83309aac5a646a3a7c6f343

Observation 932aad78-8e47-4db1-aeef-9cf77236476d · outbound

This paper cites OmniMAE: Single Model Masked Pretraining on Images and Videos.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniMAE: Single Model Masked Pretraining on Images and Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.285243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:16.598120Z digest=sha256:52f65d214a3a3c20e39b357b4a449fafb93518d71a465d1cb52e3759972bca54

Observation f09ffe82-9b83-4d02-aef9-ae2198364411 · outbound

This paper cites Omni- vore: A single model for many visual modalities.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Omni- vore: A single model for many visual modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.711982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.711982Z digest=sha256:9e789a6000582193f2dbc82f7ca875b60f31b08567e234e84e241bb424fd254e

Observation 475ea48a-1140-4d20-9ed5-3805e5e0fc1e · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.915486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.915486Z digest=sha256:5452eabac3f0d352a0fa9ad7c74d653140dc8913a82a5f1a5a686dc07ca43c55

Observation c5ee46f4-7d1b-482f-8cad-1e106d61c0a3 · outbound

This paper cites AST: Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.039560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.039560Z digest=sha256:cd9c0f4499e7654b5b712445cb1b15df67794f15debcaf7c5b9fbf343eca63de

Observation e59a712b-21b1-4fb7-81b0-22fb95d4f43f · outbound

This paper cites Uavm: Towards unifying audio and visual models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uavm: Towards unifying audio and visual models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.202842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.202842Z digest=sha256:3ff2d9be6bb518e751b4ff1dca708ac428386e16da00c98468584fbd5c354ac1

Observation 1ef8eb6e-cdb5-41bb-9fbe-0567c56ac70e · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The” something something” video database for learning and evaluating visual common sense

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.322899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.322899Z digest=sha256:baa733edc25cad3712d0fca83e91b2d5f35d0b87694aefd2b8ede44490f9a80f

Observation 4684aabb-05a1-438c-b075-9b2859961c47 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.420096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.420096Z digest=sha256:f3bb5f4e29a821e7e791cf79847e043612e90212b320ca218c6f9ae21c49b78d

Observation 718fe2f5-fe32-4d5c-84fc-5ac5709e6bdf · outbound

This paper cites Dynamic task prioritization for multitask learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Dynamic task prioritization for multitask learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.507498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.507498Z digest=sha256:55443a91df9d931da0d391b21a6ed1c390819a25e1c42cf6ac1b5f9f9bee200d

Observation 87e16809-6d0e-4eda-b546-ff98c30ed626 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.617134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.617134Z digest=sha256:32bbe925888fd7cf9b6b5a3723e9e9ded9dd4f1da38d7ee77067476ca6c33204

Observation e0b9e3d5-d3ef-4db0-92b5-c89935dbf407 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked autoencoders are scal- able vision learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.712917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.712917Z digest=sha256:7aaf7c19bf8ace8c0bc6d4ca0fa152f4cc12ae5417025c5d285f80aca646909e

Observation 0a79934a-741c-4e41-b197-0f059c137274 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Gaussian Error Linear Units (GELUs)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.826839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.826839Z digest=sha256:567488fa73aa88ab0bb1a25c93a6b4314819ea30336546afd9756a7d2eb5b864

Observation 1f89a59d-b2a3-4f03-bbb4-dad5dfd4add6 · outbound

This paper cites Spectral- former: Rethinking hyperspectral image classification with transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Spectral- former: Rethinking hyperspectral image classification with transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.932549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.932549Z digest=sha256:70030b611f7cfaa62e55a7549372d0df3c0afab989fa371834c4ee5cd3c56784

Observation 621186a5-ba8c-416d-8ff9-bd93cdde9a52 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Unit: Multimodal multitask learning with a unified transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.021377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.021377Z digest=sha256:b1e733637b4f9dd27158be4d1ad39351aa8510adeb5a8213b4006c2c147dc828

Observation 84660675-7b4b-4f5c-82f4-da2d19e23c2c · outbound

This paper cites OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.119583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.119583Z digest=sha256:e09210b91aa591e8956523bdac34bfde14a7b9a0952b992ca12e27309cec6a78

Observation 342a32dc-600d-4380-a433-3c3e8138eded · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.251387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.251387Z digest=sha256:02108de05bf86b243aa4aecc6a56c33422643c3f0700e7d86f4b5af6708e9260

Observation a5d0eb73-3086-4f59-aead-67f1a0556924 · outbound

This paper cites Perceiver: General perception with iterative attention.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver: General perception with iterative attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.379108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.379108Z digest=sha256:8899cd8942ecd1e0c784297b4c33eaf57794e1d752878c2151a1c10e7e470ae2

Observation 9d9f546e-fc24-4b2c-89fc-3a190f2f92bb · outbound

This paper cites A review of multimodal image matching: Methods and applications.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A review of multimodal image matching: Methods and applications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.514139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.514139Z digest=sha256:f00c3e5d870eb245fc4c1cf0b1c9ba5209eee4a658424ef29cc1a8724a350bbf

Observation 9cd1fe4c-4482-434f-b3e5-f327b37fc8f1 · outbound

This paper cites One Model To Learn Them All.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model To Learn Them All

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.598502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.598502Z digest=sha256:8c774440cc89a59ea16b6c427be93766945af947f0ed9fe0009985fe8af671d2

Observation 2c9ea7e5-0131-4774-9b13-11500b736714 · outbound

This paper cites The Kinetics Human Action Video Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The Kinetics Human Action Video Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.711195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.711195Z digest=sha256:c0e3c3c7163baf6c36996bd5f5337dff5fa51b4173a9a2b72be748e46af554eb

Observation 30d3f7e6-cc88-430e-8730-dc40862b5df6 · outbound

This paper cites Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.963938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:18.839182Z digest=sha256:d5f1a32096918f48995eebc68af10a20bcd09d7a3034bfa99c4b4f629078419d

Observation 0b31472a-8eda-49b8-9476-252fb415548f · outbound

This paper cites Re- former: The efficient transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- former: The efficient transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.937345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.937345Z digest=sha256:9618df1cf0f4e78cb364e385a446414a5b4887e5f3b31600951aecae16e22712

Observation e7b145c0-c1a0-4ea9-b71c-f24d22135333 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hmdb: a large video database for human motion recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.051884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.051884Z digest=sha256:0517862d01b801e2fc1433987bd9440e5d028ba440e0200668840a1c8c710790

Observation 1345fdee-e667-447e-8af6-ab88c4dd7d26 · outbound

This paper cites Modeling long-and short-term temporal patterns with deep neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modeling long-and short-term temporal patterns with deep neural networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.166703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.166703Z digest=sha256:53502b38a7c27410eadb7ed61b1e92583c7066241457afd2fd43d1caaad4f50e

Observation 2b5de6d7-497d-47c5-9c91-c72f11041bf1 · outbound

This paper cites Stratified trans- former for 3d point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Stratified trans- former for 3d point cloud segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.290419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.290419Z digest=sha256:d4fb006b055f7cc0760ab9281e95005b4d43b38552f376845f19e062ebf9737e

Observation 6230888b-edf8-4947-8121-8823afc854e5 · outbound

This paper cites Regu- larization strategy for point cloud via rigidly mixed sample.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Regu- larization strategy for point cloud via rigidly mixed sample

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.311366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:19.428323Z digest=sha256:915dba55ecaf3d3bbca1731eaaa3070afe58fab0ca59567113eb061e1a808715

Observation 1d5913ce-11ba-4c0d-8c9e-28202a97a75c · outbound

This paper cites Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.222632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:19.541910Z digest=sha256:63951e9211554e97f8cf3295804c246bd78b0b5deef53820a8529ffbff125f61

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:daa81614e46edf46727e3a65dd3d3b8a6bc332a696cbd4dabf63ba21b78de573

Observation 941ae255-0a70-4627-8651-0196319e1ce4 · outbound

This paper cites Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.074110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:19.768028Z digest=sha256:4f56e91e63841d2cbef391cb5bc11ebb80a2108094cbdf5ac03c4a53f1b2ca1a

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · outbound

This paper cites Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:2c09d8dcb2a90617475ab232a9e3fbd604f6aaba7182b507d925866f297ba6a3

Observation df00d131-ea42-45f1-bd8c-76ec203b2e1c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.892880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:19.953525Z digest=sha256:d537ed9ba5fe26117261ac9f799b6e7d732dba72cacdb9ccd9b1c3424700f08f

Observation 56008d9f-43a8-4667-b043-2ab6eeb4e977 · outbound

This paper cites OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.761552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.051811Z digest=sha256:74e6c33678f4f1e557e59d8a0491f58ec6ca82f50124b65eece61d89c255b566

Observation 8f652af2-877d-4b29-90cc-df16bbcac204 · outbound

This paper cites Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.679062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.168295Z digest=sha256:bdc3f67b54ec87087d4f7f06e4f81e222037835ad83e144e3f6cb883b44cd03a

Observation 08f78ae6-58ba-42ce-b471-ac7950235867 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.304795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.304795Z digest=sha256:e938d292b4468a4aa7e60cd706ebbfebc7eea9083816785d189b31a177a28b65

Observation 194acbc9-711a-48bb-9adc-f519ce36393a · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Moments in time dataset: one million videos for event understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.539943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.398454Z digest=sha256:3b3198296928662d970f1b0df5810f4bbc978d976bdef851e76c08c7320aec0d

Observation 43b4d318-5029-4061-b742-6bc3205a9dad · outbound

This paper cites Person recognition system based on a combination of body images from visible light and thermal cameras.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Person recognition system based on a combination of body images from visible light and thermal cameras

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.423488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.525721Z digest=sha256:9abf8000a167c35f06a6060e30d921fe8f71ae90ce95bad290d7128881e4391f

Observation 43fa27db-de94-4685-827b-56ae6cb30588 · outbound

This paper cites N-BEATS: Neural basis expansion analysis for interpretable time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning N-BEATS: Neural basis expansion analysis for interpretable time series forecasting

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.661670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.661670Z digest=sha256:13ecee2c4c2fdc92fcc2beea37aa81800efb4e6ec08fc1c061ac172a4182b8d3

Observation 60e8e3d8-5162-4883-a951-b6a1b9b0e488 · outbound

This paper cites Cats and dogs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Cats and dogs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.767129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.767129Z digest=sha256:9fea98d46cf5705fffea35de06d9fd58a419661398a9683f582e2a3360b95d9b

Observation ed785505-9347-4868-91d0-3a3ae50cc0af · outbound

This paper cites Esc: Dataset for environmental sound clas- sification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Esc: Dataset for environmental sound clas- sification

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.311763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.883668Z digest=sha256:e7c4890c0a625fb9e86e77b91979f21231eb0b1dfc3b2d69cec24c721db5e9e9

Observation 6b0d0f90-f56e-4ab7-8257-a0b3f156ec28 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.184763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:20.990939Z digest=sha256:bd601c1bf1b723e848ac7061df9edc4ae02d7974a511c7e64d23f1fdebfd19f5

Observation 8a5f51c5-9667-473c-9240-5e22df4fd717 · outbound

This paper cites OmniNet: A unified architecture for multi-modal multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniNet: A unified architecture for multi-modal multi-task learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.121011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.121011Z digest=sha256:9de5fd209c5410b92f0c67c715d4a83b2b402084c7ba81dd340e7eef05252a08

Observation 960b3122-f46a-4655-8e8c-c6780eec6014 · outbound

This paper cites Point- net++: Deep hierarchical feature learning on point sets in a metric space.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- net++: Deep hierarchical feature learning on point sets in a metric space

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.040826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.235366Z digest=sha256:29edd6f1916ac0e4d682075516eeb0bf21409d3dc074df9fdfc708d559dd451b

Observation 85a502b9-ace2-44b5-b02a-a1445bda232b · outbound

This paper cites Improving language understanding by gen- erative pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Improving language understanding by gen- erative pre-training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.380262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.380262Z digest=sha256:f6411b8219d4fb58c077e21be36d124303fcbe6ffa96133af1af126173983941

Observation cebe27b6-d809-402e-aa63-1b51e2012ace · outbound

This paper cites Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.946155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.519423Z digest=sha256:b6a8abd892152a71f93acab264b9e6c55d1432d89ec43f77cc559e244b21cbff

Observation 5f88d135-92c1-4203-823c-dee0e0d6a492 · outbound

This paper cites Zorro: the masked multimodal transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Zorro: the masked multimodal transformer

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.551900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.672452Z digest=sha256:c3648ffeb5eb17974d3fbc92ad78a363cf865bbb0850fb83e5b6f90cb57bbdc4

Observation 3f24eccf-3cc9-461e-a080-e9ab0c57190f · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Indoor segmentation and support inference from rgbd images

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.817984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.776607Z digest=sha256:8f6fbf4ba2a7e490396259a1bcb8b7b931e9f701458d465297ad24eccfccc109

Observation 21993166-f068-41a5-84ad-b57e3fca20bc · outbound

This paper cites Mpnet: Masked and permuted pre-training for lan- guage understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mpnet: Masked and permuted pre-training for lan- guage understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.698263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.861378Z digest=sha256:ce322a1c27b6adad2b335d2396b7f6f117f8627c58f484df342e1b1416828228

Observation 9a540032-3568-429f-af77-4e6fe5f3d6a0 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.580490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:21.997522Z digest=sha256:9c465023645d99a125bd44bf33919deebae3896410d9b2860d193819e7343d0d

Observation f8a1fc2e-e727-48e9-b9d7-b1eae07b61c2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.073655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.073655Z digest=sha256:a3204e46418d72738916d34fc1f790a40023352651e529a5b3f3c06be4e276e8

Observation 3fed73df-e32b-4259-aa76-b7be8cddd991 · outbound

This paper cites OmniVec: Learning robust representations with cross modal sharing.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniVec: Learning robust representations with cross modal sharing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.147490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.147490Z digest=sha256:fc4a2e4d5bb042bb4ea3afd5c14db5ee11f613654aeb8082bf578f487210c808

Observation 8049192c-b68b-414c-a96b-c586b6cd28ba · outbound

This paper cites Hierarchical multi-task learning via task affin- ity groupings.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hierarchical multi-task learning via task affin- ity groupings

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.455919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.210014Z digest=sha256:f2d0f688b0c2c44bc751c462c6d345f1bb746ddb45f3b4eaca449927eeeb0f3e

Observation f3edae28-32b7-4164-8523-4d5b73918175 · outbound

This paper cites Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.270903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.270903Z digest=sha256:ffedfaa8b41899b8e06754fe401f8e8f3cb78c418d6437eb9b85b347d5d29d04

Observation edf6e8df-359f-4bad-8d02-a2c8f3fbc02c · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.333424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.333424Z digest=sha256:dbff9e0b17d97f2d335dd0b0e80d68aab07163e759a1e82545ad8f592d573956

Observation c367eb61-daef-4882-96a4-71cc92cb0c7f · outbound

This paper cites Contrastive boundary learning for point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Contrastive boundary learning for point cloud segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.308893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.389608Z digest=sha256:98faed13c8a32efd00229b86578c1b185943fd9830abbed0f2de0c38f05f4311

Observation 3271b91b-af12-4545-bb64-f5f2e23434bb · outbound

This paper cites Small sample hyper- spectral image classification based on the random patches network and recursive filtering.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Small sample hyper- spectral image classification based on the random patches network and recursive filtering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.150417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.462428Z digest=sha256:e13ffdf5488f55a4794a8e685c52288a497816d3d03bfec710649da48e8c34a4

Observation f1b25de1-73cf-4f19-b089-6357805ff965 · outbound

This paper cites Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.959279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.537912Z digest=sha256:b47f31a6447a83a3fc98e9dc0e383d58321fc7ce25a49c9636ea32268b0b6d6d

Observation 327fdeff-ef5c-489b-9c7f-53dee17d507c · outbound

This paper cites The inaturalist species classification and detection dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The inaturalist species classification and detection dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.779706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.632705Z digest=sha256:feefcdf3d31d59a28af42996c1f95cbea8c6077c92e92fe0903bfc8ffab1c65f

Observation b6988d14-15b5-46a2-98e9-f821a035c5a3 · outbound

This paper cites Attention is all you need.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Attention is all you need

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.627959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.716235Z digest=sha256:b1e8fac940eddbf19a8af45f8372a98893c71bfdf8f175a42720d69d83e610e2

Observation 11b94fd8-c192-4958-8fbc-4900b5be83e0 · outbound

This paper cites Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.467629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.775789Z digest=sha256:ffcbe27450081e05f0fe1d8c984626eacf8ab97c745c856ada8e884a8c817e19

Observation 90af2435-5042-4201-ad4c-2e1abd6c5506 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.840001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.840001Z digest=sha256:08595a391aa89a976040a6d1b91aea55567911bf49ad3f1978859f2d9d383c9c

Observation f9b85ae5-edfd-4db3-9b20-6bd5903666be · outbound

This paper cites Masked feature pre- diction for self-supervised visual pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked feature pre- diction for self-supervised visual pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.260075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.894679Z digest=sha256:b3cae1d82a23113bfeb8405e96a9c8efe2475381e9d680109d9a7be2bb2ab9fe

Observation 349ba8bd-f659-4ebf-8a76-276997f08069 · outbound

This paper cites Syn- cretic modality collaborative learning for visible infrared person re-identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Syn- cretic modality collaborative learning for visible infrared person re-identification

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.084751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.941403Z digest=sha256:f1f98b03dacf7936c2509cc9dcb0abd1bd7d46f17a45944a205294e331a0a479

Observation 0d42cd48-9c56-4dc2-846e-3514925e8c9a · outbound

This paper cites Controllable Abstractive Dialogue Summarization with Sketch Supervision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Controllable Abstractive Dialogue Summarization with Sketch Supervision

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.321727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:22.994077Z digest=sha256:1cac85d73ee60cb1ff1813e6110e1168ab1f181504e99d11813f4d1369309708

Observation cecf0c2d-d737-4047-a8ac-88955a12dafb · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Tinyvit: Fast pretraining distillation for small vision transformers

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.942665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.044939Z digest=sha256:3702175cb0008c1b25889f610048ce54d22bfcdab5428c3e329801c5cdbfbedc

Observation 825288bf-78c8-4094-8cf0-1eaa9b17c048 · outbound

This paper cites Point transformer v2: Grouped vector at- tention and partition-based pooling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point transformer v2: Grouped vector at- tention and partition-based pooling

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.793382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.101262Z digest=sha256:d6cae00f6b7a8839579fe6df5b85e8a1833507c2f95dbd4b297efa9852d76db8

Observation 4c15fc0b-37a7-42c0-a0ca-8f888624e3ab · outbound

This paper cites 3d shapenets: A deep representation for volumetric shapes.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d shapenets: A deep representation for volumetric shapes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.654139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.180431Z digest=sha256:0cf960bb78bf784f5f0ecd7d075a0a990c663cbf4bb52e3bfdff7b01448dddbe

Observation d4f48735-2087-4877-8af9-04073b46d0bd · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audiovisual SlowFast Networks for Video Recognition

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.240129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.240129Z digest=sha256:f56428c462ad2231832de512cb0225f3631fc867718a2314cafb3813233d7d47

Observation bfb374f0-7601-4427-96df-5ee5faf71b0e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Msr-vtt: A large video description dataset for bridging video and language

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.456028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.286212Z digest=sha256:ea38bc9b7dee349f3422ca53436e420b0d53ee2006ef837330a8bea55a5b3093

Observation e0c2cd51-f150-447c-a73f-a643a7c81994 · outbound

This paper cites Multimodal Learning with Transformers: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multimodal Learning with Transformers: A Survey

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.148397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.399732Z digest=sha256:41c1970e20383ae23a0520155455998cb6d7efac71c6fd5505f897e52c08a29d

Observation cfa2cd38-10c1-42a4-8f1e-7a24a37a3490 · outbound

This paper cites Multi-modal masked pre-training for monocular panoramic depth completion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-modal masked pre-training for monocular panoramic depth completion

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.330526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.499703Z digest=sha256:00f75f7a8d9247b935a661a6f8f39dbbfd0be34268534ba69c98acf2321363e3

Observation f955fc72-ae65-4e44-bf53-cb3a6af81b53 · outbound

This paper cites Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.580652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.580652Z digest=sha256:22c7b520f2d0d5f85f822a97847f36aad3b53d4e7ce0d03766b36c6ec99ca4d3

Observation b008ece1-7827-4814-802a-93820de787d8 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.682264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.682264Z digest=sha256:b024594cd49fad25be9e66edbcbf09b24d9bb86999864c1b487bbb8e6e6ac77c

Observation 25e42e8f-0a25-4e26-a6fa-ea9b67d18dbd · outbound

This paper cites Deep Learning for Person Re-identification: A Survey and Outlook.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Deep Learning for Person Re-identification: A Survey and Outlook

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:24.969127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.740334Z digest=sha256:72aa6f90fae9e4fa0a3699304387c68c310f8c9a648b48c3b3196dc57376f911

Observation 240ad846-4d80-4e93-a896-f29ed0221efb · outbound

This paper cites Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.189988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.790277Z digest=sha256:eec651d69f8ed0e41a565369d209c9bb28802e92d21c254740500683baf3b23d

Observation 7aceffcd-fbf6-4549-99b2-da442e37cdde · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.861350Z digest=sha256:9ed18218f45e345c4f1647a9ab24cbe0fca9fd733d848babf5920a26d6191038

Observation e1e514a6-2fda-4a17-8cac-dda7c93bbe04 · outbound

This paper cites Metaformer is actually what you need for vision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Metaformer is actually what you need for vision

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.047216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.921175Z digest=sha256:edfe8e4af27049638fd43bae86018da38e57125ff753d97410c27a9d0d2b342a

Observation b4baed03-004d-4915-9ad5-4c3216d2b102 · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point-bert: Pre-training 3d point cloud transformers with masked point modeling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.928246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:23.991722Z digest=sha256:ab731de6a96e5baa4ca675b63f36d334fb5e33921290f57c3f545da6a0809f08

Observation 4ffff764-3d5f-4d5f-8564-951c54b457e6 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.066566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.066566Z digest=sha256:ed6a85bf98c69491195a988d0f6845654dba87f324f4b45efea8d34410d35840

Observation b2ebe9a9-9854-4749-93c2-f4b57288dbc8 · outbound

This paper cites Point- cutmix: Regularization strategy for point cloud classifica- tion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- cutmix: Regularization strategy for point cloud classifica- tion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.750612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:24.114385Z digest=sha256:2c4b50860c5ea32a547f2e5a6c93ae1dd05da40c47328358c8435efb799c39bb

Observation 521fa700-e1e6-4b00-8163-3c2a9f01e95c · outbound

This paper cites An overview of multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An overview of multi-task learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.594862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:24.164336Z digest=sha256:400efe3c9ee34ec4db377a11a49ce604e513c63f330f713c0bf93d076a4ba601

Observation 3953969f-63e7-4d49-91f8-48982b077969 · outbound

This paper cites Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.512561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:24.215744Z digest=sha256:c07c351d5ecdd7f28e01f8f373239dd1f669d130117d60c21793e8edec1d03dd

Observation af7aa99d-b892-47ac-a506-e0154c640fe8 · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.267156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.267156Z digest=sha256:42a25a0835c622532b5cb3562d9956e2209286f810072abdc1899756a35004fb

Observation a316fe08-b859-461a-853f-a88b28f6f36c · outbound

This paper cites Places: A 10 million image database for scene recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Places: A 10 million image database for scene recognition

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.367186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:50:24.303188Z digest=sha256:384b649a32ac2117353881442a23c2de4322095f6fb4a67abfd14a22f3a7d639

Pith citing papers

No inbound Pith citation observations are available.