Pith. sign in

Paper Citation Record · LEDGER

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

As of 7 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.13364.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13364 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:50:24.303188Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact8
  • verified fuzzy35
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d59a1872-489d-4553-9de4-41191b6890f5 · outbound

This paper cites Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.223191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.223191Z digest=sha256:5a2d32fe5c31e9bf7afec7ea9488038cd4f6aac512da9c6f9193331f176b9966

Observation 58c6bae3-fdc2-4237-8b13-d67425b5116f · outbound

This paper cites Objects that sound.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Objects that sound

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.301171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.301171Z digest=sha256:0321091310b03c15188710f52bdc67ca054f538af05a7cceb1edd16504051fa3

Observation af1c03c2-eed2-4c63-8061-3e9de9304f1e · outbound

This paper cites 3d seman- tic parsing of large-scale indoor spaces.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d seman- tic parsing of large-scale indoor spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.441879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.441879Z digest=sha256:20baaf9efab89b02e207707a43c97c3bff7570ffa892203d181a8c29cd00f77b

Observation 39d3b6b5-846b-49fd-87e5-30dc0c0f116b · outbound

This paper cites MAE-AST: Masked Autoencoding Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.572549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.572549Z digest=sha256:30c9cc668407de9c9b8f5afffa848356b9c82bba934ee4c60284782c2fa8cf97

Observation a21ca1f0-fa77-4a38-b1f9-6cd59471f3d0 · outbound

This paper cites Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Data2vec: A general frame- work for self-supervised learning in speech, vision and lan- guage

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.702418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.702418Z digest=sha256:b9aac1ad943060cda994ca35ead7b126e284e716c0fe8a60463b8c67356761f4

Observation 2bea425a-ac9b-44e4-8ad7-d70b29a3bf8d · outbound

This paper cites Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Generative adversarial networks based on transformer encoder and convolution block for hyperspectral image classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.854860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.854860Z digest=sha256:96d75a8374d1d8536ec8efb2848667467add599471b00ce74ef30209aa4d66e9

Observation 3c461085-a472-455b-901b-a4c5fd5b9a5b · outbound

This paper cites HiP: Hierarchical Perceiver.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning HiP: Hierarchical Perceiver

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:14.988580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:14.988580Z digest=sha256:32c38305a8bd8120984ab635f3f8c0dbc492c35994a3f495c4d6f868c6dd14ba

Observation 56e36f43-8424-40ea-a6a1-99d718592339 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.091237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.091237Z digest=sha256:29c988b2ba696b3ff538199de7502799520427e8fdf8fc0cce8a1b944de86878

Observation 54ea8e34-ffa0-4245-9962-ef434d3f3ae7 · outbound

This paper cites DialogSum: A Real-Life Scenario Dialogue Summarization Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning DialogSum: A Real-Life Scenario Dialogue Summarization Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.207070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.207070Z digest=sha256:9a5447f496920518fb95c710a710a4a71ef964fae0665a036672cd3a19a7ff6c

Observation c5a18620-9c84-4893-8dc8-db6d1adb2641 · outbound

This paper cites Multi-Task Learning with Deep Neural Networks: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-Task Learning with Deep Neural Networks: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.356316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.356316Z digest=sha256:5979a203b632b514123d79ac381d75be66a1759958db1c9276802ebb1596f571

Observation 56770bc8-7652-4291-a236-385ac86ae989 · outbound

This paper cites One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.526503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:15.468308Z digest=sha256:790d76cc05ac5b2986fea9af388485849a95ebc457fd888f3387d96288e2292c

Observation 9f3f72e0-b581-414b-8320-fc49d0f25b45 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Imagenet: A large-scale hierarchical im- age database

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.576775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.576775Z digest=sha256:f0541f9656f1b42684ad2e6d2732dba437d1f9a9cd3f49d7d0e634e19c469f71

Observation fd8e2161-2dab-474d-aafc-305379b80e06 · outbound

This paper cites BERT: Pre-training of deep bidirectional trans- formers for language understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning BERT: Pre-training of deep bidirectional trans- formers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.745612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.745612Z digest=sha256:2f7cc4a7c2d0d72a18937a35c2c9770c83636071e2090eabcd8ba44e5249ab53

Observation 37f95f8a-075d-45ee-b344-173b91ad14b3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:15.911782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:15.911782Z digest=sha256:a23235ae15f1097e6bc42492ab8fbc211168f0e77b7a28e3d9dfd7c0af8b54f1

Observation 55f8eeca-e4a0-4882-b8a4-0aead2363c93 · outbound

This paper cites A generaliza- tion of transformer networks to graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A generaliza- tion of transformer networks to graphs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.035240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.035240Z digest=sha256:45bdbe95eeef185278fb8fd4414b1301d937288e66e26774d6c9f480e4140322

Observation dc85f918-01e8-42ac-9e37-65f9d313ef58 · outbound

This paper cites Efficiently identifying task group- ings for multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficiently identifying task group- ings for multi-task learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.197878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.197878Z digest=sha256:b23cafb13cc08819a477300b131855339108c4b694dccbe1d00cd7f9ecaac295

Observation dae79af5-b80a-4199-8116-20bd3ce75789 · outbound

This paper cites End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.324991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.324991Z digest=sha256:df3356210253144244f22e833a347c3ef4411fee9a215ee9423ed6a1bd7b32d5

Observation d9b07949-7982-4622-b307-34c03a5d24c0 · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audio set: An ontology and human- labeled dataset for audio events

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.490217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.490217Z digest=sha256:8fbfe7f6e141cc3c3c0dadae3a743383faf8db82a83309aac5a646a3a7c6f343

Observation 932aad78-8e47-4db1-aeef-9cf77236476d · outbound

This paper cites OmniMAE: Single Model Masked Pretraining on Images and Videos.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniMAE: Single Model Masked Pretraining on Images and Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:26.285243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:16.598120Z digest=sha256:b929a644dca2b2545a5910a8d22638abeeb555fbcbf790756f9990bf61203e06

Observation f09ffe82-9b83-4d02-aef9-ae2198364411 · outbound

This paper cites Omni- vore: A single model for many visual modalities.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Omni- vore: A single model for many visual modalities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.711982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.711982Z digest=sha256:9e789a6000582193f2dbc82f7ca875b60f31b08567e234e84e241bb424fd254e

Observation 475ea48a-1140-4d20-9ed5-3805e5e0fc1e · outbound

This paper cites SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:16.915486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:16.915486Z digest=sha256:5452eabac3f0d352a0fa9ad7c74d653140dc8913a82a5f1a5a686dc07ca43c55

Observation c5ee46f4-7d1b-482f-8cad-1e106d61c0a3 · outbound

This paper cites AST: Audio Spectrogram Transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning AST: Audio Spectrogram Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.039560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.039560Z digest=sha256:cd9c0f4499e7654b5b712445cb1b15df67794f15debcaf7c5b9fbf343eca63de

Observation e59a712b-21b1-4fb7-81b0-22fb95d4f43f · outbound

This paper cites Uavm: Towards unifying audio and visual models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uavm: Towards unifying audio and visual models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.202842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.202842Z digest=sha256:3ff2d9be6bb518e751b4ff1dca708ac428386e16da00c98468584fbd5c354ac1

Observation 1ef8eb6e-cdb5-41bb-9fbe-0567c56ac70e · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The” something something” video database for learning and evaluating visual common sense

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.322899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.322899Z digest=sha256:baa733edc25cad3712d0fca83e91b2d5f35d0b87694aefd2b8ede44490f9a80f

Observation 4684aabb-05a1-438c-b075-9b2859961c47 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.420096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.420096Z digest=sha256:f3bb5f4e29a821e7e791cf79847e043612e90212b320ca218c6f9ae21c49b78d

Observation 718fe2f5-fe32-4d5c-84fc-5ac5709e6bdf · outbound

This paper cites Dynamic task prioritization for multitask learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Dynamic task prioritization for multitask learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.507498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.507498Z digest=sha256:55443a91df9d931da0d391b21a6ed1c390819a25e1c42cf6ac1b5f9f9bee200d

Observation 87e16809-6d0e-4eda-b546-ff98c30ed626 · outbound

This paper cites MaskViT: Masked Visual Pre-Training for Video Prediction.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning MaskViT: Masked Visual Pre-Training for Video Prediction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.617134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.617134Z digest=sha256:32bbe925888fd7cf9b6b5a3723e9e9ded9dd4f1da38d7ee77067476ca6c33204

Observation e0b9e3d5-d3ef-4db0-92b5-c89935dbf407 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked autoencoders are scal- able vision learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.712917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.712917Z digest=sha256:7aaf7c19bf8ace8c0bc6d4ca0fa152f4cc12ae5417025c5d285f80aca646909e

Observation 0a79934a-741c-4e41-b197-0f059c137274 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Gaussian Error Linear Units (GELUs)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.826839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.826839Z digest=sha256:567488fa73aa88ab0bb1a25c93a6b4314819ea30336546afd9756a7d2eb5b864

Observation 1f89a59d-b2a3-4f03-bbb4-dad5dfd4add6 · outbound

This paper cites Spectral- former: Rethinking hyperspectral image classification with transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Spectral- former: Rethinking hyperspectral image classification with transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:17.932549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:17.932549Z digest=sha256:70030b611f7cfaa62e55a7549372d0df3c0afab989fa371834c4ee5cd3c56784

Observation 621186a5-ba8c-416d-8ff9-bd93cdde9a52 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Unit: Multimodal multitask learning with a unified transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.021377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.021377Z digest=sha256:b1e733637b4f9dd27158be4d1ad39351aa8510adeb5a8213b4006c2c147dc828

Observation 84660675-7b4b-4f5c-82f4-da2d19e23c2c · outbound

This paper cites OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OGB-LSC: A Large-Scale Challenge for Machine Learning on Graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.119583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.119583Z digest=sha256:e09210b91aa591e8956523bdac34bfde14a7b9a0952b992ca12e27309cec6a78

Observation 342a32dc-600d-4380-a433-3c3e8138eded · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.251387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.251387Z digest=sha256:02108de05bf86b243aa4aecc6a56c33422643c3f0700e7d86f4b5af6708e9260

Observation a5d0eb73-3086-4f59-aead-67f1a0556924 · outbound

This paper cites Perceiver: General perception with iterative attention.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Perceiver: General perception with iterative attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.379108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.379108Z digest=sha256:8899cd8942ecd1e0c784297b4c33eaf57794e1d752878c2151a1c10e7e470ae2

Observation 9d9f546e-fc24-4b2c-89fc-3a190f2f92bb · outbound

This paper cites A review of multimodal image matching: Methods and applications.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning A review of multimodal image matching: Methods and applications

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.514139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.514139Z digest=sha256:f00c3e5d870eb245fc4c1cf0b1c9ba5209eee4a658424ef29cc1a8724a350bbf

Observation 9cd1fe4c-4482-434f-b3e5-f327b37fc8f1 · outbound

This paper cites One Model To Learn Them All.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning One Model To Learn Them All

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.598502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.598502Z digest=sha256:8c774440cc89a59ea16b6c427be93766945af947f0ed9fe0009985fe8af671d2

Observation 2c9ea7e5-0131-4774-9b13-11500b736714 · outbound

This paper cites The Kinetics Human Action Video Dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The Kinetics Human Action Video Dataset

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.711195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.711195Z digest=sha256:c0e3c3c7163baf6c36996bd5f5337dff5fa51b4173a9a2b72be748e46af554eb

Observation 30d3f7e6-cc88-430e-8730-dc40862b5df6 · outbound

This paper cites Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.963938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:18.839182Z digest=sha256:d8db56a12f85f54db87013265981a2e8e4239c7331e7cc73fb1c643452a6e5fd

Observation 0b31472a-8eda-49b8-9476-252fb415548f · outbound

This paper cites Re- former: The efficient transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- former: The efficient transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:18.937345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:18.937345Z digest=sha256:9618df1cf0f4e78cb364e385a446414a5b4887e5f3b31600951aecae16e22712

Observation e7b145c0-c1a0-4ea9-b71c-f24d22135333 · outbound

This paper cites Hmdb: a large video database for human motion recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hmdb: a large video database for human motion recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.051884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.051884Z digest=sha256:0517862d01b801e2fc1433987bd9440e5d028ba440e0200668840a1c8c710790

Observation 1345fdee-e667-447e-8af6-ab88c4dd7d26 · outbound

This paper cites Modeling long-and short-term temporal patterns with deep neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modeling long-and short-term temporal patterns with deep neural networks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.166703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.166703Z digest=sha256:53502b38a7c27410eadb7ed61b1e92583c7066241457afd2fd43d1caaad4f50e

Observation 2b5de6d7-497d-47c5-9c91-c72f11041bf1 · outbound

This paper cites Stratified trans- former for 3d point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Stratified trans- former for 3d point cloud segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.290419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.290419Z digest=sha256:d4fb006b055f7cc0760ab9281e95005b4d43b38552f376845f19e062ebf9737e

Observation 6230888b-edf8-4947-8121-8823afc854e5 · outbound

This paper cites Regu- larization strategy for point cloud via rigidly mixed sample.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Regu- larization strategy for point cloud via rigidly mixed sample

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.311366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:19.428323Z digest=sha256:ac7b1cf554ae3884c57f3c724868ef269985b47c21fa6338098f511b5ad5ff9f

Observation 1d5913ce-11ba-4c0d-8c9e-28202a97a75c · outbound

This paper cites Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.222632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:19.541910Z digest=sha256:9309560a56b453ab44ca3114ee07d07328f6cdcb7e447eb06ebeae2f9d335ee1

Observation f1a474c0-6c44-45d3-9420-44286c7b83b3 · outbound

This paper cites UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UniFormerV2: Spatiotemporal Learning by Arming Image ViTs with Video UniFormer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.646948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.646948Z digest=sha256:daa81614e46edf46727e3a65dd3d3b8a6bc332a696cbd4dabf63ba21b78de573

Observation 941ae255-0a70-4627-8651-0196319e1ce4 · outbound

This paper cites Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:32.074110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:19.768028Z digest=sha256:238e6b7c65c711f90198b7e1ff4a986629e42ba10e0cff44746ecd5c85c311c7

Observation 9a8f0407-498c-4677-afa6-5800b0e5c46b · outbound

This paper cites Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:19.855755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:19.855755Z digest=sha256:2c09d8dcb2a90617475ab232a9e3fbd604f6aaba7182b507d925866f297ba6a3

Observation df00d131-ea42-45f1-bd8c-76ec203b2e1c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.892880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:19.953525Z digest=sha256:cd37340334359db400c8cc0d0db0aed1abe34bf4469fc3115feb247446be8910

Observation 56008d9f-43a8-4667-b043-2ab6eeb4e977 · outbound

This paper cites OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OPT: Omni-Perception Pre-Trainer for Cross-Modal Understanding and Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.761552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.051811Z digest=sha256:255b25bd8ec101b075e87d5bbf1993664d491192bcaf8e5935a7e9812b51a594

Observation 8f652af2-877d-4b29-90cc-df16bbcac204 · outbound

This paper cites Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Pyraformer: Low- complexity pyramidal attention for long-range time series modeling and forecasting

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.679062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.168295Z digest=sha256:de7911f482a08ca598b5b8f24b0ad973989b775ea1292e15e22d325baedf2426

Observation 08f78ae6-58ba-42ce-b471-ac7950235867 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.304795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.304795Z digest=sha256:e938d292b4468a4aa7e60cd706ebbfebc7eea9083816785d189b31a177a28b65

Observation 194acbc9-711a-48bb-9adc-f519ce36393a · outbound

This paper cites Moments in time dataset: one million videos for event understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Moments in time dataset: one million videos for event understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.539943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.398454Z digest=sha256:01051ef09fe30c7939383b9aad0c1811169f4611ae03ead98494b0fb6561b424

Observation 43b4d318-5029-4061-b742-6bc3205a9dad · outbound

This paper cites Person recognition system based on a combination of body images from visible light and thermal cameras.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Person recognition system based on a combination of body images from visible light and thermal cameras

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.423488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.525721Z digest=sha256:67b997818e27cf24fa055322aa2f745e7a3f82fc2dc70bf379ed949f5c2d964f

Observation 43fa27db-de94-4685-827b-56ae6cb30588 · outbound

This paper cites N-BEATS: Neural basis expansion analysis for interpretable time series forecasting.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning N-BEATS: Neural basis expansion analysis for interpretable time series forecasting

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.661670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.661670Z digest=sha256:13ecee2c4c2fdc92fcc2beea37aa81800efb4e6ec08fc1c061ac172a4182b8d3

Observation 60e8e3d8-5162-4883-a951-b6a1b9b0e488 · outbound

This paper cites Cats and dogs.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Cats and dogs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:20.767129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:20.767129Z digest=sha256:9fea98d46cf5705fffea35de06d9fd58a419661398a9683f582e2a3360b95d9b

Observation ed785505-9347-4868-91d0-3a3ae50cc0af · outbound

This paper cites Esc: Dataset for environmental sound clas- sification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Esc: Dataset for environmental sound clas- sification

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.311763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.883668Z digest=sha256:889edef4a20791ca647232e58ae009e146bdd0aec8cd0e6337de89b8f6d3464a

Observation 6b0d0f90-f56e-4ab7-8257-a0b3f156ec28 · outbound

This paper cites Re- thinking video vits: Sparse video tubes for joint image and video learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Re- thinking video vits: Sparse video tubes for joint image and video learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.184763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:20.990939Z digest=sha256:859e87cf07dc565c745a0ad13d40c157546b6f4fc344688af189afb0ce4567ce

Observation 8a5f51c5-9667-473c-9240-5e22df4fd717 · outbound

This paper cites OmniNet: A unified architecture for multi-modal multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniNet: A unified architecture for multi-modal multi-task learning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.121011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.121011Z digest=sha256:9de5fd209c5410b92f0c67c715d4a83b2b402084c7ba81dd340e7eef05252a08

Observation 960b3122-f46a-4655-8e8c-c6780eec6014 · outbound

This paper cites Point- net++: Deep hierarchical feature learning on point sets in a metric space.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- net++: Deep hierarchical feature learning on point sets in a metric space

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:31.040826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.235366Z digest=sha256:ef2e05871382aa3082876c8ee71da2a985d670cface60822e1296b3b3de26764

Observation 85a502b9-ace2-44b5-b02a-a1445bda232b · outbound

This paper cites Improving language understanding by gen- erative pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Improving language understanding by gen- erative pre-training

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:21.380262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:21.380262Z digest=sha256:f6411b8219d4fb58c077e21be36d124303fcbe6ffa96133af1af126173983941

Observation cebe27b6-d809-402e-aa63-1b51e2012ace · outbound

This paper cites Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Reliable tuberculosis de- tection using chest x-ray with deep learning, segmentation and visualization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.946155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.519423Z digest=sha256:a8ca2b1735267489314f93db57f9d6444c8027238495087251df45feddbb4ac0

Observation 5f88d135-92c1-4203-823c-dee0e0d6a492 · outbound

This paper cites Zorro: the masked multimodal transformer.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Zorro: the masked multimodal transformer

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.551900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.672452Z digest=sha256:ab0837cf8c6797c3c13fdbccb863bc00b1cdec9d79d37f7c41eb3fa04a17205f

Observation 3f24eccf-3cc9-461e-a080-e9ab0c57190f · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Indoor segmentation and support inference from rgbd images

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.817984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.776607Z digest=sha256:079fb23096ff0a430374e1aa3665f6f7fbca7a1d145d3e8cbc2bfb4a34f11ab3

Observation 21993166-f068-41a5-84ad-b57e3fca20bc · outbound

This paper cites Mpnet: Masked and permuted pre-training for lan- guage understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Mpnet: Masked and permuted pre-training for lan- guage understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.698263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.861378Z digest=sha256:691417e6727b53e85a9dc3aa3c2ecb7f868491e4203be9387be3ab3e5ae76615

Observation 9a540032-3568-429f-af77-4e6fe5f3d6a0 · outbound

This paper cites Sun rgb-d: A rgb-d scene understanding benchmark suite.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Sun rgb-d: A rgb-d scene understanding benchmark suite

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.580490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:21.997522Z digest=sha256:91d97790df3091da2500e399b13fd1975130ff307fafa2b411cbbd6d510ed62a

Observation f8a1fc2e-e727-48e9-b9d7-b1eae07b61c2 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.073655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.073655Z digest=sha256:a3204e46418d72738916d34fc1f790a40023352651e529a5b3f3c06be4e276e8

Observation 3fed73df-e32b-4259-aa76-b7be8cddd991 · outbound

This paper cites OmniVec: Learning robust representations with cross modal sharing.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning OmniVec: Learning robust representations with cross modal sharing

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.147490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.147490Z digest=sha256:fc4a2e4d5bb042bb4ea3afd5c14db5ee11f613654aeb8082bf578f487210c808

Observation 8049192c-b68b-414c-a96b-c586b6cd28ba · outbound

This paper cites Hierarchical multi-task learning via task affin- ity groupings.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Hierarchical multi-task learning via task affin- ity groupings

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.455919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.210014Z digest=sha256:34675e531abc2448756abfc85f7811681b5cc6009651d302d713c6be68e762d4

Observation f3edae28-32b7-4164-8523-4d5b73918175 · outbound

This paper cites Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Benchmarking Robustness of 3D Point Cloud Recognition Against Common Corruptions

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.270903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.270903Z digest=sha256:ffedfaa8b41899b8e06754fe401f8e8f3cb78c418d6437eb9b85b347d5d29d04

Observation edf6e8df-359f-4bad-8d02-a2c8f3fbc02c · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.333424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.333424Z digest=sha256:dbff9e0b17d97f2d335dd0b0e80d68aab07163e759a1e82545ad8f592d573956

Observation c367eb61-daef-4882-96a4-71cc92cb0c7f · outbound

This paper cites Contrastive boundary learning for point cloud segmentation.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Contrastive boundary learning for point cloud segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.308893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.389608Z digest=sha256:2b12e1bb161db247ef6b2d4298e16211f8d23516e3fe39ad31bc325abe2bfd79

Observation 3271b91b-af12-4545-bb64-f5f2e23434bb · outbound

This paper cites Small sample hyper- spectral image classification based on the random patches network and recursive filtering.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Small sample hyper- spectral image classification based on the random patches network and recursive filtering

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:30.150417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.462428Z digest=sha256:94b3869b15ec592b45cc3ca02af36dad5f16c462788e69a6a06be9e1b4beb81f

Observation f1b25de1-73cf-4f19-b089-6357805ff965 · outbound

This paper cites Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Revisiting point cloud classification: A new benchmark dataset and classification model on real-world data

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.959279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.537912Z digest=sha256:367a6ee3e2dc03d59bd01ccb21eb50709485c1feec0b7e33799d57d75c64f7c1

Observation 327fdeff-ef5c-489b-9c7f-53dee17d507c · outbound

This paper cites The inaturalist species classification and detection dataset.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning The inaturalist species classification and detection dataset

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.779706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.632705Z digest=sha256:8004c1618a3aa15943b813c8c84326f16397d38d679942d23fdccd4024fedb52

Observation b6988d14-15b5-46a2-98e9-f821a035c5a3 · outbound

This paper cites Attention is all you need.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Attention is all you need

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.627959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.716235Z digest=sha256:6823021f8e21ca0ed163479551879dbfcd6ee8ce20825b039647f2ae8a8dd85e

Observation 11b94fd8-c192-4958-8fbc-4900b5be83e0 · outbound

This paper cites Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.467629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.775789Z digest=sha256:d9c0e3a1dfd9c6381ebe518c412aa12d6aeebc3c9b3a5f30e945879644b73d8b

Observation 90af2435-5042-4201-ad4c-2e1abd6c5506 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.840001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.840001Z digest=sha256:08595a391aa89a976040a6d1b91aea55567911bf49ad3f1978859f2d9d383c9c

Observation f9b85ae5-edfd-4db3-9b20-6bd5903666be · outbound

This paper cites Masked feature pre- diction for self-supervised visual pre-training.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Masked feature pre- diction for self-supervised visual pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.260075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.894679Z digest=sha256:d7f12ad71b78212b4958bc0815cea14c08fc7b2daaad7c4a8ba1e507c9d6f0bc

Observation 349ba8bd-f659-4ebf-8a76-276997f08069 · outbound

This paper cites Syn- cretic modality collaborative learning for visible infrared person re-identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Syn- cretic modality collaborative learning for visible infrared person re-identification

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:29.084751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.941403Z digest=sha256:b659399d267cddfe71df1d66bd23bb64024643bd03fc13ebe037f52136f5cd9a

Observation 0d42cd48-9c56-4dc2-846e-3514925e8c9a · outbound

This paper cites Controllable Abstractive Dialogue Summarization with Sketch Supervision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Controllable Abstractive Dialogue Summarization with Sketch Supervision

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.321727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:22.994077Z digest=sha256:2b5fa1b1b1706048d49b7375511a07ce275be2060c70703b73a27322d34ef049

Observation cecf0c2d-d737-4047-a8ac-88955a12dafb · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Tinyvit: Fast pretraining distillation for small vision transformers

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.942665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.044939Z digest=sha256:aecdc82fa31bf446b1edc7667d35de1afa5d67dc41b4596e9be00fb1332a85c7

Observation 825288bf-78c8-4094-8cf0-1eaa9b17c048 · outbound

This paper cites Point transformer v2: Grouped vector at- tention and partition-based pooling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point transformer v2: Grouped vector at- tention and partition-based pooling

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.793382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.101262Z digest=sha256:4c55fafdcffbdb2554c5c6b46cf2ae5a4a0d5a5c9b58fba400a59007cd36c321

Observation 4c15fc0b-37a7-42c0-a0ca-8f888624e3ab · outbound

This paper cites 3d shapenets: A deep representation for volumetric shapes.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 3d shapenets: A deep representation for volumetric shapes

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.654139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.180431Z digest=sha256:0de5285e9900738014a9f6ff2f8fcb438685b7268c7bf2241dde0f92510642b6

Observation d4f48735-2087-4877-8af9-04073b46d0bd · outbound

This paper cites Audiovisual SlowFast Networks for Video Recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Audiovisual SlowFast Networks for Video Recognition

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.240129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.240129Z digest=sha256:f56428c462ad2231832de512cb0225f3631fc867718a2314cafb3813233d7d47

Observation bfb374f0-7601-4427-96df-5ee5faf71b0e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Msr-vtt: A large video description dataset for bridging video and language

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.456028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.286212Z digest=sha256:b23d29a00d2ae43dffec6cc87ae2f2fa2eb919caae49e382e91ecacf626a6036

Observation e0c2cd51-f150-447c-a73f-a643a7c81994 · outbound

This paper cites Multimodal Learning with Transformers: A Survey.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multimodal Learning with Transformers: A Survey

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:25.148397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.399732Z digest=sha256:72abe1da89f604fb99d572fd84140d2a9f955beddf759ac81f11125843c96719

Observation cfa2cd38-10c1-42a4-8f1e-7a24a37a3490 · outbound

This paper cites Multi-modal masked pre-training for monocular panoramic depth completion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Multi-modal masked pre-training for monocular panoramic depth completion

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.330526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.499703Z digest=sha256:8154734f550c626c62ab9509f834b8bc33cf298d4be739379b72556fa45189d8

Observation f955fc72-ae65-4e44-bf53-cb3a6af81b53 · outbound

This paper cites Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Swin3D: A Pretrained Transformer Backbone for 3D Indoor Scene Understanding

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.580652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.580652Z digest=sha256:22c7b520f2d0d5f85f822a97847f36aad3b53d4e7ce0d03766b36c6ec99ca4d3

Observation b008ece1-7827-4814-802a-93820de787d8 · outbound

This paper cites XLNet: Generalized Autoregressive Pretraining for Language Understanding.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning XLNet: Generalized Autoregressive Pretraining for Language Understanding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.682264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.682264Z digest=sha256:b024594cd49fad25be9e66edbcbf09b24d9bb86999864c1b487bbb8e6e6ac77c

Observation 25e42e8f-0a25-4e26-a6fa-ea9b67d18dbd · outbound

This paper cites Deep Learning for Person Re-identification: A Survey and Outlook.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Deep Learning for Person Re-identification: A Survey and Outlook

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:50:24.969127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.740334Z digest=sha256:5b4665c061776eb32507385866955489d6d043cbbdbe3ee961f8020d92ac67c1

Observation 240ad846-4d80-4e93-a896-f29ed0221efb · outbound

This paper cites Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Do transformers really perform badly for graph representation? In Thirty-Fifth Conference on Neural Information Process- ing Systems, 2021

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.189988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.790277Z digest=sha256:b25be5c0e86d953d9077f629de8c0fb20a8ebc0540822a37a578df5e25286d40

Observation 7aceffcd-fbf6-4549-99b2-da442e37cdde · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:23.861350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:23.861350Z digest=sha256:9ed18218f45e345c4f1647a9ab24cbe0fca9fd733d848babf5920a26d6191038

Observation e1e514a6-2fda-4a17-8cac-dda7c93bbe04 · outbound

This paper cites Metaformer is actually what you need for vision.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Metaformer is actually what you need for vision

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:28.047216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.921175Z digest=sha256:d8c5560e7eaa36d45b895ecbd025a93dfa928a90be30951f184d9f2734ea3859

Observation b4baed03-004d-4915-9ad5-4c3216d2b102 · outbound

This paper cites Point-bert: Pre-training 3d point cloud transformers with masked point modeling.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point-bert: Pre-training 3d point cloud transformers with masked point modeling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.928246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:23.991722Z digest=sha256:0f56fdd95079e6829471d9ca8cb0bd6d765fcf9cb7e937b5d6aefcb128984b57

Observation 4ffff764-3d5f-4d5f-8564-951c54b457e6 · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.066566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.066566Z digest=sha256:ed6a85bf98c69491195a988d0f6845654dba87f324f4b45efea8d34410d35840

Observation b2ebe9a9-9854-4749-93c2-f4b57288dbc8 · outbound

This paper cites Point- cutmix: Regularization strategy for point cloud classifica- tion.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Point- cutmix: Regularization strategy for point cloud classifica- tion

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.750612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:24.114385Z digest=sha256:bfbf2dce541bad89229ddf9c222b4a266c98e1d01dbc67ba658b355ba0695d85

Observation 521fa700-e1e6-4b00-8163-3c2a9f01e95c · outbound

This paper cites An overview of multi-task learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning An overview of multi-task learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.594862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:24.164336Z digest=sha256:771e0f2e005459c620dd33f3599ed5a729ec81b47eb190d805843ba1cc6114ca

Observation 3953969f-63e7-4d49-91f8-48982b077969 · outbound

This paper cites Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Modality synergy complement learning with cascaded aggregation for visible-infrared person re- identification

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.512561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:24.215744Z digest=sha256:ebc244a9b47ce7ec200aa21db355ed8860e31ab45ab46e57fd68343b111dd625

Observation af7aa99d-b892-47ac-a506-e0154c640fe8 · outbound

This paper cites Meta-Transformer: A Unified Framework for Multimodal Learning.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Meta-Transformer: A Unified Framework for Multimodal Learning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:24.267156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:24.267156Z digest=sha256:42a25a0835c622532b5cb3562d9956e2209286f810072abdc1899756a35004fb

Observation a316fe08-b859-461a-853f-a88b28f6f36c · outbound

This paper cites Places: A 10 million image database for scene recognition.

OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning Places: A 10 million image database for scene recognition

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:50:27.367186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:50:24.303188Z digest=sha256:c3a3defab75368e95a341cd19c2d1c8412ea44dc55c6735b426c7a2e94075deb

Pith citing papers

No inbound Pith citation observations are available.