Pith. sign in

Paper Citation Record · LEDGER

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.13056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13056 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:23.947375Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact6
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78de6edd-6ba5-4416-a0f8-b5a9de3591b8 · outbound

This paper cites write newline.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.782447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.782447Z digest=sha256:966f86da6980ce05a6ef76d403189c05e46098bb7601ac2509391a2f26d5e752

Observation c2614720-692f-476a-8d3f-41d5899fd900 · outbound

This paper cites Arteta, V.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Arteta, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.052902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.787353Z digest=sha256:52a39012c590488ea5241636975393a07f7bfa0fbc782669c533cf8bde1db964

Observation 9336d576-683b-4a9e-a2a2-e54644fc83db · outbound

This paper cites A spatio-temporal attentive network for video-based crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark A spatio-temporal attentive network for video-based crowd counting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.791186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.791186Z digest=sha256:8543b5a91ffa396b78df55eda03aabe4577f3720cb1073cdab75832f13da7cad

Observation 52a8ef27-0ddc-4792-b827-281bb1618396 · outbound

This paper cites MultiMAE : Multi-modal multi-task masked autoencoders.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark MultiMAE : Multi-modal multi-task masked autoencoders

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.042712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.795090Z digest=sha256:d9157150bae0f71d32591a0d24198cda7e7dc5eb7a805e1f68e905ad6b62442c

Observation 78f9d0d9-06a9-4221-9460-8fd303c44804 · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:57:25.032456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.798655Z digest=sha256:5aaae40b6c8aa0eac98ee77f2f5196f46e23921ab0493a5d0067fa44b0c56fa5

Observation b1d05056-85ce-473d-ae0e-94583a8d7f67 · outbound

This paper cites BE it: BERT pre-training of image transformers.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark BE it: BERT pre-training of image transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.802121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.802121Z digest=sha256:6f5e41f275e134ce2c290a32b7045395dee9fdfc69cda6f18052aa58b186500e

Observation ab34073d-7ad5-4ec8-85b8-cc4932a56fcc · outbound

This paper cites Generative pretraining from pixels.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Generative pretraining from pixels

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.016749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.805406Z digest=sha256:c098fc14fb9380c386bd230cca0de412f9d5362d850e840a00f439ecf01a09e9

Observation 7559e6d6-2689-4933-babd-80b49dffba1c · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 8

Resolution
malformed identifier
no resolver link, observed 2026-08-12T16:57:23.809064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.809064Z digest=sha256:51b6492e8a428e2c7a6f37a5393edc2f83a78aea66c570aa22637607ecff8906

Observation 440ed128-4068-4d88-aa34-4a7999eff820 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.812413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.812413Z digest=sha256:ab081eacea689ce72b8450be126cfa0dad3c6525f73dd7008d5c1d4428f678e1

Observation 9685b215-d4aa-48d5-a5d6-15a924bd2266 · outbound

This paper cites Redesigning multi-scale neural network for crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Redesigning multi-scale neural network for crowd counting

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:25.007105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.815966Z digest=sha256:973cf2a18db775b730ea9d4be79eab6ffb482a43b7be922eae908e01b1609a94

Observation 626e37d2-a91a-4636-9903-0addb52abe50 · outbound

This paper cites Locality-constrained Spatial Transformer Network for Video Crowd Counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Locality-constrained Spatial Transformer Network for Video Crowd Counting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.819308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.819308Z digest=sha256:e7c9e4971bd15c49c36ee85fad38b13a4d40a281019b2da87bf69826c5690082

Observation 371c6409-ac7b-4fe8-9441-4d53923a1442 · outbound

This paper cites Multi-level feature fusion based locality-constrained spatial transformer network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Multi-level feature fusion based locality-constrained spatial transformer network for video crowd counting

Reference 12

Resolution
verified exact
doi, observed 2026-08-12T16:57:24.018963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.823305Z digest=sha256:cff2849bb5fecf5de2c14882670aa1d9e3bec66a20da84266bf9d7731ddff4cf

Observation d4e7bb18-c7e9-4865-be09-3b03414112da · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Masked Autoencoders Are Scalable Vision Learners

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.826976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.826976Z digest=sha256:567de428a056854d3a1216da1f292004bb3f0e7d00dbbba542b8ac42f9d4608e

Observation dd01dfb4-1de7-4279-a4ee-3027a3802743 · outbound

This paper cites Video-based crowd counting using a multi-scale optical flow pyramid network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video-based crowd counting using a multi-scale optical flow pyramid network

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.996512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.830574Z digest=sha256:7904ab5d3918493d67bc2ab9bd962c9cd965b33e2ac308f3675f0b4d1c18443e

Observation 9db75f98-6b03-47ea-a154-dc0b13e0e245 · outbound

This paper cites Frame-recurrent video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Frame-recurrent video crowd counting

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.833699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.833699Z digest=sha256:1c2efbe06c636995e2ba2ac79a9166135863099db666ac040402e8fd1460f7fa

Observation d57eee47-f21d-4ffc-81d4-41e1418f6ed8 · outbound

This paper cites Clip-count: Towards text-guided zero-shot object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Clip-count: Towards text-guided zero-shot object counting

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.837157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.837157Z digest=sha256:e744790ebe7a7881b8d4406b414d1f4c9c4c18651aa4e05f2c795f61292c1627

Observation b07f420c-f39f-4c4b-bbdc-2760dec31ff1 · outbound

This paper cites Vlcounter: Text-aware visual representation for zero-shot object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Vlcounter: Text-aware visual representation for zero-shot object counting

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.985644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.840443Z digest=sha256:f6ec2b80c0f48d3d36f621a947698ce7b1723dea484383c4dd5f2692539327dd

Observation 44abdf7a-f553-4568-9b82-d0e3f0096172 · outbound

This paper cites Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video crowd localization with multifocus gaussian neighborhood attention and a large-scale benchmark

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T16:57:24.609996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.843597Z digest=sha256:ab06f8ca277ca33634e092c73acba676459ed8f2d92d60cd8dabcf34423fe9c3

Observation 43b5dbcc-1851-4bbe-b4aa-145e236db61d · outbound

This paper cites Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.846966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.846966Z digest=sha256:798c9971f8db6c5c0d99266b91988300f73bf604e16a6988ce657f771545e6cb

Observation 7e7bb7f6-0872-48bc-9c94-cd381b828bad · outbound

This paper cites Transcrowd: weakly-supervised crowd counting with transformers.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Transcrowd: weakly-supervised crowd counting with transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.975681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.853888Z digest=sha256:964b9c804870bcf40e24938affbcb8cc36774b47905a0c5ba919851b3339d500

Observation f8569ae5-1d0d-48be-a570-e427f6b725be · outbound

This paper cites Boosting crowd counting via multifaceted attention.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Boosting crowd counting via multifaceted attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.965341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.857463Z digest=sha256:8afdc41a696bd24c19796ad52fdb87473e234ad1886bc360d2e193e28c5b9ddd

Observation cd0653db-3a7d-43e6-9cdd-0a97e685dace · outbound

This paper cites Gramformer: Learning Crowd Counting via Graph-Modulated Transformer.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Gramformer: Learning Crowd Counting via Graph-Modulated Transformer

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:57:24.474037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.860634Z digest=sha256:7eddf822c0cbbb7249b865b2c814a218d8cd661ec9aa58d019671ca183a54c67

Observation d159f02c-9c61-4f5d-b4a5-f0a345eddb72 · outbound

This paper cites Point-query quadtree for crowd counting, localization, and more.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Point-query quadtree for crowd counting, localization, and more

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.955272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.863958Z digest=sha256:e751d162a756b7c628bb8cd657ce21452b425efb2a5eaa5ef7d712ae7616a0c0

Observation 7a0b16c9-afd0-4c71-9896-e1652db2a2ee · outbound

This paper cites Context-aware crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Context-aware crowd counting

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.945253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.867202Z digest=sha256:3910b066cbed9e913bcf34443339618c84af83d8bbbf31a3872ee5b690ef0700

Observation 664e97d6-0f59-46a2-990b-63d94dcb3050 · outbound

This paper cites Estimating people flows to better count them in crowded scenes.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Estimating people flows to better count them in crowded scenes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.934936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.870428Z digest=sha256:3761bcfeb9cb33991c82520879e9711086e0a20760079b317ace0275abf3165f

Observation aa9eff38-6cd3-4e1a-aec1-b52ec02577b8 · outbound

This paper cites From semi-supervised to transfer counting of crowds.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark From semi-supervised to transfer counting of crowds

Reference 26

Resolution
verified exact
doi, observed 2026-08-12T16:57:24.008980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.873558Z digest=sha256:7e340fbed06ea77d9064a382b43662764fcfcf1b97605ac61aff724296b5e7ef

Observation 166a18b5-fc40-4574-9de1-af9f9e8641b0 · outbound

This paper cites Bayesian loss for crowd count estimation with point supervision.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Bayesian loss for crowd count estimation with point supervision

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.924155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.876933Z digest=sha256:5b9dc93d88174d26bedf92418455fd22c1e7d2c7f3382d963276d30a923e4c5b

Observation ecad169a-ae56-4f8c-bd91-ec3133ddc848 · outbound

This paper cites Phnet: Parasite-host network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Phnet: Parasite-host network for video crowd counting

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.880503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.880503Z digest=sha256:46fbc540d629cb7ed2a6231d64e28ca51c223f287a08bc78d7ca3446f69c8943

Observation e3544976-f0bd-4017-a5b4-3356690211c3 · outbound

This paper cites Rudin, Stanley Osher, and Emad Fatemi.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Rudin, Stanley Osher, and Emad Fatemi

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.883550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.883550Z digest=sha256:4a8afea1b14996baa02f4b7e1de83ba764a34cfb79ac0d25cf01389ee2ebd198

Observation 5bb89ac6-e73a-40be-8d23-01b8fa385da9 · outbound

This paper cites Convolutional lstm network: a machine learning approach for precipitation nowcasting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Convolutional lstm network: a machine learning approach for precipitation nowcasting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.913583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.886796Z digest=sha256:94b6a30a41d7351bcbfdf7cc9455b97e4f023d063f0a54031f987be3e361f1ea

Observation 91c7fe4f-fb45-4a2d-be96-89aa511175b8 · outbound

This paper cites Crowd counting in the frequency domain.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Crowd counting in the frequency domain

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.903338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.889842Z digest=sha256:35cdd6a7c8123a5417fbf1aa2c0559f47378a2648c36c1cba8f739165bbf697e

Observation 42c4479e-a51b-402e-9b9b-5cd1d2c1b31a · outbound

This paper cites PWC-Net : CNNs for optical flow using pyramid, warping, and cost volume.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark PWC-Net : CNNs for optical flow using pyramid, warping, and cost volume

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.893026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.893078Z digest=sha256:d06d8b854ed86f96fb585e3d16f50c49a4bc64237e1a907204153ba2d1d262ec

Observation d51332c6-d2ce-48cf-a815-15be8225e664 · outbound

This paper cites CCTrans: Simplifying and Improving Crowd Counting with Transformer.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark CCTrans: Simplifying and Improving Crowd Counting with Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.896271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.896271Z digest=sha256:7e362e39ec0a8a661e82c421f805f6bd110a22d8343dbcefcaa737c99e8fe4fc

Observation 3e5a2c63-1336-4c00-9265-d6ff21e9ba3b · outbound

This paper cites Video MAE : Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Video MAE : Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.882248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.899774Z digest=sha256:a8f659492f377a0fb35a002f05e89d568fd37ee9e107b61d006a5cd03dff5ab4

Observation 8c05006d-6a9c-4f66-a70d-d29040192035 · outbound

This paper cites Attention is all you need.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Attention is all you need

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.903168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.903168Z digest=sha256:77fd20e88015625b006f15e31e45fd7208ccc6944df87e94b8ea63f044774db9

Observation 0b6566b4-245c-4135-9a4b-6423cc651fa1 · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Extracting and composing robust features with denoising autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.906726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.906726Z digest=sha256:f1bbd0aa695ec2bbb56e28c17f52f935ddceee665d013391c98325e386f3a70a

Observation c431bf60-4e15-4b29-accb-0a01d9669159 · outbound

This paper cites Bird-count: a multi-modality benchmark and system for bird population counting in the wild.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Bird-count: a multi-modality benchmark and system for bird population counting in the wild

Reference 37

Resolution
verified exact
doi, observed 2026-08-12T16:57:23.998950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.910115Z digest=sha256:05cfeb9f903a0b53da610276164c5b42d41f1f0ec95a8247d7e5ecde40028584

Observation a2bd0bf0-a4b8-4aa7-a698-fce5016d0455 · outbound

This paper cites Fast video crowd counting with a temporal aware network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Fast video crowd counting with a temporal aware network

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.866031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.913415Z digest=sha256:2b5bb00d491f7aa946696fc3903f01982c550e455c5389b3abd2189536542129

Observation c3b0ec67-e2ae-441f-b134-c66cba37238b · outbound

This paper cites Spatial-temporal graph network for video crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Spatial-temporal graph network for video crowd counting

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.916569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.916569Z digest=sha256:2cb63d440f800c118e57fbf71120b27ba73b903188a6abb20ce8404188ed4422

Observation 8c06c77d-2b8a-43bc-a1d8-7d2e5070bc3a · outbound

This paper cites Spatiotemporal modeling for crowd counting in videos.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Spatiotemporal modeling for crowd counting in videos

Reference 40

Resolution
verified exact
doi, observed 2026-08-12T16:57:23.985587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.919832Z digest=sha256:6fbed1b817348efe638fd314624e8c42239687a8549546e981b1e452171a2800

Observation 470da1b4-0df2-4d25-9f9c-e3fa5f45fd4f · outbound

This paper cites Reverse perspective network for perspective-aware object counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Reverse perspective network for perspective-aware object counting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.855306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.923135Z digest=sha256:eddfeb272ac1c9554205d35cd40e8fc4869bbff84110ceb1284b96801e819925

Observation 7a3fbda3-32c5-486a-9fdc-48b6ff5a7707 · outbound

This paper cites Single-image crowd counting via multi-column convolutional neural network.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Single-image crowd counting via multi-column convolutional neural network

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:57:24.844758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.926246Z digest=sha256:68cf3ee0585735f784178cf942a8ca74b325d86da641ad0a9b864e08d66987a2

Observation c75526b9-4f58-4993-b322-03008a703bc4 · outbound

This paper cites Locality-aware crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Locality-aware crowd counting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.929315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.929315Z digest=sha256:a60e93e47b8a400256faa000c9542f49fdc4bd2d5a672d4f325ffdf306cb33d4

Observation b7e0ffea-fd1d-4f8a-bf2b-0649ab05a780 · outbound

This paper cites Graph regularized flow attention network for video animal counting from drones.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Graph regularized flow attention network for video animal counting from drones

Reference 44

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T16:57:24.135086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.932906Z digest=sha256:82daf42b383c765a357e46d90149f65e1c42aa23fe3cc51cf0d21dd54a8785c8

Observation f90697e8-e4f4-4dab-9989-1d5d94f0adc5 · outbound

This paper cites Enhanced 3D convolutional networks for crowd counting.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Enhanced 3D convolutional networks for crowd counting

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:57:24.039410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-12T16:57:23.936167Z digest=sha256:2798310ee56e60b77a3a85667b7768cb1266f382e11807b8b5765304ae0aaf9a

Observation 7d9bc860-b41b-481b-bf69-a368c78472d8 · outbound

This paper cites @esa (Ref.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.939779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.939779Z digest=sha256:9c3c628513ca21f3942b2c5708a29a27ad5935a5df238d53d68fc6de13a728bb

Observation 2c6a71cc-5246-406f-9fea-608bfe245980 · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.943688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.943688Z digest=sha256:8ae0ab91ee809f2f46d2cc4d7bdc1f1e434a47a7fd2c5e79bd8ca663de0ec13f

Observation 0fe7fb19-91b2-479c-842f-0ee17427ea7c · outbound

This paper cites an unresolved cited work.

Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:23.947375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:57:23.947375Z digest=sha256:6da3aa2d29a2c480e6e0e6058ae74db657f4e1435140bfb31d59371a56435095

Pith citing papers

No inbound Pith citation observations are available.