Pith. sign in

super hub Mixed citations

PoseNet: A convolutional network for real-time 6-dof camera relocalization

Mixed citation behavior. Most common role is background (67%).

123 Pith papers citing it
Background 67% of classified citations

hub tools

citation-role summary

background 14 method 6 dataset 4

citation-polarity summary

claims ledger

  • method single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu
  • dataset (C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative
  • background Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, "Recognizing text with perspective distortion in natural scenes," in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, "Toward understanding wordart: Corner-guided transformer for scene text recognition," in Pr
  • method . Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat
  • method To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss
  • background Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model

authors

co-cited works

representative citing papers

MATCH: Flow Matching for Multi-View Anomaly Detection

cs.CV · 2026-06-23 · unverdicted · novelty 7.0

MATCH is the first flow matching method for multi-view anomaly detection, reporting SOTA results on Real-IAD and the first comprehensive evaluation on MANTA-Tiny while enabling real-time use by omitting the divergence term.

Adaptive Volumetric Mechanical Property Fields Invariant to Resolution

cs.CV · 2026-06-16 · unverdicted · novelty 7.0

AdaVoMP predicts accurate dense spatially-varying Young's modulus, Poisson's ratio and density for 3D objects using an adaptive sparse voxel structure generated by a sparse transformer encoder-decoder at 16^3 higher resolution than prior fixed-voxel methods.

Brain-IT-VQA: From Brain Signals to Answers

cs.CV · 2026-05-28 · unverdicted · novelty 7.0

Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.

citing papers explorer

Showing 50 of 123 citing papers.