{"work":{"id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","openalex_id":null,"doi":"10.1109/iccv","arxiv_id":null,"raw_key":null,"title":"Detection of Interest Points Using Symmetry","authors":null,"authors_text":"Reisfeld, D","year":1990,"venue":null,"abstract":null,"external_url":"https://doi.org/10.1109/iccv","cited_by_count":null,"metadata_source":"doi_reference","metadata_fetched_at":"2026-07-10T08:26:58.784811+00:00","pith_arxiv_id":null,"created_at":"2026-05-08T18:33:59.247625+00:00","updated_at":"2026-07-10T08:26:58.784811+00:00","title_quality_ok":true,"display_title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization","render_title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization"},"hub":{"state":{"work_id":"135418b1-cafd-49fd-803d-1ca6433d4b1b","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":123,"external_cited_by_count":null,"distinct_field_count":20,"first_pith_cited_at":"2017-08-22T17:31:54+00:00","last_pith_cited_at":"2026-07-09T17:59:58+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T22:49:20.477563+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":14},{"context_role":"method","n":6},{"context_role":"dataset","n":4}],"polarity_counts":[{"context_polarity":"background","n":16},{"context_polarity":"use_method","n":6},{"context_polarity":"support","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization","claims":[{"claim_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative ","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, \"Recognizing text with perspective distortion in natural scenes,\" in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, \"Toward understanding wordart: Corner-guided transformer for scene text recognition,\" in Pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks PoseNet: A convolutional network for real-time 6-dof camera relocalization because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (14 contexts).","role_counts":[{"n":14,"context_role":"background"},{"n":6,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-28T14:57:47.223142+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"ad0a11de-6175-4339-a8fa-1f60baf1c99e","orcid":null,"display_name":"doi: 10"}]},"error":null,"updated_at":"2026-06-28T14:57:47.219662+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-22T21:44:12.099634+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":13},{"title":"Adabins: Depth estimation using adap- tive bins","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":12},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":12},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":11},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":10},{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":10},{"title":"Poly- max: General dense prediction with mask transformer","work_id":"498f8786-e863-4217-b210-bf4fe976c779","shared_citers":8},{"title":"URLhttps://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":8},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"In: 2021 IEEE/CVF In- ternational Conference on Computer Vision (ICCV)","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":6},{"title":"doi: 10.18653/v1/ 2024.findings-acl.586","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":5},{"title":"EW Dijkstra","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":5},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":4},{"title":"BERT : Pre-training of deep bidirectional transformers for language understanding","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":4},{"title":"Conditional prompt learning for vision- language models","work_id":"025819dc-724a-4ff8-ba0a-0ba72c046d8c","shared_citers":4},{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":4},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":4},{"title":"","work_id":"ad0a8ee9-e814-486c-a25b-40126e136b0f","shared_citers":3},{"title":"2023.Parallel Programming for Multicore and Cluster Systems(3 ed.)","work_id":"cf4c4e77-acaa-46b4-b066-ddf045165d05","shared_citers":3},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":3},{"title":"GeoNum: Bridging numerical continuity and language semantics via geometric embedding","work_id":"9f349f1f-0e39-446f-8ac6-694a06c25de5","shared_citers":3}],"time_series":[{"n":1,"year":2017},{"n":1,"year":2023},{"n":1,"year":2024},{"n":15,"year":2025},{"n":48,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"LiBrA-Net: Lie-Algebraic Bilateral Affine Fields for Real-Time 4K Video Dehazing","primary_cat":"cs.CV","context_text":"at least linearly with pixel count, capping practical training at 1080p. Native 4K video dehazing therefore remains at the unattended intersection of these two lines: the UHD branch offers spatial scalability without temporal coherence, and the video branch offers temporal coherence without resolution scalability. 2.2 Bilateral Grids and Locally Affine Color Transforms The bilateral filter [34] and the guided filter [ 15] introduced edge-aware smoothing; the bilateral grid [4] cast the same operation as a 3D splat-blur-slice pipeline, and Bilateral Guided Upsampling [5] showed that locally affine transforms fitted in this space can approximate a wide class of image operators. HDRNet [ 11] made the fitting step learnable via differentiable slicing, and subsequent","citing_arxiv_id":"2605.11508"},{"n":1,"role":"dataset","polarity":"background","paper_title":"Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception","primary_cat":"cs.CV","context_text":"None captures authentic first-person spatial narratives of social media. Urban perception datasets including Place Pulse 2.0 [ 9], MMS-VPR [30], and UrbanFeel [ 13] focus on exterior street-level imagery, provide no textual modality, and lack multi-task annotation. Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated billion-scale vision-language","citing_arxiv_id":"2605.09936"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images","primary_cat":"cs.CV","context_text":"1), we construct and releaseCrossGeo, an automatically curated dataset of 85 globally distributed scenes sourced from Google Maps, Google Earth, and Google Street View. Evaluated on 2 Table 1: Datasets Comparison Name Data Source Images Grd UA V Map Grd Pose UA V Observation Scenes AnyVisLoc [55] Real 18,000✗✓✗- Multi-view Multiple University-1652 [58] Real 50,218✓ ✓ ✓✗no pose Buildings CVUSA [48] Real 44,516✓✗✓3DoF - Streets CV ACT [23] Real 44,416✓✗✓3DoF - Streets KITTI [7] Real 14,999✓✗✓3DoF - Streets VIGOR [60] Real 195,832✓✗✓3DoF - Multiple ULTRRA Challenge [14] Synthetic 1,207✓ ✓✗6DoF Multi-view Buildings AerialMD [37] Real+Synth 132,137✓ ✓✗6DoF Multi-view Multiple Ours (CrossGeo) Real+Synth 277,812✓ ✓ ✓6DoF Multi-view Multiple Figure 2: CrossGeo data sources.","citing_arxiv_id":"2605.07978"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Information-Preserving Domain Transfer with Unlabeled Data in Misspecified Simulation-Based Inference","primary_cat":"cs.LG","context_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulator labels: LNPE =L s(qψ) +λ infoE(θ,xs)∼ps(θ,xs) [−logq ψ(θ|sg(x srs))]. Here, sg(xsrs) keeps the transported observations fixed. This update extends posterior training to transported simulator-domain inputs, while preventing the posterior loss from changing the transport","citing_arxiv_id":"2605.05652"},{"n":1,"role":"method","polarity":"use_method","paper_title":"NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation","primary_cat":"cs.NI","context_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distributions, in our case, the observed application metrics from live 5G network and emulation environment. A lower EMD indicates a closer dis- tance between the two distributions, hence better emulation fidelity. We calculateemulation distribution errorin percent- age by dividing the EMD with the mean application metrics","citing_arxiv_id":"2604.26080"},{"n":1,"role":"method","polarity":"use_method","paper_title":"How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models","primary_cat":"cs.LG","context_text":"[36] Nicholas Roberts, Sungjun Cho, Zhiqi Gao, Tzu-Heng Huang, Albert Wu, Gabriel Orlanski, Avi Trost, Kelly Buchanan, Aws Albarghouthi, and Frederic Sala. Test-time scaling makes overtraining compute-optimal, 2026. URLhttps://arxiv.org/abs/2604.01411. [37] Gemma Team. Gemma 2: Improving open language models at a practical size, 2024. URL https://arxiv.org/abs/2408.00118. [38] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In2015 IEEE International Conference on Computer Vision (ICCV), pages 1026-1034, 2015. doi: 10.1109/ICCV .2015.123. 12 [39] Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre,","citing_arxiv_id":"2604.21106"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Weak-to-Strong Knowledge Distillation Accelerates Visual Learning","primary_cat":"cs.CV","context_text":"softened posteriors Ldistill(x, u) =T(u) 2 KL \u0010 pT(u) t (x)∥p T(u) s (x) \u0011 ,(3) whereT(u)is the temperature schedule, andp T t (x) = softmax(g ϕ(x)/T)and pT s (x) = softmax(f θ(x)/T). We use forward KL following [17], with temperature decayed from 6 to 1. Object Detection.For object detection,Lbase is the original detector loss (classifi- cation + box regression) [27,38]. Our distillation term aligns teacher and student prediction heads on the same images. We distill classification logits with temper- ature scaling and confidence masking (teacher-score threshold), and optionally add a box-regression alignment term, that is Ldet distill =L cls-distill +βL box-distill, whereL cls-distill distills classification logits,Lbox-distill distills box regression, and","citing_arxiv_id":"2604.15451"},{"n":1,"role":"method","polarity":"use_method","paper_title":"A Resource-Efficient Hybrid CNN-LSTM network for image-based bean leaf disease classification","primary_cat":"cs.CV","context_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridisscale)andshowsthat the feature extraction by the custom models becomes more specific in deeper layers. 4.2. Results: Bean-CNN vs. Bean-CNN-LSTM Figures 5 and 6 visualise the training progress of the best models trained on the original training set (905 samples).","citing_arxiv_id":"2604.13835"},{"n":1,"role":"dataset","polarity":"background","paper_title":"Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding","primary_cat":"cs.CV","context_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative & Rhetorical Analysis Humor/Sarcasm/Satire MUStARD (C) [96], UR-FUNNY (C) [97], MHD (C) [98], WITS (C) [99], ExFunTube (Cap) [100], YesBut (C) [101], V-FLUTE (Cap,C) [102],AVH (R) [103], FOR (C) [103] Visual Metaphors VMC (Cap) [104], V-FLUTE (Cap,C) [102], Mul-","citing_arxiv_id":"2508.20765"}]},"error":null,"updated_at":"2026-05-22T21:43:57.981267+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-22T21:44:05.768521+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"PoseNet: A convolutional network for real-time 6-dof camera relocalization","claims":[{"claim_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative ","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, \"Recognizing text with perspective distortion in natural scenes,\" in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, \"Toward understanding wordart: Corner-guided transformer for scene text recognition,\" in Pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks PoseNet: A convolutional network for real-time 6-dof camera relocalization because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (14 contexts).","role_counts":[{"n":14,"context_role":"background"},{"n":6,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-28T14:57:46.192428+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Enforcing geometric constraints of vir- tual normal for depth prediction","claims":[{"claim_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative ","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, \"Recognizing text with perspective distortion in natural scenes,\" in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, \"Toward understanding wordart: Corner-guided transformer for scene text recognition,\" in Pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Enforcing geometric constraints of vir- tual normal for depth prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":6,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-22T21:43:57.914992+00:00"}},"summary":{"title":"Enforcing geometric constraints of vir- tual normal for depth prediction","claims":[{"claim_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distribu","claim_type":"method","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative ","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Available: https://api.semanticscholar.org/CorpusID:15559857 [84] T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, \"Recognizing text with perspective distortion in natural scenes,\" in Proceedings of IEEE/CVF International Conference on Computer Vision . IEEE Computer Society, 2013, pp. 569-576. [Online]. Available: https://doi.org/10.1109/ICCV .2013.76 [85] X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, \"Toward understanding wordart: Corner-guided transformer for scene text recognition,\" in Pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulat","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridiss","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Class Activation Mapping methods were introduced to ad- dress the deployment trust gap by making CNN spatial rea- soning visible and auditable [12]. However, a series of foundational studies has revealed that CAM methods them- selves suffer from reliability failures that are independent of - and invisible to - classification performance met- rics. Model Parameter Randomization[13]: Several widely used explanation methods produce nearly identical heatmaps for a fully trained model and for a model","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Enforcing geometric constraints of vir- tual normal for depth prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (13 contexts).","role_counts":[{"n":13,"context_role":"background"},{"n":6,"context_role":"method"},{"n":3,"context_role":"dataset"}]},"graph":{"co_cited":[{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":13},{"title":"Adabins: Depth estimation using adap- tive bins","work_id":"7083a41e-5666-435b-ab26-c753f6490b9a","shared_citers":12},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":12},{"title":"In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","work_id":"9da51225-b7bd-4032-b7db-ca577971dafe","shared_citers":11},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":10},{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":10},{"title":"Poly- max: General dense prediction with mask transformer","work_id":"498f8786-e863-4217-b210-bf4fe976c779","shared_citers":8},{"title":"URLhttps://doi.org/10.48550/arXiv","work_id":"5c2060c6-427c-4321-be22-49ccae439d80","shared_citers":8},{"title":"URLhttp://dx.doi.org/10.1109/CVPR.2016.90","work_id":"b353bda2-591d-479a-9c8b-22dfcba12431","shared_citers":7},{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":6},{"title":"In: 2021 IEEE/CVF In- ternational Conference on Computer Vision (ICCV)","work_id":"3820f598-11b0-45c3-8c99-0079181ac0a7","shared_citers":6},{"title":"doi: 10.18653/v1/ 2024.findings-acl.586","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":5},{"title":"EW Dijkstra","work_id":"effdb28b-742e-4840-b3ca-d89502a6cd4d","shared_citers":5},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":4},{"title":"BERT : Pre-training of deep bidirectional transformers for language understanding","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":4},{"title":"Conditional prompt learning for vision- language models","work_id":"025819dc-724a-4ff8-ba0a-0ba72c046d8c","shared_citers":4},{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":4},{"title":"SAM 2: Segment Anything in Images and Videos","work_id":"acc13f66-d814-44f9-9688-375688bf2d4a","shared_citers":4},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":4},{"title":"","work_id":"ad0a8ee9-e814-486c-a25b-40126e136b0f","shared_citers":3},{"title":"2023.Parallel Programming for Multicore and Cluster Systems(3 ed.)","work_id":"cf4c4e77-acaa-46b4-b066-ddf045165d05","shared_citers":3},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":3},{"title":"doi:10.1109/CVPR46437","work_id":"ac0d5a4b-3ab1-462a-bed7-00be8d403f67","shared_citers":3},{"title":"GeoNum: Bridging numerical continuity and language semantics via geometric embedding","work_id":"9f349f1f-0e39-446f-8ac6-694a06c25de5","shared_citers":3}],"time_series":[{"n":1,"year":2017},{"n":1,"year":2023},{"n":1,"year":2024},{"n":15,"year":2025},{"n":48,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"LiBrA-Net: Lie-Algebraic Bilateral Affine Fields for Real-Time 4K Video Dehazing","primary_cat":"cs.CV","context_text":"at least linearly with pixel count, capping practical training at 1080p. Native 4K video dehazing therefore remains at the unattended intersection of these two lines: the UHD branch offers spatial scalability without temporal coherence, and the video branch offers temporal coherence without resolution scalability. 2.2 Bilateral Grids and Locally Affine Color Transforms The bilateral filter [34] and the guided filter [ 15] introduced edge-aware smoothing; the bilateral grid [4] cast the same operation as a 3D splat-blur-slice pipeline, and Bilateral Guided Upsampling [5] showed that locally affine transforms fitted in this space can approximate a wide class of image operators. HDRNet [ 11] made the fitting step learnable via differentiable slicing, and subsequent","citing_arxiv_id":"2605.11508"},{"n":1,"role":"dataset","polarity":"background","paper_title":"Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception","primary_cat":"cs.CV","context_text":"None captures authentic first-person spatial narratives of social media. Urban perception datasets including Place Pulse 2.0 [ 9], MMS-VPR [30], and UrbanFeel [ 13] focus on exterior street-level imagery, provide no textual modality, and lack multi-task annotation. Instance Segmentation.Cityscapes [ 6], ADE20K [42], LVIS [12], and Mapillary Vistas [28] cover outdoor driving and general scenes but apply no domain-specific vocabulary tailored to commercial spaces-the escalators, retail shelves, display cases, hotel beds, and food presentations that define the majority of Urban-ImageNet's images. Scaling Behaviour.ImageNet [ 7] established scale as a performance driver; GPT-3 [4] and scaling laws [18] showed predictable growth; LAION-5B [35] demonstrated billion-scale vision-language","citing_arxiv_id":"2605.09936"},{"n":1,"role":"dataset","polarity":"use_dataset","paper_title":"Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images","primary_cat":"cs.CV","context_text":"1), we construct and releaseCrossGeo, an automatically curated dataset of 85 globally distributed scenes sourced from Google Maps, Google Earth, and Google Street View. Evaluated on 2 Table 1: Datasets Comparison Name Data Source Images Grd UA V Map Grd Pose UA V Observation Scenes AnyVisLoc [55] Real 18,000✗✓✗- Multi-view Multiple University-1652 [58] Real 50,218✓ ✓ ✓✗no pose Buildings CVUSA [48] Real 44,516✓✗✓3DoF - Streets CV ACT [23] Real 44,416✓✗✓3DoF - Streets KITTI [7] Real 14,999✓✗✓3DoF - Streets VIGOR [60] Real 195,832✓✗✓3DoF - Multiple ULTRRA Challenge [14] Synthetic 1,207✓ ✓✗6DoF Multi-view Buildings AerialMD [37] Real+Synth 132,137✓ ✓✗6DoF Multi-view Multiple Ours (CrossGeo) Real+Synth 277,812✓ ✓ ✓6DoF Multi-view Multiple Figure 2: CrossGeo data sources.","citing_arxiv_id":"2605.07978"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Information-Preserving Domain Transfer with Unlabeled Data in Misspecified Simulation-Based Inference","primary_cat":"cs.LG","context_text":". Here, sg(·) denotes the stop-gradient operator, and qsg(ψ,ω) indicates that the summary network and posterior estimator are held fixed during the generator update. Thus, the information-preservation term updates only the transport networksG rs andG sr. Discriminator loss.The discriminators use hinge adversarial losses with spectral normalization [35]: LD =L D adv. Posterior loss.The posterior estimator is trained on both original simulated observations and transported observations with simulator labels: LNPE =L s(qψ) +λ infoE(θ,xs)∼ps(θ,xs) [−logq ψ(θ|sg(x srs))]. Here, sg(xsrs) keeps the transported observations fixed. This update extends posterior training to transported simulator-domain inputs, while preventing the posterior loss from changing the transport","citing_arxiv_id":"2605.05652"},{"n":1,"role":"method","polarity":"use_method","paper_title":"NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation","primary_cat":"cs.NI","context_text":"single-user scenarios, because they only support the single- user scenario. We also demonstrateNeuralEmu's ability in a multi-user scenario, which is the first of its kind. Emulation error metrics.We quantify the emulation errors using the normalized distributions difference between net- work environments: one from the live 5G network and the other from the emulation. We use Earth Mover's Distance (EMD) [39], defined as: EMD(L,T) = R ∞ −∞ |L(x)−T(x)|dx , where L and T are the CDF of two distributions, in our case, the observed application metrics from live 5G network and emulation environment. A lower EMD indicates a closer dis- tance between the two distributions, hence better emulation fidelity. We calculateemulation distribution errorin percent- age by dividing the EMD with the mean application metrics","citing_arxiv_id":"2604.26080"},{"n":1,"role":"method","polarity":"use_method","paper_title":"How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models","primary_cat":"cs.LG","context_text":"[36] Nicholas Roberts, Sungjun Cho, Zhiqi Gao, Tzu-Heng Huang, Albert Wu, Gabriel Orlanski, Avi Trost, Kelly Buchanan, Aws Albarghouthi, and Frederic Sala. Test-time scaling makes overtraining compute-optimal, 2026. URLhttps://arxiv.org/abs/2604.01411. [37] Gemma Team. Gemma 2: Improving open language models at a practical size, 2024. URL https://arxiv.org/abs/2408.00118. [38] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In2015 IEEE International Conference on Computer Vision (ICCV), pages 1026-1034, 2015. doi: 10.1109/ICCV .2015.123. 12 [39] Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre,","citing_arxiv_id":"2604.21106"},{"n":1,"role":"method","polarity":"use_method","paper_title":"Weak-to-Strong Knowledge Distillation Accelerates Visual Learning","primary_cat":"cs.CV","context_text":"softened posteriors Ldistill(x, u) =T(u) 2 KL \u0010 pT(u) t (x)∥p T(u) s (x) \u0011 ,(3) whereT(u)is the temperature schedule, andp T t (x) = softmax(g ϕ(x)/T)and pT s (x) = softmax(f θ(x)/T). We use forward KL following [17], with temperature decayed from 6 to 1. Object Detection.For object detection,Lbase is the original detector loss (classifi- cation + box regression) [27,38]. Our distillation term aligns teacher and student prediction heads on the same images. We distill classification logits with temper- ature scaling and confidence masking (teacher-score threshold), and optionally add a box-regression alignment term, that is Ldet distill =L cls-distill +βL box-distill, whereL cls-distill distills classification logits,Lbox-distill distills box regression, and","citing_arxiv_id":"2604.15451"},{"n":1,"role":"method","polarity":"use_method","paper_title":"A Resource-Efficient Hybrid CNN-LSTM network for image-based bean leaf disease classification","primary_cat":"cs.CV","context_text":"To ensure the best performance and the balance between overfitting and underfitting, we optimised each model's capacity,suchasthenumberofhiddenlayersandunits.Table 1describestheoptimalparameterandhyperparametervalues that we found during the iterative fine-tuning process for our custom models. We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [39] to visualise the features extracted in the convolutional layers. The weight distribution is represented inaheat-mapinFigure4(intheViridisscale)andshowsthat the feature extraction by the custom models becomes more specific in deeper layers. 4.2. Results: Bean-CNN vs. Bean-CNN-LSTM Figures 5 and 6 visualise the training progress of the best models trained on the original training set (905 samples).","citing_arxiv_id":"2604.13835"},{"n":1,"role":"dataset","polarity":"background","paper_title":"Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding","primary_cat":"cs.CV","context_text":"(C) [11], Ekman Emotion Dataset (C) [11], VAAD (C) [79], iMiGUE (C) [80], EALD (Q) [81], VCE (C) [82], V2V (R) [82], VEATIC (R) [83], MERR (C,Cap) [14], 3MASSIV (C) [70], LAMBDA (Q) [63], ArtEmis (C,Cap) [84], EmoSet (C) [85] Relationships SRIV (C) [86], ViSR (C) [87], PERR (C) [88], MovieGraphs (Q) [89], LVU (C) [66], VideoAds [69], Social Relation Dataset (C) [90], PISC (C) [91], PIPA (C) [92] Situation Analysis MovieGraphs (Q) [89], HLVU (Q) [93], Social-IQ (Q) [94], DeSIQ (Q) [95] Narrative & Rhetorical Analysis Humor/Sarcasm/Satire MUStARD (C) [96], UR-FUNNY (C) [97], MHD (C) [98], WITS (C) [99], ExFunTube (Cap) [100], YesBut (C) [101], V-FLUTE (Cap,C) [102],AVH (R) [103], FOR (C) [103] Visual Metaphors VMC (Cap) [104], V-FLUTE (Cap,C) [102], Mul-","citing_arxiv_id":"2508.20765"}]},"authors":[{"id":"ad0a11de-6175-4339-a8fa-1f60baf1c99e","orcid":null,"display_name":"doi: 10","source":"manual","import_confidence":0.72}]}}