{"work":{"id":"da360c40-6481-4088-bd96-8e73e0280a6b","openalex_id":null,"doi":null,"arxiv_id":null,"raw_key":"raw:6a0a6258a17fc6b9201fc70c","title":"Proceedings of the IEEE conference on computer vision and pattern recognition , pages=","authors":null,"authors_text":"Deep residual learning for image recognition , author=","year":null,"venue":null,"abstract":null,"external_url":null,"cited_by_count":null,"metadata_source":"raw_reference","metadata_fetched_at":"2026-07-11T03:47:54.372474+00:00","pith_arxiv_id":null,"created_at":"2026-05-12T11:11:32.235729+00:00","updated_at":"2026-07-11T03:47:54.372474+00:00","title_quality_ok":false,"display_title":"Proceedings of the IEEE conference on computer vision and pattern recognition , pages=","render_title":"Proceedings of the IEEE conference on computer vision and pattern recognition , pages="},"hub":{"state":{"work_id":"da360c40-6481-4088-bd96-8e73e0280a6b","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":103,"external_cited_by_count":null,"distinct_field_count":17,"first_pith_cited_at":"2018-11-27T13:17:45+00:00","last_pith_cited_at":"2026-07-09T15:29:04+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T02:29:35.211853+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":7},{"context_role":"method","n":2}],"polarity_counts":[{"context_polarity":"background","n":4},{"context_polarity":"unclear","n":3},{"context_polarity":"use_method","n":2}],"runs":{"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-22T10:03:54.383822+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":16},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":15},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":13},{"title":"2016 , publisher=","work_id":"cf0899e0-53ee-4591-aae4-f38fa5ac12ad","shared_citers":12},{"title":"International conference on machine learning , pages=","work_id":"1031f0f0-cd67-4170-83ad-9a088b363c67","shared_citers":12},{"title":"and Osindero, Simon and Teh, Yee Whye , journal =","work_id":"0a5921e3-ac4e-46f1-85ae-866119a87be0","shared_citers":11},{"title":"Scaling Learning Algorithms Towards","work_id":"bb2761cc-98d0-411b-92f6-803773d64460","shared_citers":11},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":10},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":10},{"title":"2009 IEEE conference on computer vision and pattern recognition , pages=","work_id":"0287b3d7-3294-4d36-bdf0-64add6c6c84b","shared_citers":9},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":8},{"title":"Proceedings of the IEEE/CVF international conference on computer vision , pages=","work_id":"c9e852df-fd5d-4c02-8f60-78e4f614d3f4","shared_citers":8},{"title":"2009 , publisher=","work_id":"75c9142a-a015-42fa-b181-87ac5c0ba487","shared_citers":7},{"title":"Advances in neural information processing systems , volume=","work_id":"12f5a236-ef7a-4d13-b4de-b51465a6f977","shared_citers":7},{"title":"nature , volume=","work_id":"e6b92db6-c2c1-4e28-9c46-8d7bfebba8f1","shared_citers":7},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":7},{"title":"Advances in neural information processing systems , volume=","work_id":"d0910d33-1f9a-4803-a985-78f07ec4afd5","shared_citers":6},{"title":"Advances in neural information processing systems , volume=","work_id":"48172a5a-0dfc-45cf-9fdc-988f99c16450","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":6},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":6},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":6},{"title":"Advances in neural information processing systems , volume=","work_id":"f2a4f68e-48af-40c6-90d4-76b155136aa5","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":5}],"time_series":[{"n":1,"year":2020},{"n":1,"year":2022},{"n":4,"year":2023},{"n":4,"year":2024},{"n":2,"year":2025},{"n":53,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Spatial Adapter: Structured Spatial Decomposition and Closed-Form Covariance for Frozen Predictors","primary_cat":"stat.ML","context_text":"This appendix supplies the predictive-variance machinery deferred from Section 2.5. Kriging conditional variance.Given an observation set Oj⊆{1,...,N}at sample j, the fixed-rank kriging conditional variance (Cressie and Johannesson, 2008) at a query locations∗is ˆv(s∗,j) :=ˆσ2 + ˆϕ(s∗)⊤ˆΛ cond(Oj)ˆϕ(s∗),(16) with plug-in conditional score covariance ˆΛ cond(Oj) := ( ˆΛ−1+ˆσ−2Φ⊤ OjΦOj )−1 .(17) WhenOj =∅(new-sample regime),ˆΛ cond =ˆΛ and ˆvreduces to the marginal variance; when Oj is the training-station set (Weather2K spatial-holdout, Section 4.1),ˆΛ cond shrinks strictly below ˆΛand recovers the standard kriging variance. Proposition 2(Plug-in conditional Gaussian predictive law).Under the optional Gaussian working modelαj∼N(0, Λ), ϵj∼N(0,σ2I), plugging the ADMM estimates(ˆΦ, ˆΛ,ˆσ2,ˆθ,ˆψ)","citing_arxiv_id":"2605.11394"}]},"error":null,"updated_at":"2026-05-22T10:03:54.564301+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-22T10:03:48.343188+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Proceedings of the IEEE conference on computer vision and pattern recognition , pages=","claims":[{"claim_text":"This appendix supplies the predictive-variance machinery deferred from Section 2.5. Kriging conditional variance.Given an observation set Oj⊆{1,...,N}at sample j, the fixed-rank kriging conditional variance (Cressie and Johannesson, 2008) at a query locations∗is ˆv(s∗,j) :=ˆσ2 + ˆϕ(s∗)⊤ˆΛ cond(Oj)ˆϕ(s∗),(16) with plug-in conditional score covariance ˆΛ cond(Oj) := ( ˆΛ−1+ˆσ−2Φ⊤ OjΦOj )−1 .(17) WhenOj =∅(new-sample regime),ˆΛ cond =ˆΛ and ˆvreduces to the marginal variance; when Oj is the trainin","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Multi-Generation Regression Outputs.For a minibatch of inputs {x1, x2, . . . , xB}, GRPO samples K independent generation trajectories for each input. This yields a set of numeric predictions: q(xi) = \u0002 q1(xi), q2(xi), . . . , qK (xi) \u0003⊤ ,(1) which naturally encode prediction variability. We summa- rize these outputs using their empirical mean: µ(xi) = 1 K KX k=1 qk(xi),(2) which provides a stable, low-variance estimate for each compared sample during reward computation. Batch-Level Relational C","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"(28) In particular,Γ j,R is not a fold trace (e.g., it excludesz(y) =|y 1|along{y 1 = 0}), and the trace locally separatesU j,R into two nonempty strict-sign sides. (A3) Non-redundancy of intersecting traces (essentiality).For everyj∈ Kwith Zj ∩int(P)̸=∅, there exists a point y∈Z j ∩int(P)such thatz k(y)̸= 0∀k∈ K \\ {j}.(29) Equivalently, Zj ∩int(P)̸⊆ [ k∈K\\{j} Zk.(30) Lemma 3 (Counting cells intersectingPvs.int(P))LetP⊂R d be a convex closed polytope withint(P)̸=∅, and letPbe any collection of f","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"},{"claim_text":"In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217-2226, 2019. JulianIbarz, JieTan, ChelseaFinn, MrinalKalakrishnan, PeterPastor, andSergeyLevine. Howtotrainyour robot with deep reinforcement l","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"},{"claim_text":"For FEDADAGRAD , we setβ1 =β2 = 0 (as typical versions of ADAGRAD do not use momentum). For FEDADAM and FEDYOGI we setβ1 = 0.9,β 2 = 0.99. While these parameters are generally 22 Published as a conference paper at ICLR 2021 Algorithm 5 FEDADAGRADFEDADAGRADFEDADAGRAD , FEDYOGIFEDYOGIFEDYOGI , and FEDADAMFEDADAMFEDADAM - Batched data Input:x0,v−1≥τ 2, optionalβ1,β 2∈ [0, 1) for FEDYOGI and FEDADAM fort = 0,··· ,T − 1 do Sample a subsetS of clients xt i =xt for each clienti∈S in parallel do fore = ","claim_type":"background","confidence":0.4,"evidence_strength":"citation_context"},{"claim_text":"Jingkang Yang, Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, Wenxuan Peng, Haoqi Wang, Guangyao Chen, Bo Li, Yiyou Sun, et al. Openood: Benchmarking generalized out-of- distribution detection.Advances in Neural Information Processing Systems, 35:32598-32611, 2022. Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey.International Journal of Computer Vision, 132(12):5635-5662, 2024. 11 A TRAININGCONFIGURATION All experiments share a com","claim_type":"background","confidence":0.3,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Proceedings of the IEEE conference on computer vision and pattern recognition , pages= because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (5 contexts).","role_counts":[{"n":5,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-22T10:03:48.348090+00:00"}},"summary":{"title":"Proceedings of the IEEE conference on computer vision and pattern recognition , pages=","claims":[{"claim_text":"This appendix supplies the predictive-variance machinery deferred from Section 2.5. Kriging conditional variance.Given an observation set Oj⊆{1,...,N}at sample j, the fixed-rank kriging conditional variance (Cressie and Johannesson, 2008) at a query locations∗is ˆv(s∗,j) :=ˆσ2 + ˆϕ(s∗)⊤ˆΛ cond(Oj)ˆϕ(s∗),(16) with plug-in conditional score covariance ˆΛ cond(Oj) := ( ˆΛ−1+ˆσ−2Φ⊤ OjΦOj )−1 .(17) WhenOj =∅(new-sample regime),ˆΛ cond =ˆΛ and ˆvreduces to the marginal variance; when Oj is the trainin","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Multi-Generation Regression Outputs.For a minibatch of inputs {x1, x2, . . . , xB}, GRPO samples K independent generation trajectories for each input. This yields a set of numeric predictions: q(xi) = \u0002 q1(xi), q2(xi), . . . , qK (xi) \u0003⊤ ,(1) which naturally encode prediction variability. We summa- rize these outputs using their empirical mean: µ(xi) = 1 K KX k=1 qk(xi),(2) which provides a stable, low-variance estimate for each compared sample during reward computation. Batch-Level Relational C","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"(28) In particular,Γ j,R is not a fold trace (e.g., it excludesz(y) =|y 1|along{y 1 = 0}), and the trace locally separatesU j,R into two nonempty strict-sign sides. (A3) Non-redundancy of intersecting traces (essentiality).For everyj∈ Kwith Zj ∩int(P)̸=∅, there exists a point y∈Z j ∩int(P)such thatz k(y)̸= 0∀k∈ K \\ {j}.(29) Equivalently, Zj ∩int(P)̸⊆ [ k∈K\\{j} Zk.(30) Lemma 3 (Counting cells intersectingPvs.int(P))LetP⊂R d be a convex closed polytope withint(P)̸=∅, and letPbe any collection of f","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"},{"claim_text":"In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770-778, 2016. Patrick Helber, Benjamin Bischke, Andreas Dengel, and Damian Borth. Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification.IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 12(7):2217-2226, 2019. JulianIbarz, JieTan, ChelseaFinn, MrinalKalakrishnan, PeterPastor, andSergeyLevine. Howtotrainyour robot with deep reinforcement l","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"},{"claim_text":"For FEDADAGRAD , we setβ1 =β2 = 0 (as typical versions of ADAGRAD do not use momentum). For FEDADAM and FEDYOGI we setβ1 = 0.9,β 2 = 0.99. While these parameters are generally 22 Published as a conference paper at ICLR 2021 Algorithm 5 FEDADAGRADFEDADAGRADFEDADAGRAD , FEDYOGIFEDYOGIFEDYOGI , and FEDADAMFEDADAMFEDADAM - Batched data Input:x0,v−1≥τ 2, optionalβ1,β 2∈ [0, 1) for FEDYOGI and FEDADAM fort = 0,··· ,T − 1 do Sample a subsetS of clients xt i =xt for each clienti∈S in parallel do fore = ","claim_type":"background","confidence":0.4,"evidence_strength":"citation_context"},{"claim_text":"Jingkang Yang, Pengyun Wang, Dejian Zou, Zitang Zhou, Kunyuan Ding, Wenxuan Peng, Haoqi Wang, Guangyao Chen, Bo Li, Yiyou Sun, et al. Openood: Benchmarking generalized out-of- distribution detection.Advances in Neural Information Processing Systems, 35:32598-32611, 2022. Jingkang Yang, Kaiyang Zhou, Yixuan Li, and Ziwei Liu. Generalized out-of-distribution detection: A survey.International Journal of Computer Vision, 132(12):5635-5662, 2024. 11 A TRAININGCONFIGURATION All experiments share a com","claim_type":"background","confidence":0.3,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Proceedings of the IEEE conference on computer vision and pattern recognition , pages= because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (5 contexts).","role_counts":[{"n":5,"context_role":"background"},{"n":1,"context_role":"method"}]},"graph":{"co_cited":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":16},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":15},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":13},{"title":"2016 , publisher=","work_id":"cf0899e0-53ee-4591-aae4-f38fa5ac12ad","shared_citers":12},{"title":"International conference on machine learning , pages=","work_id":"1031f0f0-cd67-4170-83ad-9a088b363c67","shared_citers":12},{"title":"and Osindero, Simon and Teh, Yee Whye , journal =","work_id":"0a5921e3-ac4e-46f1-85ae-866119a87be0","shared_citers":11},{"title":"Scaling Learning Algorithms Towards","work_id":"bb2761cc-98d0-411b-92f6-803773d64460","shared_citers":11},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":10},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":10},{"title":"2009 IEEE conference on computer vision and pattern recognition , pages=","work_id":"0287b3d7-3294-4d36-bdf0-64add6c6c84b","shared_citers":9},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":8},{"title":"Proceedings of the IEEE/CVF international conference on computer vision , pages=","work_id":"c9e852df-fd5d-4c02-8f60-78e4f614d3f4","shared_citers":8},{"title":"2009 , publisher=","work_id":"75c9142a-a015-42fa-b181-87ac5c0ba487","shared_citers":7},{"title":"Advances in neural information processing systems , volume=","work_id":"12f5a236-ef7a-4d13-b4de-b51465a6f977","shared_citers":7},{"title":"nature , volume=","work_id":"e6b92db6-c2c1-4e28-9c46-8d7bfebba8f1","shared_citers":7},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":7},{"title":"Advances in neural information processing systems , volume=","work_id":"d0910d33-1f9a-4803-a985-78f07ec4afd5","shared_citers":6},{"title":"Advances in neural information processing systems , volume=","work_id":"48172a5a-0dfc-45cf-9fdc-988f99c16450","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":6},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":6},{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","shared_citers":6},{"title":"Advances in neural information processing systems , volume=","work_id":"f2a4f68e-48af-40c6-90d4-76b155136aa5","shared_citers":5},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":5}],"time_series":[{"n":1,"year":2020},{"n":1,"year":2022},{"n":4,"year":2023},{"n":4,"year":2024},{"n":2,"year":2025},{"n":53,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Spatial Adapter: Structured Spatial Decomposition and Closed-Form Covariance for Frozen Predictors","primary_cat":"stat.ML","context_text":"This appendix supplies the predictive-variance machinery deferred from Section 2.5. Kriging conditional variance.Given an observation set Oj⊆{1,...,N}at sample j, the fixed-rank kriging conditional variance (Cressie and Johannesson, 2008) at a query locations∗is ˆv(s∗,j) :=ˆσ2 + ˆϕ(s∗)⊤ˆΛ cond(Oj)ˆϕ(s∗),(16) with plug-in conditional score covariance ˆΛ cond(Oj) := ( ˆΛ−1+ˆσ−2Φ⊤ OjΦOj )−1 .(17) WhenOj =∅(new-sample regime),ˆΛ cond =ˆΛ and ˆvreduces to the marginal variance; when Oj is the training-station set (Weather2K spatial-holdout, Section 4.1),ˆΛ cond shrinks strictly below ˆΛand recovers the standard kriging variance. Proposition 2(Plug-in conditional Gaussian predictive law).Under the optional Gaussian working modelαj∼N(0, Λ), ϵj∼N(0,σ2I), plugging the ADMM estimates(ˆΦ, ˆΛ,ˆσ2,ˆθ,ˆψ)","citing_arxiv_id":"2605.11394"}]},"authors":[]}}