{"work":{"id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","openalex_id":"https://openalex.org/W1686810756","doi":"10.48550/arxiv.1409.1556","arxiv_id":"1409.1556","raw_key":null,"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","authors":null,"authors_text":"Karen Simonyan and Andrew Zisserman","year":2014,"venue":"cs.CV","abstract":"In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respectively. We also show that our representations generalise well to other datasets, where they achieve state-of-the-art results. We have made our two best-performing ConvNet models publicly available to facilitate further research on the use of deep visual representations in computer vision.","external_url":"https://arxiv.org/abs/1409.1556","cited_by_count":75540,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1409.1556","created_at":"2026-05-08T21:39:25.317250+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","render_title":"Very Deep Convolutional Networks for Large-Scale Image Recognition"},"hub":{"state":{"work_id":"1c4b4409-c14b-488b-a086-c57a5aab8a29","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":451,"external_cited_by_count":75540,"distinct_field_count":42,"first_pith_cited_at":"2015-05-18T11:28:37+00:00","last_pith_cited_at":"2026-07-09T15:29:04+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T11:49:18.432173+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":24},{"context_role":"method","n":12},{"context_role":"dataset","n":2}],"polarity_counts":[{"context_polarity":"background","n":22},{"context_polarity":"use_method","n":11},{"context_polarity":"use_dataset","n":2},{"context_polarity":"baseline","n":1},{"context_polarity":"support","n":1},{"context_polarity":"unclear","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","claims":[{"claim_text":"In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respective","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Very Deep Convolutional Networks for Large-Scale Image Recognition because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T23:44:00.338098+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"21a9b4f1-1888-4f6d-a31f-e0c9e30a993c","orcid":null,"display_name":"Karen Simonyan and Andrew Zisserman"}]},"error":null,"updated_at":"2026-05-13T23:44:00.997090+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-13T23:44:04.820541+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":25},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":18},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications","work_id":"3870239a-c950-4625-bf33-c4f902d14175","shared_citers":11},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":9},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":7},{"title":"Deep Residual Learning for Image Recognition","work_id":"ae9e5671-23e8-4853-82a4-699b5b8dd639","shared_citers":7},{"title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift","work_id":"05484516-8937-4cdf-9176-7f8329ef0221","shared_citers":6},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":6},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":6},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":6},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size.arXiv2016","work_id":"755e91c7-dab5-418e-9316-bda99a153fed","shared_citers":6},{"title":"An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion","work_id":"ca618c21-3ba6-448e-bd86-bcecff3cdeb5","shared_citers":5},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":5},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":5},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":5},{"title":"arXiv preprint arXiv:1511.07122 , year=","work_id":"4784b5b3-7851-455f-ba76-9446e31e6b96","shared_citers":4},{"title":"BEiT: BERT Pre-Training of Image Transformers","work_id":"d74eda3c-bf7e-45f1-a8f1-a0137ecca3f4","shared_citers":4},{"title":"& Darrell, T","work_id":"237de0db-0556-4b65-b014-bb29418efc91","shared_citers":4},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":4}],"time_series":[{"n":3,"year":2015},{"n":2,"year":2016},{"n":2,"year":2017},{"n":1,"year":2018},{"n":2,"year":2020},{"n":1,"year":2021},{"n":1,"year":2022},{"n":1,"year":2024},{"n":3,"year":2025},{"n":113,"year":2026}]},"error":null,"updated_at":"2026-05-13T23:44:00.429546+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"fixed":1,"items":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-13T23:43:59.165715+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","claims":[{"claim_text":"In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respective","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Very Deep Convolutional Networks for Large-Scale Image Recognition because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T23:43:54.571976+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","claims":[{"claim_text":"In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respective","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Very Deep Convolutional Networks for Large-Scale Image Recognition because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T23:44:01.003659+00:00"}},"summary":{"title":"Very Deep Convolutional Networks for Large-Scale Image Recognition","claims":[{"claim_text":"In this work we investigate the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting. Our main contribution is a thorough evaluation of networks of increasing depth using an architecture with very small (3x3) convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers. These findings were the basis of our ImageNet Challenge 2014 submission, where our team secured the first and the second places in the localisation and classification tracks respective","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Very Deep Convolutional Networks for Large-Scale Image Recognition because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","work_id":"e96730e3-129b-4db6-b981-15ab7932e297","shared_citers":25},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":18},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":11},{"title":"MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications","work_id":"3870239a-c950-4625-bf33-c4f902d14175","shared_citers":11},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":9},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":7},{"title":"Deep Residual Learning for Image Recognition","work_id":"ae9e5671-23e8-4853-82a4-699b5b8dd639","shared_citers":7},{"title":"Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift","work_id":"05484516-8937-4cdf-9176-7f8329ef0221","shared_citers":6},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":6},{"title":"DINOv3","work_id":"c8b07deb-8fe7-4e18-9620-f3569d3529ce","shared_citers":6},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":6},{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size.arXiv2016","work_id":"755e91c7-dab5-418e-9316-bda99a153fed","shared_citers":6},{"title":"An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion","work_id":"ca618c21-3ba6-448e-bd86-bcecff3cdeb5","shared_citers":5},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":5},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":5},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":5},{"title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","work_id":"8abcfe4f-e0fb-44b7-9123-448fac95f90a","shared_citers":5},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":5},{"title":"arXiv preprint arXiv:1511.07122 , year=","work_id":"4784b5b3-7851-455f-ba76-9446e31e6b96","shared_citers":4},{"title":"BEiT: BERT Pre-Training of Image Transformers","work_id":"d74eda3c-bf7e-45f1-a8f1-a0137ecca3f4","shared_citers":4},{"title":"& Darrell, T","work_id":"237de0db-0556-4b65-b014-bb29418efc91","shared_citers":4},{"title":"Distilling the Knowledge in a Neural Network","work_id":"d927ab1f-17b8-4002-9d09-c3d55764fbad","shared_citers":4}],"time_series":[{"n":3,"year":2015},{"n":2,"year":2016},{"n":2,"year":2017},{"n":1,"year":2018},{"n":2,"year":2020},{"n":1,"year":2021},{"n":1,"year":2022},{"n":1,"year":2024},{"n":3,"year":2025},{"n":113,"year":2026}]},"authors":[{"id":"21a9b4f1-1888-4f6d-a31f-e0c9e30a993c","orcid":null,"display_name":"Karen Simonyan and Andrew Zisserman","source":"manual","import_confidence":0.72}]}}