REVIEW 4 major objections 6 minor 39 references
Temporally Extending Existing Web Archive Collections for Longitudinal Analysis
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Most Trump-era website deletions removed Obama-era additions
desk verdict The extension methodology is a real contribution, but the headline 87% finding is not reproducible from the paper's own counts or its unspecified term-classification method. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a temporal-extension methodology for web archive collections, built around the 2008–2016–2020 archived triplet as the unit of analysis. The method starts with the EDGI 2016–2020 collection, checks each page for 2008 mementos through TimeMap aggregation with MemGator, verifies archived HTTP statuses through CDX API lookups, crawls the past web with a sticky time policy to find pages that existed in 2008 and persisted to 2020, performs full-domain CDX queries for small poorly covered domains, combines the End of Term 2008 crawl, and finishes with iterative manual augmentation using the Wayback Machine's URLs tool and manual link replay. The methodology is what makes the 87 percent finding possible, because no single archive, crawl, or automated brute-force pass supplied enough 2008 captures.
What would settle it
Independently draw a random sample of the 15.8 million 2008 End of Term URLs from the 30 agencies, extend each forward to 2016 and 2020 without the paper's manual augmentation, and compute the share of Trump-deleted terms that were Obama-added; if that share falls far below 87 percent, the result is an artifact of the curation process rather than a property of the federal environmental web.
Extended reading notes
Core claim
The central claim is that the Trump administration's deletions from federal environmental websites were predominantly removals of terms first added during the Obama administration, and that this fact becomes visible only when a collection designed for 2016–2020 is temporally extended back to 2008. The authors contribute a dataset of 1,220 archived triplets, each with successful captures in 2008, 2016, and 2020, and report that 990 pages changed between 2008 and 2020 while 740 pages had Trump-era term deletions. Of those 740 pages, 87 percent had at least one deleted term added under Obama, whereas only 55 pages had deleted terms exclusively from the Bush era. They also find that repeated deletion patterns, called change trends, concentrated in agencies such as OSHA, NIH, and NOAA, and that among the 56 tracked terms, regulation-related terms were deleted more often than climate-related terms.
Load-bearing premise
The load-bearing premise is that the 1,220 pages with successful captures in 2008, 2016, and 2020 fairly represent the federal environmental pages whose content changed across administrations, even though pages deleted by 2020 and pages without 2008 captures are excluded by the way the sample was built.
Editorial extensions
If this is right
- The 1,220-triplet dataset provides a reusable longitudinal collection covering three US presidential administrations, and the method behind it can extend other existing collections backward in time.
- The 81 percent change rate between 2008 and 2020 shows that a URL persisting over twelve years does not imply that its content persisted.
- The 87 percent figure indicates that Trump-era deletions were largely removals of Obama-era additions, with only a small number of pages showing deletions exclusively from Bush-era content.
- Agency-level deletion patterns vary sharply: OSHA, NIH, and NOAA showed many repeated term deletions, while 17 of the 30 agencies showed no change trends at all.
- Among the 56 tracked terms, regulation-related deletions outnumbered climate-related deletions, suggesting the administration's rollback was broader than climate content alone.
Reading between the lines
- Beyond the paper: because the dataset excludes pages fully deleted by 2020 and pages with no 2008 capture, the 87 percent figure may understate the true share of Trump-era deletions that removed Obama-era additions, since the most aggressively changed pages are exactly the ones most likely to be missing from such a sample.
- Beyond the paper: two-thirds of the 2008 mementos came from Alexa crawls, so the 2008 snapshot reflects the interests and coverage of one commercial crawler; a different 2008 archive source could yield a different content mix and potentially a different deletion attribution.
- Beyond the paper: the same temporal-extension method could be applied to the 2020–2024 transition to test whether the pattern reverses, and the paper's own exemplar analysis already shows some Obama-era terms restored by 2024.
- Beyond the paper: the finding that manual augmentation was essential suggests that this method has a human-in-the-loop cost that will scale poorly to much larger collections, so automated approaches alone may not reproduce the result on broader domains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a methodology for temporally extending existing web archive collections by aggregating captures from multiple sources, and it applies this methodology to extend the EDGI collection of US federal environmental webpages back to 2008. The authors construct a dataset of 1,220 archived triplets (2008, 2016, 2020) across 30 agencies and use it to ask whether terms deleted from these pages during the Trump administration had been added during the Obama administration. They report that 81% of the pages changed between 2008 and 2020 and that 87% of pages with terms deleted under Trump had at least one deleted term added during the Obama administration. The paper also analyzes the provenance of the 2008 mementos, identifies change trends by agency, and examines whether deleted terms were restored by 2024.
Significance. If the empirical result were reliable, the paper would be a useful contribution to web-archiving methodology and to the political-science question of whether Trump-era deletions represented a rollback of Obama-era content. The dataset-assembly work is substantial: the authors document the need to aggregate captures from 16 organizations, quantify the limited contribution of the End of Term archive, and demonstrate that a single crawl source is insufficient for this period. The provenance analysis is a strength, as is the honest reporting of the manual augmentation steps. However, the headline longitudinal claim is currently not reproducible: the reported category counts do not sum to the stated denominator, the method for classifying when a term was added is not specified, and the temporal boundaries used do not align with the actual presidential administrations. These issues are load-bearing because the 87% figure is the central empirical claim in the abstract.
major comments (4)
- [Section 4.2, Figure 11] The three reported categories for the 740 pages with terms deleted during the Trump administration—373 with deleted terms added in both the Obama and Bush administrations, 274 with deleted terms only from Obama, and 55 with deleted terms only from Bush—sum to 702, not 740. The abstract's 87% figure is 647/740, but the reported counts yield 647/702 = 92.2%. The 38-page discrepancy is never explained. If the three categories are intended to be exhaustive, the denominator should be 702; if they are not exhaustive, the text must state how the remaining 38 pages were classified and why they are excluded. Without this, the paper's central empirical claim is not reproducible.
- [Section 4.2] The manuscript never specifies how a deleted term was determined to have been "added during the Obama administration." It cites prior work [9] for change-text search, but it does not describe the term-extraction algorithm, the criteria for "fully deleted," the handling of partially rendered or mis-captured mementos, or the exact capture dates used in the comparison. Because the dataset contains only a single 2008 memento per page, a term absent from that memento but present in a later 2008 memento would be classified as "Obama-added" even if it was introduced during the Bush administration. The stated classification window—"between July 2008 and June 2016"—also does not match the Obama administration, which began in January 2009. These omissions make the 87% headline result unverifiable.
- [Section 4.2] The analysis treats terms deleted "between July 2016 and July 2020" as deletions during the Trump administration, but the Trump administration did not begin until January 2017. Changes observed between July 2016 and January 2017 occurred while Obama was still president. The paper does not report the actual capture dates of the 2016 mementos, so the reader cannot determine how much of the observed deletion predates the Trump administration. This conflation directly affects the interpretation of the 87% result as evidence about Trump-era deletions.
- [Section 3, Section 5] The 87% result is computed on a non-random, quota-based convenience sample. Section 3 describes a target of 15 high and 15 deep links per agency, stopping criteria once quotas are met, and iterative manual augmentation, and the final dataset is heavily skewed toward domains with many archived captures (e.g., noaa.gov, ferc.gov, osha.gov). Because the percentages in Section 4.2 are unweighted, they describe this curated sample rather than the population of US federal environmental webpages. If the authors intend the result to answer the general research question about the Trump administration, the paper needs a representativeness analysis or a clearly hedged framing. As written, the conclusion in Section 5 ("show that the Republican president Trump deleted terms mostly added during the previous Democrat President Obama's administration") overstates what the dataset can support.
minor comments (6)
- [Section 4] The first sentence of Section 4 says the final dataset contains "1,200" triplets, while the rest of the paper (including the abstract, Figure 11, and Section 4.2) says "1,220." This should be harmonized.
- [Section 4.1.1, Table 5] The provenance counts in Table 5 sum to 1,211, not 1,220. In addition, several agency rows in Table 5 do not match the corresponding totals in Table 4 (for example, justice.gov shows 19 mementos in Table 5 but 24 triplets in Table 4, and nasa.gov shows 32 vs. 30). These discrepancies need to be resolved or explicitly explained.
- [Section 3.3] The sentence "Web crawlers like Heritrix CITE are available for crawling the live web" contains a dangling "CITE" that appears to be a placeholder for a citation.
- [Section 3.6] The sentence "The web interface gives the earliest and most recent archival dates of each page, , as shown in Figure 4" contains a doubled comma that should be removed.
- [Section 4.2.1] The word "surpising" should be "surprising."
- [Section 4.2.3] The text says that one-third of the 33 exemplars persisted until October 2024, which implies 11 pages, but the subsequent breakdown (6 pages with restored terms and 4 pages with no restored terms) sums to 10. This should be clarified.
Circularity Check
No significant circularity: the headline percentages are direct descriptive counts on a newly curated dataset, not model outputs or self-citation-dependent derivations.
full rationale
The paper's central claims, that 81 percent of pages changed between 2008 and 2020 and that 87 percent of pages with Trump-era deletions had deleted terms added during the Obama administration, are descriptive counts computed by comparing archived captures at three fixed time points. No parameter is fitted, no model is calibrated, and no quantity is predicted from the same data that is then used to explain it; the percentages are summaries of the curated triplet dataset. The dataset construction does involve target quotas and iterative manual augmentation, but inclusion criteria are not defined in terms of the conclusion: pages are selected for having successful 2008, 2016, and 2020 mementos, not for having Obama-added terms. This can introduce selection bias, but bias is a soundness concern, not circular reasoning. The self-citations in the paper are to prior tools for change-text search and CDX query methodology; these support the data-collection process and are not invoked as uniqueness theorems or as evidence that the deletions were Obama additions. The internal arithmetic discrepancy in Section 4.2, where 373 plus 274 plus 55 equals 702 rather than the stated 740, and the underspecified method for labeling a term as added during the Obama administration, are reproducibility and correctness risks; they do not constitute circularity because the classification uses distinct time intervals and does not reduce by definition to the paper's conclusion. Overall, the derivation chain is self-contained as a measurement study, and no circular step is exhibited.
Assumptions & free parameters
free parameters (3)
- Target quota of 15 high and 15 deep links per agency =
15 per category
- Change-trend threshold =
5 pages
- EDGI tracked term list =
56 terms
assumptions (4)
- domain assumption A successful archived HTTP status code at the Internet Archive in 2008, 2016, and 2020 indicates the same webpage was live at those times and is comparable across years.
- domain assumption The absence of a memento in the Internet Archive CDX for a URL means the page did not exist or was never archived, and the authors can infer non-existence.
- domain assumption Text extracted from archived mementos faithfully represents the page content at capture time for term-level diffing.
- domain assumption The 56 EDGI terms and their climate and regulation categorization are a valid operationalization of politically significant changes.
Cite this review
Pith. "Pith review of Temporally Extending Existing Web Archive Collections for Longitudinal Analysis." pith.science (2026). https://pith.science/paper/GTYEFBLO
@misc{pith2026250524091,
author = {Pith},
title = {Pith review of: Temporally Extending Existing Web Archive Collections for Longitudinal Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/GTYEFBLO}},
note = {Machine review of arXiv:2505.24091}
}
read the original abstract
The Environmental Governance and Data Initiative (EDGI) regularly crawled US federal environmental websites between 2016 and 2020 to capture changes between two presidential administrations. However, because it does not include the previous administration ending in 2008, the collection is unsuitable for answering our research question, Were the website terms deleted by the Trump administration (2017--2021) added by the Obama administration (2009--2017)? Thus, like many researchers using the Wayback Machine's holdings for historical analysis, we do not have access to a complete collection suiting our needs. To answer our research question, we must extend the EDGI collection back to January, 2008. This includes discovering relevant pages that were not included in the EDGI collection that persisted through 2020, not just going further back in time with the existing pages. We pieced together artifacts collected by various organizations for their purposes through many means (Save Page Now, Archive-It, and more) in order to curate a dataset sufficient for our intentions. In this paper, we contribute a methodology to extend existing web archive collections temporally to enable longitudinal analysis, including a dataset extended with this methodology. We use our new dataset to analyze our question, Were the website terms deleted by the Trump administration added by the Obama administration? We find that 81 percent of the pages in the dataset changed between 2008 and 2020, and that 87 percent of the pages with terms deleted by the Trump administration were terms added during the Obama administration.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Samantha Abrams, Zakiya Collier, Elena Colón-Marrero, keondra bills freemyn, Nick Krabbenhoeft, Melissa E. Wertheimer, and Amy Wickner. 2022. 2022 Web Archiving Survey Report. https://www.diglib.org/results-of-the-2022-ndsa-web-archiving-survey-report-now-available/
work page 2022
-
[2]
2009.Temporal-Informatics of the WWW
Eytan Adar. 2009.Temporal-Informatics of the WWW. Ph. D. Dissertation. University of Washington
work page 2009
-
[3]
Scott G Ainsworth and Michael L Nelson. 2013. Evaluating sliding and sticky target policies by measuring temporal drift in acyclic walks through a web archive. InProceedings of the 13th ACM/IEEE-CS joint conference on Digital libraries. 39–48
work page 2013
-
[4]
Sawood Alam and Michael L. Nelson. 2016. MemGator - A Portable Concurrent Memento Aggregator: Cross-Platform CLI and Server Binaries in Go. InProceedings of the 16th ACM/IEEE-CS on Joint Conference on Digital Libraries(Newark, New Jersey, USA)(JCDL ’16). Association for Computing Machinery, New York, NY, USA, 243–244. https://doi.org/10.1145/2910896.2925452
-
[5]
Anat Ben-David and Adam Amram. 2018. The Internet Archive and the socio-technical construction of historical facts.Internet Histories2, 1-2 (2018), 179–201
work page 2018
-
[6]
Jacob Bickford. 2013. Save Our Spiders: Crawler Traps and Sustainability at the UK Government Web Archive. https://www.dpconline.org/blog/ wdpd/wdpd2023-bickford
work page 2013
-
[7]
Roy Fielding, Mark Nottingham, and Julian Reschke. 2022. RFC 9110: HTTP Semantics. https://doi.org/10.17487/RFC9110
doi:10.17487/rfc9110 2022
-
[8]
Lesley Frew. 2023. Animating Changes in Webpages, Featuring George Santos’s Biography. https://ws-dl.blogspot.com/2023/02/2023-02-26- animating-changes-in.html
work page 2023
Show all 39 references
-
[10]
Nelson, and Michele C
Lesley Frew, Michael L. Nelson, and Michele C. Weigle. 2024. Retrogressive Document Manipulation of US Federal Environmental Websites. In Proceedings of the 33rd ACM International Conference on Information & Knowledge Management
2024
-
[11]
K. Garg, S. Alam, M. C. Weigle, M. L. Nelson, and D. Ayala. 2025. Not Here, Go There: Analyzing Redirection Patterns on the Web. . InProceedings of the 17th ACM Web Science Conference. https://doi.org/10.1145/3717867.3717925
2025
-
[12]
2021.The Past Web: Exploring Web Archives
Daniel Gomes, Elena Demidova, Jane Winters, and Thomas Risse (Eds.). 2021.The Past Web: Exploring Web Archives. Springer
2021
-
[13]
Daniel Gomes, João Miranda, and Miguel Costa. 2011. A Survey on Web Archiving Initiatives. InProceedings of Theory and Practice of Digital Libraries (TPDL). 408–420
2011
-
[14]
Gerhard Gossen, Thomas Risse, and Elena Demidova. 2020. Towards extracting event-centric collections from web archives.International Journal on Digital Libraries21, 1 (2020), 31–45
2020
-
[15]
Matt Grossmann and David A Hopkins. 2015. Ideological republicans and group interest democrats: The asymmetry of American party politics. Perspectives on Politics13, 1 (2015), 119–139
2015
-
[16]
Karen L Hanson and Karen Hanson. 2022. Preserving Innovation: Ensuring the Future of Today’s Scholarship.The Journal of Electronic Publishing 25, 1 (2022)
2022
-
[17]
2021.Improving collection understanding for web archives with storytelling: shining light into dark and stormy archives
Shawn M Jones. 2021.Improving collection understanding for web archives with storytelling: shining light into dark and stormy archives. Ph. D. Dissertation. Old Dominion University
2021
-
[18]
Frank Jotzo, Joanna Depledge, and Harald Winkler. 2018. US and international climate policy under President Trump. , 813–817 pages
2018
-
[19]
Graciela Kincaid and J Timmons Roberts. 2013. No talk, some walk: Obama administration first-term rhetoric on climate change and US international climate budget commitments.Global Environmental Politics13, 4 (2013), 41–60
2013
-
[20]
Martin Klein, Lyudmila Balakireva, and Herbert Van de Sompel. 2018. Focused crawl of web archives to build event collections. InProceedings of the 10th ACM Conference on Web Science. 333–342
2018
-
[21]
Martin Klein and Michael L Nelson. 2014. Moved but not gone: an evaluation of real-time methods for discovering replacement web pages. International Journal on Digital Libraries14, 1 (2014), 17–38
2014
-
[22]
Oren Kurland and Moshe Tennenholtz. 2022. Competitive Search. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2838–2849. https://doi.org/10.1145/3477495.3532771 Temporally Extending Existing Web Archive Collect...
2022
-
[23]
Emily Maemura. 2023. Sorting URLs out: Seeing the web through infrastructural inversion of archival crawling.Internet Histories7, 4 (2023), 386–401
2023
-
[24]
Oliver Milman and Sam Morris. 2017. Trump is deleting climate change, one site at a time. https://www.theguardian.com/us-news/2017/may/14/ donald-trump-climate-change-mentions-government-websites
2017
-
[25]
Gordon Mohr, Michael Stack, Igor Rnitovic, Dan Avery, and Michele Kimpton. 2004. Introduction to heritrix. In4th International Web Archiving Workshop. Citeseer, 109–115
2004
-
[26]
Eric Nost, Gretchen Gehrke, Grace Poudrier, Aaron Lemelin, Marcy Beck, Sara Wylie, on behalf of the Environmental Data, and Governance Initiative. 2021. Visualizing changes to US federal environmental agency websites, 2016–2020.PLOS ONE16, 2 (02 2021), 1–27. https://doi.org/10...
2021
-
[27]
Jessica Ogden, Edward Summers, and Shawn Walker. 2024. Know (ing) Infrastructure: The Wayback Machine as object and instrument of digital research.Convergence30, 1 (2024), 167–189
2024
-
[28]
Arnold Overwijk, Chenyan Xiong, and Jamie Callan. 2022. ClueWeb22: 10 Billion Web Documents with Rich Information. InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 3360–3362
2022
-
[29]
Mark Phillips, Dan Chudnov, and James Jacobs. 2016. Exploratory Analysis of the End of Term Web Archive: Comparing two collections. Presented at the ACM/IEEE JCDL 2016 Workshop on Web Archiving and Digital Libraries (WADL)
2016
-
[30]
Mark E Phillips and Kristy K Phillips. 2017. End of Term 2016 Presidential Web Archive.Against the Grain29, 6 (2017), 10
2017
-
[31]
Phillips, Kristy K
Mark E. Phillips, Kristy K. Phillips, and Sawood Alam. 2023. End of Term Web Archive Dataset: Longitudinal Web Archive of .GOV and .MIL Domains. In2023 ACM/IEEE Joint Conference on Digital Libraries (JCDL). 98–101. https://doi.org/10.1109/JCDL57899.2023.00024
2023
-
[32]
Rachel Augustine Potter, Andrew Rudalevige, Sharece Thrower, and Adam L Warber. 2019. Continuity trumps change: The first year of Trump’s administrative presidency.PS: Political Science & Politics52, 4 (2019), 613–619
2019
-
[33]
Anna Rakityanskaya. 2023. Belarusian Politics and Society Web Archive: Preserving the Belarusian Grassroots Protest.Journal of Belarusian Studies 12, 1-2 (2023), 81–94
2023
-
[34]
Tracy Seneca, Abbie Grotke, Cathy Nelson Hartman, and Kris Carpenter. 2012. It takes a village to save the web: The End of Term Web Archive. Documents to the People (DttP)40 (2012), 16
2012
-
[35]
Kristinn Sigurðsson, Michael Stack, and Igor Ranitovic. 2006. Heritrix user manual: sort-friendly URI reordering transform. http://crawler.archive. org/articles/user_manual/glossary.html#surt
2006
-
[36]
Nicholas Taylor. 2023. Beyond the Affidavit: Towards Better Standards for Web Archive Evidence. InIIPC Web Archiving Conference, Hilversum, Netherlands
2023
-
[37]
Herbert Van de Sompel, Michael Nelson, and Robert Sanderson. 2013. RFC 7089 - HTTP framework for time-based access to resource states–Memento. https://tools.ietf.org/html/rfc7089
2013
-
[38]
Michele C. Weigle. 2024. Some URLs Are Immortal, Most Are Ephemeral. https://ws-dl.blogspot.com/2024/09/2024-09-20-some-urls-are-immortal- most.html
2024
-
[39]
Michele C Weigle, Michael L Nelson, Sawood Alam, and Mark Graham. 2023. Right HTML, wrong JSON: challenges in replaying archived webpages built with client-side rendering. In2023 ACM/IEEE Joint Conference on Digital Libraries (JCDL). IEEE, 82–92
2023
-
[40]
Attribution-NonCommercial-ShareAlike 4.0 International
Diyi Yang, Aaron Halfaker, Robert Kraut, and Eduard Hovy. 2017. Identifying semantic edit intentions from revisions in Wikipedia. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. 2000–2010. https://doi.org/10.18653/v1/D17-1213 This work...
2017 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.