Every link is a promise that someone else's server will keep saying the same thing. The research says that promise is broken often, and faster than most people assume.
This page collects the most-cited studies of link rot and content drift, with what each one measured, so you can quote them accurately.
The short version
| Finding | Source |
|---|---|
| 38% of web pages that existed in 2013 were no longer accessible a decade later | Pew Research Center, 2024 |
| 25% of all pages that existed at some point between 2013 and 2023 were no longer accessible | Pew Research Center, 2024 |
| Even among pages from 2023, 8% were no longer accessible by late 2023 | Pew Research Center, 2024 |
| 23% of news pages and 21% of government pages contain at least one broken link | Pew Research Center, 2024 |
| 54% of Wikipedia articles have at least one broken link in their references | Pew Research Center, 2024 |
| 50% of links in US Supreme Court opinions no longer show the cited material | Zittrain, Albert and Lessig, 2014 |
| More than 70% of links in three Harvard law journals no longer show the cited material | Zittrain, Albert and Lessig, 2014 |
| One in five science, technology and medicine articles suffers from reference rot, rising to seven in ten among those that cite web pages | Klein et al., 2014 |
| For more than 75% of web references in scholarly articles, the content has changed since it was cited | Jones et al., 2016 |
Link rot and content drift are different problems
Link rot is the familiar one: you follow a link and get an error, a parked domain or a redirect to a home page. It is annoying, but at least it is obvious.
Content drift is worse. The link still works, but the page says something different from what it said when it was cited. Nothing tells you. A reader who follows the link reasonably assumes they are seeing what the author saw.
For anyone who uses web pages as evidence, drift is the dangerous one. A terms page that returns an error is a missing exhibit. A terms page that has been quietly rewritten looks like a working link to the version the other side prefers.
How much of the web disappears
The largest recent study is Pew Research Center's 2024 report, When Online Content Disappears. Pew sampled pages from Common Crawl snapshots taken each year from 2013 to 2023 and checked whether they could still be reached in late 2023.
- 38% of pages that existed in 2013 were no longer accessible.
- 25% of all pages collected between 2013 and 2023 were no longer accessible.
- Even recent pages go: 8% of pages from 2023 had already disappeared.
The loss isn't limited to obscure corners of the web. Pew found that 23% of news pages and 21% of government pages contained at least one broken link. On Wikipedia, 54% of articles had at least one link in their references pointing to a page that no longer exists, and 11% of all references were no longer accessible.
Citations in law and science
In 2014, Jonathan Zittrain, Kendra Albert and Lawrence Lessig studied links cited in legal writing. They found that 50% of the URLs in US Supreme Court opinions suffered reference rot, meaning they no longer produced the information originally cited. In the Harvard Law Review, the Harvard Journal of Law and Technology and the Harvard Human Rights Journal, the figure was more than 70%.
The same study also measured drift directly. Of the Supreme Court links that still loaded successfully, only 76% still led to the material cited. A working link was no guarantee of the right content.
In science, Klein and colleagues (2014) examined over a million references in almost 400,000 science, technology and medicine articles published between 1997 and 2012. One in five articles suffered from reference rot. Among articles that referenced web resources at all, it was seven in ten.
A follow-up by Jones and colleagues (2016) looked at drift specifically. For more than 75% of references to web resources, the content had drifted away from what it was when it was cited, even where the link still resolved.
Why pages disappear
The studies agree that most loss is ordinary, not sinister:
- Redesigns and CMS migrations change every address on a site, and old addresses are rarely redirected one by one.
- Companies and projects end. Domains lapse and are re-registered by someone else.
- Content moves behind logins and paywalls, or into apps.
- Pages are edited in place. Terms, prices, policies and product claims are updated without keeping the old version, and there is usually no public history.
The last one is content drift, and nothing about it is unusual. Businesses change their terms, and they're under no obligation to publish what the old ones said.
What to do about it
If you cite for readers, such as in a paper, an article or a brief, pair each important link with an archived copy. Perma.cc, created for legal citations after the Zittrain study, and the Internet Archive's Save Page Now both create a public, citable archived copy.
If you may need to prove what a page said, a public archive isn't enough. You don't control when it captured the page, whether it captured it at all, or whether the copy stays available. We go into this in Wayback Machine captures as evidence. Keep your own copy of the pages that matter, with an independent timestamp, at the moment they matter. Our guide on how to preserve a web page as evidence has a checklist.
If the page is likely to change, one capture isn't enough. Capture it on a schedule, so that when it changes, the version from before is already on file.
Sources
- Pew Research Center, When Online Content Disappears, May 2024.
- Jonathan Zittrain, Kendra Albert and Lawrence Lessig, Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations, Harvard Law Review Forum, vol. 127, 2014.
- Martin Klein et al., Scholarly Context Not Found: One in Five Articles Suffers from Reference Rot, PLOS ONE, 2014.
- Shawn M. Jones et al., Scholarly Context Adrift: Three out of Four URI References Lead to Changed Content, PLOS ONE, 2016.