Archivepanel

Link rot by the numbers: how much of the web disappears

A quarter of the web pages that existed between 2013 and 2023 are gone, and half the links in US Supreme Court opinions no longer show what they cited. The research on link rot and content drift, and what it means for anything you cite.

Every link is a promise that someone else's server will keep saying the same thing. The research says that promise is broken often, and faster than most people assume.

This page collects the most-cited studies of link rot and content drift, with what each one measured, so you can quote them accurately.

The short version

Finding Source
38% of web pages that existed in 2013 were no longer accessible a decade later Pew Research Center, 2024
25% of all pages that existed at some point between 2013 and 2023 were no longer accessible Pew Research Center, 2024
Even among pages from 2023, 8% were no longer accessible by late 2023 Pew Research Center, 2024
23% of news pages and 21% of government pages contain at least one broken link Pew Research Center, 2024
54% of Wikipedia articles have at least one broken link in their references Pew Research Center, 2024
50% of links in US Supreme Court opinions no longer show the cited material Zittrain, Albert and Lessig, 2014
More than 70% of links in three Harvard law journals no longer show the cited material Zittrain, Albert and Lessig, 2014
One in five science, technology and medicine articles suffers from reference rot, rising to seven in ten among those that cite web pages Klein et al., 2014
For more than 75% of web references in scholarly articles, the content has changed since it was cited Jones et al., 2016

Link rot is the familiar one: you follow a link and get an error, a parked domain or a redirect to a home page. It is annoying, but at least it is obvious.

Content drift is worse. The link still works, but the page says something different from what it said when it was cited. Nothing tells you. A reader who follows the link reasonably assumes they are seeing what the author saw.

For anyone who uses web pages as evidence, drift is the dangerous one. A terms page that returns an error is a missing exhibit. A terms page that has been quietly rewritten looks like a working link to the version the other side prefers.

How much of the web disappears

The largest recent study is Pew Research Center's 2024 report, When Online Content Disappears. Pew sampled pages from Common Crawl snapshots taken each year from 2013 to 2023 and checked whether they could still be reached in late 2023.

  • 38% of pages that existed in 2013 were no longer accessible.
  • 25% of all pages collected between 2013 and 2023 were no longer accessible.
  • Even recent pages go: 8% of pages from 2023 had already disappeared.

The loss isn't limited to obscure corners of the web. Pew found that 23% of news pages and 21% of government pages contained at least one broken link. On Wikipedia, 54% of articles had at least one link in their references pointing to a page that no longer exists, and 11% of all references were no longer accessible.

Citations in law and science

In 2014, Jonathan Zittrain, Kendra Albert and Lawrence Lessig studied links cited in legal writing. They found that 50% of the URLs in US Supreme Court opinions suffered reference rot, meaning they no longer produced the information originally cited. In the Harvard Law Review, the Harvard Journal of Law and Technology and the Harvard Human Rights Journal, the figure was more than 70%.

The same study also measured drift directly. Of the Supreme Court links that still loaded successfully, only 76% still led to the material cited. A working link was no guarantee of the right content.

In science, Klein and colleagues (2014) examined over a million references in almost 400,000 science, technology and medicine articles published between 1997 and 2012. One in five articles suffered from reference rot. Among articles that referenced web resources at all, it was seven in ten.

A follow-up by Jones and colleagues (2016) looked at drift specifically. For more than 75% of references to web resources, the content had drifted away from what it was when it was cited, even where the link still resolved.

Why pages disappear

The studies agree that most loss is ordinary, not sinister:

  • Redesigns and CMS migrations change every address on a site, and old addresses are rarely redirected one by one.
  • Companies and projects end. Domains lapse and are re-registered by someone else.
  • Content moves behind logins and paywalls, or into apps.
  • Pages are edited in place. Terms, prices, policies and product claims are updated without keeping the old version, and there is usually no public history.

The last one is content drift, and nothing about it is unusual. Businesses change their terms, and they're under no obligation to publish what the old ones said.

What to do about it

If you cite for readers, such as in a paper, an article or a brief, pair each important link with an archived copy. Perma.cc, created for legal citations after the Zittrain study, and the Internet Archive's Save Page Now both create a public, citable archived copy.

If you may need to prove what a page said, a public archive isn't enough. You don't control when it captured the page, whether it captured it at all, or whether the copy stays available. We go into this in Wayback Machine captures as evidence. Keep your own copy of the pages that matter, with an independent timestamp, at the moment they matter. Our guide on how to preserve a web page as evidence has a checklist.

If the page is likely to change, one capture isn't enough. Capture it on a schedule, so that when it changes, the version from before is already on file.

Sources

More from the blog

Keep the version that matters.

Start free with 50 MB of storage. No card required.