Finding Archived Web Content: Government Websites and Columbia University Pages

Web content changes constantly. Departmental websites are reorganized, URLs change, institutions and organizations rename themselves, and content may be revised or removed over time. The result is broken links and "404 Not Found" errors, which can make previously available information inaccessible.

Broken links are especially common on government and university websites, where periodic updates, redesigns, and site reorganizations can alter or remove previously available materials. Faculty listings, course descriptions, announcements, and program information that once existed on departmental sites can become difficult or impossible to retrieve through the current site. This article explains the primary tools for accessing this content.

The Wayback Machine (Internet Archive)

The Wayback Machine is your first stop for any archived web content. The Internet Archive’s Wayback Machine allows users to view archived versions of a URL over time.

How to use it: Enter any URL into the search bar. If the page has been archived, you will see a timeline of captures; click any date to view the page as it appeared at that moment.

A note on limitations: The Wayback Machine does not archive everything. Pages that appear and disappear quickly may not be recorded. The Wayback Machine is also poor at capturing pages that use JavaScript, other interactive programming, or legacy Flash animations. Similarly, audio, video, and form-submission data cannot be fully captured.

Tip: If a URL has changed over the years, you may only find captures going back to the most recent version of the URL. See the Archive-It section below for how to work around this.

United States government websites  

The following tools can help locate federal web content:

Wayback Machine: Allows users to enter a URL, including .gov websites, into the Internet Archive’s Wayback Machine to view archived versions of webpages over time. Helps locate content that has been changed, relocated, or removed.

End of Term Web Archive: Archives selected U.S. federal government websites and digital content during presidential administration transitions, including .gov websites, selected .mil domains, and selected social media content.

Data Rescue Project Portal: Provides access to rescued U.S. federal datasets and links to archived or relocated versions of data resources when available.

Federal Environmental Web Tracker (Environmental Data & Governance Initiative): Helps users identify changes to federal environmental websites and datasets, including content that has been removed, altered, or relocated, using the tracker and its documentation.

Electronic Records Archive (National Archives and Electronic Records Archives): Enables users to search for permanently preserved executive branch records and other federal agency materials.

CDC and federal public health resources

The following specialized resources are particularly useful for health researchers:

CDC Archive: Preserves CDC webpages and datasets through large-scale web crawls, including End of Term and Internet Archive captures.

ACA Signups CDC Page Index: Provides a curated, multi-part index of archived CDC webpages compiled from Internet Archive captures and helps users locate historical CDC content when original URLs are unknown or no longer accessible.

Clinical Guidelines via Professional Organizations:  Identifies access to clinical guidelines maintained by professional societies when federal versions are unavailable, relocated, or no longer accessible.

DataLumos: Provides access to a crowdsourced repository of U.S. federal datasets hosted by the Inter-university Consortium for Political and Social Research.

Archive-It and Columbia University Libraries

In 2006, the Internet Archive launched Archive-It, a subscription-based web-archiving service designed for institutional web preservation. Columbia University Libraries joined in 2010 and has since maintained an active web-archiving program covering Columbia.edu and its subdomains (e.g., cumc.columbia.edu).  

Although both Archive-It and the Wayback Machine are operated by the Internet Archive, they serve different purposes. Archive-It is designed for institutionally managed web archiving and preserves curated, policy-based web crawls at the domain or collection level rather than relying on broad, heterogeneous web crawling.

Columbia’s archived holdings can be searched directly through the Columbia Libraries Archive-It collection. This resource often preserves earlier versions of departmental and research center websites that are no longer accessible through Columbia's current site structure.

Strategies when you don’t have the exact URL

  • Start by checking whether the page is available in Archive-It or the Wayback Machine.
  • For Columbia sites, use Archive-It first to identify older or alternate URLs, then use those URLs in the Wayback Machine.
  • Navigate within archived sites by starting at a top-level domain (e.g., columbia.edu or cdc.gov) and following internal links.
  • Check whether content has simply moved by comparing the current site structure with archived versions.
  • Use Archive-It keyword search when URLs are unknown or difficult to reconstruct.

Preserving pages before they disappear

If you find a page you believe may be at risk, you can archive it yourself before it is taken down.

Save Page Now: This Wayback Machine feature lets you submit any URL for immediate archiving, free of charge and without an account.

Perma.cc: Developed by Harvard Law Library Innovation Center, Perma.cc allows users to create permanent links to archived web pages, making it particularly well-suited for citation and long-term reference.

Which tool should I use?

SituationRecommended Tool
Exact URL knownWayback Machine
Columbia University page (no URL)Columbia Libraries Archive-It
U.S. federal government pageEnd of Term Web Archive
Federal datasetData Rescue Project Portal
Need to preserve a pageSave Page Now or Perma.cc

Citing archived web pages

It has become standard practice to cite archived versions of web pages rather than the live URL because archived links remain stable even when the original page disappears. Most citation styles (APA, MLA, Chicago) accommodate archived URLs from the Wayback Machine. Check with your style guide of choice for the preferred format and use the specific timestamped Wayback Machine URL (e.g., https://web.archive.org/web/20240115154210/example.com) as your cited source.

Browser extension

The Internet Archive offers the Wayback browser extension for several browsers. It automatically alerts you when you land on a broken page and provides access to the most recent archived versions, making it a useful tool for researchers encountering dead links.

Was this article helpful?
What made the article not helpful?