Guides · updated 4 Oct 2026
The Wayback Machine: a website's history
How the Internet Archive's Wayback Machine captures websites, how to read its calendar, what a first capture means, and the gaps to watch for.
The Wayback Machine is a free service that lets you type in a web address and browse copies of that site as it looked on earlier dates. It is often the quickest way to answer a simple question about a domain: what was here before?
What the Internet Archive and the Wayback Machine are
The Wayback Machine is run by the Internet Archive, a non-profit digital library based in San Francisco. According to the Archive's own anniversary timeline, Brewster Kahle founded it in 1996, and its crawlers captured their first web pages in October that year.
The public could not browse the collection until the Wayback Machine was unveiled at the University of California, Berkeley on 24 October 2001. Its name is a nod to the WABAC ("way-back") machine in the Rocky and Bullwinkle cartoons.
In October 2025 the Internet Archive announced that it had passed one trillion archived web pages.
How captures happen
A capture (or snapshot) is a copy of one web address saved at one moment. Captures arrive in a few ways:
- Automated crawls. The Archive says hundreds of web crawls contribute captures every day, from its own crawlers and from crawls donated by Alexa Internet and others. Crawlers find pages by following links, so well-linked sites tend to be captured more often. Hovering over a capture shows a "why" link naming the crawl it came from.
- Save Page Now. Anyone can ask the Wayback Machine to save a single page using Save Page Now, a feature added in October 2013. It saves that one page once; it does not add the site to future crawls.
The Archive's help pages say there is a lag of 3 to 10 hours between a crawl and the capture appearing.
Every capture has a timestamp built into its address, in the form yyyymmddhhmmss. A capture address containing 20000229123340, for example, was taken on 29 February 2000 at 12:33:40.
How to browse the calendar
- Go to the Wayback Machine, type a domain such as
example.co.ukand press Enter. - Pick a year. The calendar shows a dot on each day that has captures.
- Check the colour of each dot. Blue means the server returned a 2xx (successful) HTTP status code, green a 3xx redirect, orange a 4xx client error and red a 5xx server error. Blue is usually what you want.
- Hover over a dot to see the capture times for that day, then click one to open it.
- To list every archived address the Wayback Machine holds for a site, use the pattern
web.archive.org//example.co.uk/.
Not every capture is complete. If something on a page was not captured that day, the Wayback Machine fills the gap from the nearest date it has, or even from the live web. Watch the date code in the address bar as you click around: one "page" can be stitched together from several dates.
What "first capture" means
The first capture is simply the earliest timestamp the archive holds for a domain. It is not the date the domain was registered, and not necessarily the date a website first went live. The two can differ in either direction.
It can come after registration. A new site may go unnoticed by crawlers for months, especially if few other sites link to it. Some domains are registered and left unused for years.
It can come before registration. The registry's creation date belongs to the current registration record. If a domain expired, was deleted and was later registered again, the new record normally has a new creation date, while the archive may still hold the earlier owner's pages. The two sources simply record different things. As a real example, when we checked on 4 October 2026, the earliest Wayback capture of nominet.uk was dated 18 July 2013, while the .uk registry's RDAP record gave a registration date of 10 June 2014.
So treat the first capture as a clue, not proof of age. Compare it with the registration date (see domain age) and open the earliest captures to see what was there.
Limitations
- Exclusions. Some sites are missing because a robots.txt file blocked the crawlers, or because the site owner asked for them to be excluded.
- Pages crawlers cannot reach. The Archive collects publicly available pages. Pages behind a password, or only reachable by submitting a form, are not collected.
- Dynamic content. JavaScript is often hard to archive. Anything that has to contact the original server to work, such as forms, search boxes or server-side image maps, will not work in the archived copy. Plain HTML archives best.
- Orphan pages. Crawlers do not type into search boxes, so pages with no links pointing to them are often never found.
- Gaps. Images can be missing, and captures may be years apart.
- Busy periods. In September 2026 the Archive explained that it had added protections against high-volume automated traffic, and that some real visitors were being blocked by mistake with HTTP 429 ("too many requests") errors. If the Wayback section of a report is temporarily empty, this may be why; try again later.
Asking for a site to be excluded
The Internet Archive's help centre says that site owners who want archives of their site excluded should email info@archive.org with:
- the URL or URLs concerned;
- the time period they want excluded;
- the period during which they controlled the site or account, if relevant;
- anything else that helps explain the request.
This starts a review, and the Archive makes no guarantee in advance about the outcome. Copyright complaints follow a separate process, described in the Wayback Machine FAQ.
For developers: the APIs
The Internet Archive offers two interfaces for developers:
- The Wayback Availability API answers a simple question: is this address archived, and what is the closest capture? You pass a
urland, optionally, atimestampof 1 to 14 digits (YYYYMMDDhhmmss). It returns JSON describing the nearest snapshot, or an empty result if there is none. - The CDX Server API gives lower-level access to the capture index, with filtering, collapsing and paging, in plain text or JSON. Its documentation is still labelled beta.
The Wayback Machine also supports the Memento protocol. Given the 2026 traffic protections, keep automated requests modest.
How Domain Statistics uses it
The Wayback Machine section of a Domain Statistics report summarises what the archive holds for a domain, including the earliest capture it returns, so you can then explore the captures yourself. If you are thinking of buying an expired domain, its archive history is one of the most useful checks you can make. Technical terms are explained in the glossary.
Sources
- anniversary.archive.org — Internet Archive timeline: Brewster Kahle founded the Archive in 1996; its crawlers captured their first web pages in October 1996; Wayback Machine unveiled at the University of California, Berkeley on 24 October 2001
- help.archive.org — what the Wayback Machine is; web archiving began in 1996 and was opened to the public five years later; crawl donations from Alexa Internet and others; name taken from the WABAC machine; "why" links show the crawl a capture came from; dynamic pages; password-protected and form-only pages not collected; copyright process
- help.archive.org — calendar dot colours (2xx blue, 3xx green, 4xx orange, 5xx red); yyyymmddhhmmss date code; Save Page Now saves one page once; 3-10 hour lag; reasons sites are missing (robots.txt, passwords, owner requests, unknown to crawlers); JavaScript, server-side image maps and orphan pages; nearest-date and live-web fill-in; listing pattern web.archive.org//site/; exclusion request by email to info@archive.org and what to include; no guarantee of outcome
- blog.archive.org — "2001 – The Wayback Machine is launched"; Save Page feature added October 2013
- blog.archive.org — Wayback Machine passed one trillion archived web pages in October 2025; non-profit library headquartered in San Francisco
- blog.archive.org — September 2026: protections against high-volume automated traffic; some real users blocked with HTTP 429 errors; contact info@archive.org if blocked in error
- archive.org — Wayback Availability JSON API: url parameter, optional timestamp (1-14 digits, YYYYMMDDhhmmss) and callback; empty archived_snapshots when nothing is found; Memento support; CDX API exists
- github.com — CDX Server API (labelled beta): index of captures, plain text or JSON output, filtering, collapsing, pagination
- rfc-editor.org — EPP domain mapping: crDate is the date and time the domain object was created
- web.archive.org — checked 4 October 2026: earliest capture of nominet.uk dated 20130718153214 (18 July 2013)
- rdap.nominet.uk — checked 4 October 2026: RDAP registration event for nominet.uk dated 2014-06-10
Facts checked on 4 Oct 2026. Rules and figures change — check the source if it matters.