Internet Archive
A nonprofit digital library that preserves and provides access to web pages, books, audiovisual media, software, and other cultural materials.
Last updated August 28, 2026
Overview
The Internet Archive is a United States nonprofit digital library and preservation organization founded in 1996 by Brewster Kahle. Its stated mission is to provide universal access to knowledge by collecting, preserving, and making available digital and digitized cultural materials. The organization is best known for the Wayback Machine, a web-archiving service that lets users retrieve historical captures of websites and web pages. It also operates or supports collections covering books, periodicals, movies, television programs, audio recordings, music, software, images, and other forms of born-digital and digitized culture. The organization emerged during the early public development of the World Wide Web, when Kahle argued that a library-like memory of the online world was necessary because web content could disappear or change without notice. The Internet Archive consequently developed large-scale crawling and storage systems intended to preserve successive versions of web pages. The Wayback Machine became a major public reference tool for researchers, journalists, historians, lawyers, technologists, and ordinary users seeking earlier versions of online material. Its captures can document changes to organizations, products, public statements, news coverage, and websites that are no longer available from their original hosts. Beyond web preservation, the Internet Archive operates a broad digital-library platform. The archive.org site provides access to public-domain and openly licensed material, user-contributed files, digitized media, and lending collections. The Open Library project offers a catalog and online reading interface intended to create a web page for every book, while Internet Archive book-scanning and lending programs make many digitized volumes available under the rules of the relevant collection. Archive-It provides web-archiving tools and hosting services for institutions that want to preserve selected websites and online collections under their own collecting programs. Its collections also include historical software and computer games, live music and other audio, films and videos, educational materials, images, and community uploads. Availability varies by item, copyright status, geographic restrictions, collection policy, and lending or access controls. The organization therefore combines public access with preservation infrastructure, metadata, format migration, storage management, and partnerships with libraries, archives, universities, cultural institutions, and other rights holders. The Internet Archive is not a conventional commercial technology company or content publisher. It operates as a nonprofit organization and relies on donations, institutional relationships, grants, service activity, and other support rather than public equity-market financing. Its public-interest role has made it an important part of the digital preservation ecosystem, but its activities have also generated legal and policy disputes. These have included copyright litigation concerning digital lending and objections from rights holders to the availability of archived or uploaded material. Such disputes reflect a continuing tension between preservation and access on one side and copyright enforcement and licensing models on the other. Today, the Internet Archive remains internationally used through archive.org and related services. Its importance derives from the combination of a large historical web corpus, a wide multimedia collection, public-facing discovery tools, and infrastructure designed to keep digital cultural records available over time.
History
The Internet Archive was founded in 1996 by Brewster Kahle as a nonprofit effort to preserve and provide access to the rapidly expanding body of digital information. Its early work centered on collecting web pages, addressing the problem that online material could disappear or change without leaving a stable public record. The organization developed automated crawling and storage systems and built a large historical corpus of web captures. The Wayback Machine became the Internet Archive's most recognizable service. Publicly available from 2001, it enabled users to enter a web address and inspect available snapshots from earlier dates. The service became useful across journalism, scholarship, cultural history, legal research, technical troubleshooting, and fact-checking. Because web pages are frequently edited or removed, the archive's captures can provide evidence of how public online information appeared at a particular time, although capture coverage is not complete and archived content may be subject to access restrictions. The organization subsequently expanded beyond web pages into a broad digital-library model. Its collections include digitized books, public-domain texts, films, television material, audio recordings, live music, software, games, images, and community-uploaded files. The Open Library project provides a large bibliographic catalog and online reading functions. Internet Archive book projects have worked with libraries and other partners to scan physical books, create searchable digital copies, and make selected titles available through lending or other access arrangements. Archive-It extended the organization's preservation model to institutions that wanted to collect and manage specific portions of the web. Libraries, universities, government bodies, museums, and other organizations can use institutional web-archiving workflows to define collections, preserve captures, and present them to users. This helped establish web preservation as a regular activity for organizations whose records increasingly exist online. The Internet Archive's growth has also brought recurring questions about copyright, licensing, privacy, and the responsibilities of a public digital repository. Its lending and access practices have been challenged by rights holders, particularly where digitized books were made available without conventional licenses. The COVID-19-era National Emergency Library intensified these disputes and was followed by litigation from publishers and authors. A later court ruling against the Internet Archive's controlled digital lending program placed limits on the organization's interpretation of library digitization rights. Despite these disputes, the Internet Archive continues to operate as a nonprofit digital library and preservation organization. Its services combine public discovery interfaces with large-scale storage, metadata, digitization, crawling, and preservation infrastructure. The organization remains particularly important because it preserves material that may otherwise be lost when websites close, files become obsolete, or online services change their policies.
- 2023Court rules against Internet Archive in book-lending case
A federal court found that the Internet Archive's controlled digital lending of books infringed publishers' copyrights in the litigation brought by major publishers.
- 2020National Emergency Library created and closed
The Internet Archive temporarily expanded digital-book access during the COVID-19 pandemic before ending the program ahead of schedule following criticism and litigation.
- 2006Archive-It introduced
Archive-It provided institutions with tools and hosted infrastructure for creating curated web archives.
- 2005Open Library project launched
The Open Library project began developing a freely accessible online catalog and reading platform for books.
- 2001Wayback Machine made publicly available
The Internet Archive opened public access to its historical web-archiving service.
- 1996Internet Archive founded
Brewster Kahle founded the Internet Archive as a nonprofit digital preservation and access organization.
Products and positioning
A nonprofit, public-interest digital library and preservation infrastructure provider focused on long-term access to online and digitized cultural heritage.
Wayback MachineWeb archive2001
The Wayback Machine is a public interface to the Internet Archive's web-capture collection. Users can search for websites and, where captures exist, view earlier versions of pages across different dates. It is used for historical research, journalism, citation checking, preservation, and recovery of information that has disappeared from the live web. Coverage depends on crawling, robots directives, technical accessibility, and other collection constraints.
Open LibraryDigital library and book catalog2005
Open Library is an Internet Archive project that aims to create an openly accessible web page for every published book. It combines bibliographic records, editions, metadata, covers, links, and reading features. Some books can be read directly, while others are represented primarily through catalog information or links to external and library resources.
Archive-ItInstitutional web archiving2006
Archive-It is a hosted web-archiving service for libraries, universities, government agencies, museums, and other organizations. It supports collection selection, web crawling, metadata management, preservation, and public presentation of archived online materials. The service allows institutions to preserve their own thematic or organizational records while using Internet Archive infrastructure.
Internet Archive BooksDigitized books and lending
Internet Archive Books provides access to digitized books and periodicals from library partnerships, public-domain collections, and other sources. Depending on rights and collection rules, users may download a file, read it online, or borrow it through controlled access. The project combines scanning, optical character recognition, metadata, preservation storage, and public discovery.
Internet Archive AudioAudio archive
The audio collections include music, spoken-word recordings, podcasts, radio material, field recordings, and live concert recordings. Availability and permitted use vary by item. The collection serves both as a public listening resource and as a repository for audio files contributed by users, artists, institutions, and community projects.
Internet Archive Software CollectionSoftware and games archive
The software collection preserves computer programs, operating systems, games, utilities, and related digital artifacts. Some titles can be run through browser-based emulation or accessed as downloadable files, subject to rights and technical limitations. It supports research into software history, computing culture, and the preservation of obsolete formats.
Flagship businesses
- Wayback Machine
- Open Library
- Archive-It
- Internet Archive Books
- Internet Archive Audio
- Internet Archive Software Collection
Marketing campaigns
- 2020National Emergency Library
Global
During the COVID-19 pandemic, the Internet Archive temporarily removed waiting lists from parts of its digitized book lending program to support students and readers who could not access physical libraries.
Outcome. The program was closed earlier than initially planned after strong criticism from rights holders and the filing of copyright litigation.
Brand decisions
- 2023Continue defending controlled digital lendingStrategy
Publishers challenged the Internet Archive's lending model for digitized books in federal court.
What changed. The organization continued to defend its library and preservation position through the litigation and public policy debate.
Aftermath. The court ruled against the Internet Archive on the challenged lending practices, creating a major legal constraint on its book-access model.
- 2020Expand book access during the pandemicStrategy
Physical library closures during COVID-19 restricted access to printed books and increased demand for remote educational resources.
What changed. The Internet Archive created the National Emergency Library by temporarily changing access conditions for portions of its digitized book collection.
Aftermath. The decision prompted copyright criticism and litigation, and the program ended before its announced end date.
Leadership
| Name | Title | Tenure |
|---|---|---|
| Brewster Kahle | Founder and Digital Librarian | 1996– |
Controversies
- 2023Controlled digital lending rulingControversy
A federal court ruled against the Internet Archive in a case concerning its lending of digitized books, finding that the challenged practices infringed publishers' copyrights.
- 2020National Emergency Library copyright disputeControversy
Publishers and authors criticized the temporary lending expansion as an unauthorized distribution of copyrighted books, leading to litigation and the program's early closure.
Recent events
- 2024Internet Archive reports cyberattack and service disruption
The organization reported a cyberattack and associated service interruptions affecting parts of its public services. It communicated restoration and security measures through its public updates.
Other - 2001Internet Archive launches the Wayback Machine
The Internet Archive made its historical web-capture service publicly available as the Wayback Machine, allowing users to consult preserved versions of websites.
Product launch
Sources
Cite this profile: Cite the canonical profile. /brand-wiki/internet-archive · Editorial policy · How profiles are compiled