The internet feels permanent until the page you need returns a 404.
Every broken link is a small data loss event.
A citation stops working. A government document disappears. A company quietly rewrites a policy. A useful forum closes. A personal website vanishes because its owner stops paying for the domain.
We usually treat these as isolated failures. They are not. They are the normal behaviour of a medium built to publish information, but not necessarily to preserve it.
The Internet Archive exists because the web does not remember itself.
A Library Built for Digital Material
The Internet Archive is a nonprofit digital library founded in 1996. Its best known service is the Wayback Machine, which stores captures of websites and allows us to revisit them as they appeared at different points in time.
Its collections extend far beyond webpages. They include digitised books, films, television broadcasts, radio programmes, live concert recordings, photographs, magazines, academic papers, manuals, old software, computer games, and material uploaded by individuals and institutions.
In October 2025, the Wayback Machine passed one trillion archived webpages.
That number is almost impossible to picture. Those captures include abandoned blogs, deleted news articles, old government pages, discontinued documentation, company announcements, community websites, and personal projects that may no longer exist anywhere else.
The Archive is often described as a collection of old internet pages, but it is an infrastructure for recovering information after the original source has failed.
The Web Is an Unstable Medium
A printed book can survive for centuries without requiring a software update, an active domain name, or a monthly hosting payment. A webpage depends on a domain, a server, a database, working software, security updates, external services, and somebody continuing to pay the bills. Modern sites may also depend on JavaScript frameworks, embedded media, authentication systems, advertising networks, APIs, and cloud platforms controlled by other companies.
Remove one of those dependencies and part of the site may stop working. Remove several and the whole thing can disappear.
The Pew Research Center examined webpages that existed between 2013 and 2023. By October 2023, roughly one quarter were no longer accessible. Among pages that existed in 2013, 38 percent had disappeared within a decade.
Link rot also damages material that remains online. Pew found broken references on news sites and government pages, as well as on more than half of the Wikipedia pages it examined.
A hyperlink may look like a permanent citation, but it is really a dependency on somebody else’s server.
Without an archive, much of the web has the historical reliability of a whiteboard.
History Needs Version Control
Programmers use version control because the current version of a file does not tell the whole story. The same principle applies to public information.
A government department can update a page without keeping the old wording available. A company can revise its terms of service. A politician can delete an old promise. A news organisation can alter an article after publication. A product description can change after customers have already bought the product.
The live page tells us what an organisation says now. An archived copy may show what it said before. That makes the Wayback Machine useful to journalists, researchers, historians, lawyers, and anyone else trying to establish what appeared online at a particular time.
An archive does not decide which version is correct. It preserves the evidence needed to ask the question. Accountability becomes difficult when yesterday can be silently overwritten.
The Small Web Is Part of the Record
The historical value of the Internet Archive is not limited to national newspapers, government agencies, or universities.
The web was built from millions of smaller contributions: personal homepages, local news sites, hobby forums, technical tutorials, community organisations, fan sites, and projects created by people who never expected their work to become historically important. Those sites often preserve material that larger institutions never collected.
They contain local photographs, obituaries, community announcements, newsletters, sermons, meeting recordings, oral histories, technical solutions, and first-hand accounts that may exist nowhere else.
In my own work, the Internet Archive has been more than somewhere to browse old websites. It has helped me locate older recordings, organise conference audio, recover forgotten material, and provide access to files that might otherwise remain on a single ageing hard drive.
That experience changed how I think about preservation.
Not everything worth saving was produced by a large institution. Some of the most irreplaceable material is precisely the material nobody thought was important enough to save.
A local recording may have little commercial value, yet contain the only surviving message from a particular event. A small church website may seem ordinary until it becomes the last accessible record of an assembly that has closed. A personal blog may preserve how someone understood an event before historians and institutions shaped the accepted version.
History is not made only from official records. It is also assembled from the ordinary things people leave behind.
Software Is Culture Too
Software is often treated as disposable.
A programme is released, updated, replaced, and eventually made incompatible with newer systems. Games disappear when authentication servers are shut down. Mobile applications vanish from app stores. Online services cease to exist when the company behind them closes.
Yet software records how people worked, communicated, created, and played during a particular period.
The Internet Archive preserves old programmes, games, shareware collections, manuals, magazines, screenshots, and documentation. Many programmes can even be run directly in a browser through emulation.
A screenshot can show what an old programme looked like. A working copy can show how it behaved.
There is a difference between preserving a photograph of a tool and preserving the tool itself.
Future historians should be able to study more than the advertisements and screenshots left behind by the computer age. They should be able to experience the software, understand its limitations, and see how people actually interacted with it.
Without deliberate preservation, huge portions of digital culture can vanish within a single hardware generation.
A Library Must Be Able to Keep What It Lends
The Internet Archive’s legal battles exposed a larger problem: digital media has changed what it means for a library to own something.
A traditional library can buy a printed book and keep it for decades. It can lend that copy repeatedly, repair it, move it to another branch, or preserve it after the publisher stops selling it.
Digital books often operate under different rules.
A library may receive a limited licence instead of ownership. Access may expire, require approved software, be restricted to a fixed number of loans, or disappear when a vendor changes its terms. A book can appear to be part of the library’s collection while remaining under the publisher’s technical and contractual control.
When a library owns a physical item, it can preserve it.
When it rents access to a digital item, preservation depends on the company providing the licence.
The Internet Archive operated a programme known as controlled digital lending. It scanned books it physically owned and generally limited circulation so that only one corresponding copy (physical or digital) could be loaned at a time.
During the COVID-19 shutdowns, it temporarily removed those waiting list restrictions through the National Emergency Library, saying the change was intended to help people who had lost access to physical libraries.
Four major publishers sued.
In September 2024, the United States Court of Appeals for the Second Circuit upheld a lower court ruling that the Archive’s lending of the disputed books was not protected as fair use. The Internet Archive later decided not to seek review by the Supreme Court and agreed to continue removing books from lending when requested by publishers covered by its agreement with the Association of American Publishers.
The litigation resulted in the removal of public lending access to more than 500,000 digitised books.
The Archive also faced a separate lawsuit from major record labels over the Great 78 Project, which preserves fragile 78-rpm recordings. That case ended in a confidential settlement in September 2025.
By November 2025, an Internet Archive spokesperson said the organisation faced no major active lawsuits and no immediate legal threat to its collections.
But the cases showed how vulnerable a nonprofit library can become when preservation, copyright law, and commercial licensing collide. They also left lasting damage. A lawsuit can end while the material removed because of it remains unavailable.
There are legitimate disagreements over copyright, licensing, authors’ rights, and how digital lending should work. Pretending those questions are simple does not help anyone.
But there is also a broader question that cannot be ignored:
What happens to culture when libraries are no longer able to keep independent copies of what they provide?
The New Threat: Blocking the Archivists
The next major problem is already taking shape.
Publishers are understandably concerned that artificial intelligence companies may scrape their journalism without permission or payment. Reporting costs money. Writers and publishers deserve meaningful control over how commercial systems use their work.
In response, some publishers have begun blocking automated crawlers.
Unfortunately, those restrictions do not always distinguish between commercial AI systems and preservation tools operated by the Internet Archive.
A May 2026 investigation by the Nieman Journalism Lab examined 382 news websites that restricted at least one Internet Archive affiliated crawler. Of those, 342 were local news sites.
That is not every news website, and it does not mean every page from those sites has already vanished from the Archive. It does show, however, that a significant portion of local journalism is becoming harder to preserve.
Local news is especially vulnerable. Newspapers close, merge, change owners, replace publishing systems, and migrate material onto new platforms. Many do not maintain complete public archives. Some have never properly digitised their older reporting.
When one of those websites disappears, there may be no paper archive, institutional repository, or reliable backup waiting somewhere else.
The concern about AI scraping is legitimate. But blocking a preservation crawler does not necessarily stop a well funded AI company. It may simply prevent a library from preserving the public record.
Commercial extraction and public interest preservation are not the same activity.
Treating every crawler as identical may be convenient, but it risks creating permanent holes in our historical record.
The answer should include better access controls, sensible rate limits, agreements with publishers, and technical methods for distinguishing preservation crawlers from bulk commercial harvesting.
The answer cannot simply be to stop recording history.
Preservation Is an Operating System, Not a Backup File
It is tempting to imagine the Internet Archive as a very large pile of hard drives.
In reality, preservation is an ongoing process.
Files must be copied, indexed, checked for corruption, moved onto new hardware, and kept readable as formats and software change. Services must remain online. Security systems must be maintained. Accounts must be protected. Multiple copies must exist in different locations.
In October 2024, the Internet Archive was hit by distributed denial-of-service attacks and a security breach that exposed patron email addresses and encrypted passwords. The organisation temporarily took services offline while rebuilding and strengthening its systems. It reported that the archived collections themselves remained safe.
The incident demonstrated an important point: storing information once is not the same as preserving it.
An archive is not a backup forgotten in a drawer. It is a live system designed to keep old data usable.
That system costs money.
The Internet Archive does not rely on behavioural advertising or place its main collections behind a subscription. It still needs staff, buildings, servers, storage hardware, bandwidth, electricity, security, and replacements for equipment that eventually fails.
It also needs people capable of maintaining old formats, operating massive storage systems, responding to legal requests, repairing broken services, protecting accounts, and keeping the Wayback Machine useful as the web becomes increasingly complicated.
Most internet companies are built to sell advertising, subscriptions, software, data, or access to a marketplace.
The Internet Archive’s product is continued public memory.
There is no simple commercial model for remembering everything on behalf of everyone.
One Copy Is Not Preservation
Anyone who has worked with old tapes, discs, photographs, or hard drives eventually learns the same lesson: one copy is no copy.
A recording stored on one cassette can be lost when the tape stretches or mold appears. A file stored on one external drive can disappear when the drive fails. A website hosted on one server can vanish when the account closes.
The Internet Archive is an important part of preservation, but it should not become our excuse for neglecting our own copies.
Material worth preserving should exist in more than one place, under more than one person’s control, and preferably on more than one kind of storage.
The Archive is not a substitute for local backups. Local backups are not a substitute for an independent archive.
They strengthen each other.
How to Help
Supporting the Internet Archive does not require becoming a professional archivist.
Use It
Search the collections. Explore an older version of a website. Borrow books that remain available. Listen to recordings. Run old software. Share useful material with other people.
A library becomes easier to defend when people understand what it provides.
Save Important Pages
The Wayback Machine’s Save Page Now tool lets anyone submit a public webpage for preservation.
It normally saves the specific page submitted rather than crawling an entire website, so save the exact articles, documents, announcements, and references that matter.
Do not assume somebody else has already done it.
When publishing an article of your own, consider archiving the sources you link to. A reference is much more useful when a copy still exists five or ten years later.
Preserve Your Own Work
Keep local copies of your website, writing, photographs, recordings, databases, and project files.
Use common formats where practical. Include dates, names, descriptions, and other metadata. A file can survive perfectly while becoming nearly useless because nobody knows what it contains.
Where you own the material and are comfortable making it public, consider uploading a copy to the Internet Archive.
A public website should never be the only copy of your work.
Donate
The Internet Archive accepts donations to support its staff and infrastructure. Recurring donations can be especially useful because storage, electricity, bandwidth, staffing, and security are recurring expenses.
Canadian donors should note that Internet Archive Canada is a not-for-profit corporation but is not currently registered as a Canadian public charity. A contribution should not be assumed to qualify for a Canadian charitable tax credit.
Public Memory Is Infrastructure
We tend to notice infrastructure only when it stops working. We notice the power grid when the lights go out. We notice backups when a drive fails. We notice archives when the document we need has disappeared.
The Internet Archive is part of the web’s public infrastructure. It gives researchers a record to study, journalists a source to verify, programmers a way to recover old documentation, and ordinary people a chance to revisit work that would otherwise be gone.
It is not perfect. It cannot capture every page, script, video, database, or interactive service. Some archived pages are incomplete. Some files are missing. Modern websites are increasingly difficult to reproduce outside the systems that originally generated them.
But the alternative to an imperfect archive is not a perfect archive. It is no archive at all.
We are producing more information than any earlier generation while storing much of it on systems designed around short term traffic, commercial access, and continuous replacement. A civilisation can record almost everything it does and still leave very little behind if nobody accepts responsibility for keeping the record.
The Internet Archive matters because memory should not depend entirely on whether yesterday remains profitable.
Use it. Save the pages that matter. Preserve your own work. Support the people maintaining the machines.
The web forgets by default. We do not have to let it.
Sources and Further Reading
- Internet Archive: General Information
- Internet Archive: One Trillion Web Pages Archived
- Pew Research Center: When Online Content Disappears
- Internet Archive: End of Hachette v. Internet Archive
- Internet Archive: Response to the Appellate Opinion
- Internet Archive: An Update on the Great 78s Lawsuit
- Ars Technica: The Internet Archive’s Legal Fights Are Over
- Nieman Journalism Lab: More Than 340 Local News Outlets Are Limiting Internet Archive Access
- Electronic Frontier Foundation: Blocking the Internet Archive Will Not Stop AI
- Internet Archive: FAQ on Publishers Blocking the Wayback Machine
- Internet Archive: Services Update Following the 2024 Attack
- Internet Archive: Save Pages in the Wayback Machine
- Internet Archive: Donations and Gifts FAQ
- Internet Archive: Donating to Internet Archive Canada