Skip to content
Tech News
← Back to articles

As AI eats the web, the internet’s collective memory is disappearing

read original more articles
Why This Matters

This article highlights how AI-driven search engines are contributing to the loss of the internet’s collective memory by making it harder to access accurate and original information. As AI summaries and hallucinations become more prevalent, both consumers and the tech industry face challenges in maintaining reliable access to factual data, risking a decline in the integrity of online knowledge. This shift underscores the importance of preserving original sources and improving AI accuracy to sustain trustworthy digital information.

Key Takeaways

It turns out a lot of people ask Google when the sun will set. It’s a good fact to know. Getting a ballpark on twilight can help you plan a bike ride, barbecue, or a perfectly timed promposal. The occurrence lasts only between two and a half and four minutes, and it happens dynamically depending on your proximity to the equator, making it a tricky event to keep tabs on.

Recently, sunset chasers have been missing the main event because Google’s new AI summaries started inventing times. “I had the projector set up outside and was waiting for the sun to set,” wrote one Facebook user in Colorado Springs, “but to my surprise I was simply living in the past. AI informed me the sunset had already happened.” These episodes point to something many have noticed: the world’s dominant search engine now struggles over basic facts. The company that built its reputation returning the right answer has, as one tech expert observed, “lost its edge.”

Still, we keep googling because we keep expecting the internet to know things. That stubborn hope underpins almost every argument about information access. When search worsens, we grumble that Google has been enshittified. When AI hallucinates, we say the model is sloppy. When misinformation spreads, we reassure ourselves the truth is still “out there,” waiting patiently for us to string together a better query. Even as the web is polluted by AI slop, we cling to the belief it’s a skills problem and share search tips.

But what if the “truth” is harder to find online because the infrastructure that once stored it is breaking down? Some of that is wear and tear. Link rot erases pages every day. Key sections of the United States Constitution briefly disappeared from the Library of Congress website because of a coding error. But by interposing an error-prone AI between us and an original source, Google has ensured that if an underlying page exists, it can become practically undiscoverable. Meanwhile, pollution is now moving upstream: 404 Media reports companies are planting content on Reddit to influence the answers generated by AI search, contaminating the public record from which those systems draw their summaries.

The corpus is collapsing in real time. This degraded digital space should force a broader understanding of cultural sovereignty. Right now, the debate is framed largely as a fight over whether foreign mega platforms like Netflix, YouTube, or Spotify should support Canadian content. That charged conversation sidesteps the deeper question: Who preserves and controls access to our cultural record—not just in the present but for generations after?

Search can no longer pretend to be a neutral gateway to a stable body of knowledge. While the web has always been organized around intermediaries that shape what survives online and who sees it, the internet’s archival function is today breaking down under relentless pressure from stakeholders with very different—and often conflicting—priorities. At least book banning happens in your face: it has villains, school board meetings, sensational headlines. Digital erasure is more insidious and, in some ways, more devastating. A banned book can still be found. A scrubbed webpage can disappear so completely few people would ever realize it was there.

Consider what happened to FiveThirtyEight, an American website that focused on opinion poll analysis, politics, economics, and sports blogging. The Walt Disney Company had owned the blog through its subsidiaries for more than a decade. Its founder, Nate Silver, left in 2023, and by March 2025, the remaining staff had been laid off. Once Disney concluded the site was no longer an active asset—no new journalism and no meaningful advertising revenue—it straight up deleted nearly the entire archive.

The retrieval crisis has reached even Wikipedia, one of the world’s most significant volunteer-run public knowledge resources. For years, search engines sent billions of viewers to its pages. But now, AI systems scrape and ingest Wikipedia’s content directly when presenting their results, eliminating the need for users to click through. Wikipedia has become the infrastructure of its own demise: dwindling traffic means attention and donations no longer reliably flow back to the encyclopedia to keep it alive.

What has been happening to the Internet Archive is just as alarming. A non-profit digital library, it collects and makes available materials that might otherwise vanish—from audio recordings to historical documents to books. It is best known for the Wayback Machine, a service that has captured hundreds of billions of snapshots of the web at different points in time. Journalists, researchers, historians, lawyers, and the public use it to recover deleted pages, verify past statements, and document changes to online content. The Wayback Machine is the closest thing the web has to a fail-safe backup memory.

The overall project, however, is not only buckling under the engineering strain of indexing and storing an ever-growing repository of information but it is also being battered by cyberattacks and costly litigation. After publishers successfully sued the Internet Archive over its digital lending program, calling it unauthorized copying, news organizations started blocking the Wayback Machine’s crawlers out of fear that archived pages can provide AI companies with an indirect source of copyrighted material. Each new restriction limits the archive’s ability to act as a comprehensive backstop.

... continue reading