The Real Dead Internet Theory
You may not know what the open web is, but you'll miss it when it's gone

Reddit used to have one of the most open APIs on the Internet. Until 2023, API access was free with a fairly generous rate-limit. The platform made itself scrapable, with every page having an alternate .json twin that returned the same contents on an unauthenticated endpoint. This paved the way for easy to use third party tooling, like PRAW, a popular python API wrapper, or Pushshift, a publicly searchable archive of every post, across every subreddit, that published dumps monthly as torrents. For fifteen years, Reddit’s API quietly underwrote a significant amount of computational social science.
In 2023, this started to rapidly change. First, the platform announced that API access would now have a paid tier, calling out data collection for machine learning as an unauthorized use case that needed to be prevented. In 2024, leading up to an IPO, Reddit listed data licensing as a revenue vertical. In June 2025, they sued Anthropic. In Aug 2025, they blocked access by the Internet Archive’s Wayback Machine. By 2026, open API signups had completely closed. Access now requires manual approval under Reddit’s Responsible Builder Policy — a name that manages both moral performance and bureaucratic deflection at once. For researchers, a separate program called Reddit for Researchers and requiring official institutional affiliation is the only path. Community projects like Pushshift have either shut down, stopped syncing, or themselves developed gates. This is deeply ironic, as Reddit was at an early stage co-owned and architected by Aaron Swartz, a famous freedom of information advocate, who later committed suicide after being grotesquely scapegoated for running a script to bulk download articles from JSTOR. And unfortunately, it’s a case study in where the open web is going.
What exactly is the open web? It’s a system of protocols and standards, from TCP to HTTP to HTML, that operates in a decentralized, interoperable way. It is what is surfaced by Google, but it is not Google. It is your website. It’s my website. It’s Wikipedia, and online magazines and every blog ever created, and until very recently it was Reddit. The most important part of the open web is that it’s open. You do not need to log in to see what is on the open web. You do not need to use a particular app, platform, or device to create something there (yes, Reddit has logins, but that was a handwave before 2023). A server is just a weird computer, and though most people pay for hosting, with some fiddling, you do not need to. The most basic layers of the open web are not complicated: I teach HTML workshops sometimes, and you can impart enough knowledge for someone to edit a simple website template in an afternoon. Yes, it will not look good on mobile and the divs may not be centered, but it will render and perhaps even be charming.
You may not think very often about websites. There is a whole apparatus of standards bodies and network engineers that keep things running so you don’t have to. Websites are mundane. Indeed, so much so that I had doubts about whether anyone would want to read a column about them. Why care about websites in 2026? The online experiences engendered by them are so commonplace that it may blind you to their significance. Do I need to do the silly dance where I remind you what the world was like before the Internet? Why Aaron Swartz wrote the Guerilla Open Access Manifesto, and was scraping JSTOR in the first place? We are now several generations of social and economic problems downstream, but don’t let it obscure the majesty that for a brief, brilliant moment, we really did think information could be free.
Instead everywhere, the open internet is closing. Towards the end of 2025, the New York Times started hard-blocking the Internet Archive’s crawlers. Historically, a robots.txt file was used to opt-out of being crawled, but this was a request, not a wall. Now some news sites like the NYT have put more robust prevention methods in place. The blocking is also widespread; A Nieman Journalism Lab report investigated the robots.txt files of 1167 news sites and found that “241 news sites from nine countries explicitly disallow at least one out of the four Internet Archive crawling bots.” A similar later analysis found 382. It’s true that the Internet Archive is not the be all and end all of the open web, but the requests it makes as it crawls sites encompass the same open protocols the rest of the open web works on. It is emblematic of the access to information dream the early Internet was born into. Not only that, despite what news outlets today might think, it’s often a crucial source for tomorrow’s journalists.
AI is specifically implicated in the blocks. A NYT spokesperson went on record, saying, “We are blocking the Internet Archive’s bot from accessing the Times because the Wayback Machine provides unfettered access to Times content — including by AI companies — without authorization.” You thought LLM training runs using your data without permission was the problem? The cure may very well be worse than the disease. And it’s true that the Internet Archive seems to be a frequent source for those building training datasets. A 2023 Washington Post analysis found that the Wayback Machine was the 187th most commonly represented web domain (out of 15 million) in a dataset used to train Google’s T5 and Meta’s LLaMa. And a documented mass-scraping incident in 2023 was enough to briefly take the Internet Archive offline.
Why are these news sites blocking Internet Archive? Do they hope to license their dataset, the first step of which is blocking the channels through which it could be drunk up for free? Is it hostility towards AI for other reasons, maybe AI-threatened job loss, or the existential crisis of generated fake news? Sometimes the reasons may be more complex — as the CTO of the Baltimore Banner, which blocks IA but not scrapers used by ChatGPT and Claude, explained, their concern is that if their content is consumed via a third party like Internet Archive, it might lose proper attribution. The Internet Archive itself is perhaps not free from incentives either: After the 2023 scraping event that resulted in an outage, Mark Graham, head of the Wayback Machine said, “We got in contact with them. They ended up giving us a donation.” The hunger of every AI lab for large swathes of organic free-range written language has turned every text-based archive into an unlatched chicken coop. But it turns out there’re foxes on all sides.
Navigating the Internet has also been fundamentally transformed in the last two years, in ways which threaten the open web and the business models of news sites. A recent analysis by SparkToro, an online audience-research firm, titled “In 2026, Less than One Third of Google Searches Still Send a Click” traced — alongside the headline finding — a declining clickthrough rate since 2016. AI search summaries are widely blamed for the phenomenon, and a 2025 Pew Research Center report corroborates that searches with an AI summary result in roughly half the clickthroughs of those without (8 vs 15 percent). But to focus too much on the AI search summary alone is to miss a big part of the story, which is the rise of AI apps themselves — a significant fraction of LLM users are looking up information that might have previously been a Google search — or turning to in-platform search on apps like TikTok and Instagram. This trend is harder to quantify, but we can look at ChatGPT’s runaway growth, as the fastest app to ever reach a billion monthly active users. Or the recent comment by Eddy Cue, an executive at Apple, testifying that in April 2025, search traffic on Safari had dropped for the first time in 22 years. He blamed AI.
The SparkToro piece, which seems sincere in its attempt to quantify the trend despite being ultimately the sales blog of an online marketing platform, points to the rise of “Zero Click Marketing” aka, marketing tactics that still work, even when no one clicks on your website at all. It advises building a presence on platforms you don’t own (in Apps) and shifting to thinking of your website primarily as something that will be indexed via LLMs and regurgitated from training data. This already even has its own liturgy: llms.txt, the opposite of the classic robots.txt, that explicitly invites LLMs and instructs them on what to look at first. Roughly 10% of websites implement it. The problem is no one seems to be looking: One monitoring firm’s analysis of more than half a billion AI crawler requests found only 408 requests for llms.txt. Is this the future of the open web? A choice between platforms like fortresses pulling up the drawbridge, and the websites like supplicants, trying to give keys to someone who’s already taken the town?
There are movements that hope otherwise. Earlier this year, 2026 was being branded the year of “going analog” by a number of influencers, mostly Gen Z (and somewhat ironically spreading the message via Instagram reels and YouTube). Some of this looked like a turn towards offline single use devices, like digital cameras and MP3 players. Some of it looked like adoption of dumb phones. Some of it encouraged people to make handmade websites, on platforms like Neocities. In the last year a number of magazines and platforms have cited Neocities as proof the Internet can still be a place for art, creativity, and weird websites, from The Verge to Plaster Magazine. Neocities is indeed a refreshing countertrend, though with a little over 1.6 million sites, it is still tiny in relative terms. Another touchstone for the analog movement might be The Luddite Club, a no-smartphones irl meetup founded by a group of New York teenagers in 2021, with chapters now in over 30 cities, from Seattle to Stockholm. Culturally, we’ve obviously reached peak smart phone. And while I actually support all of this to a degree, and the Neocities user is indeed doing something meaningful to continue the craft of the open web, it seems equally true that encountering something (a website, a dumb phone) through the lens of nostalgia itself indicates the death of the form. These trends can claim novelty only when the thing they imitate is no longer a primary form. And none of this will be enough to halt the ongoing enclosures. In the 2010s, my generation revived cassette tapes but it did nothing to halt the rise of Spotify.
The threats to the open web are obviously structural. They began long before LLMs, though they have been massively accelerated by them. I would place the antecedent in the dominant business models of the open web. Namely, a conflation of open protocols and freedom of information with ad revenue as the main, or only, monetization path. This is how various news platforms ended up so structurally vulnerable to a declining click through rate, isn’t it? It’s also what burned out a generation (mine) of online creative people from organizing indie online spaces. Maybe the transformation charted here is inevitable — we are always creating new commons to enclose after all — but I think it bears asking instead, what would it take to be different? The counterfactual — a Reddit that kept the basic API open but continued to monetize through paid features, premium tiers, in platform marketplaces, etc — is certainly possible to imagine. The tactics needed are twofold: an analysis of the problems that doesn’t collapse into a handwave about “capitalism” and ongoing experimentation at the layer of protocols and finance. It means competing on more than rhetoric and instead addressing incentives. It means actually serving the user. A structural problem needs structural solutions. The open web can’t run on vibes alone. What would it take to actually win?

