如果你在电影行业或相关领域工作,那么本月你很有可能使用过 The Numbers 的成果,无论你是否意识到这一点。
它通过人工调研获取的数据质量最高,追踪了超过 78,000 部电影的票房收入、预算、家庭录像和流媒体数据,涉及 236,000 名从业者。该网站每年访问量超过 800 万,被记者、学者、电影制作人、预测市场甚至吉尼斯世界纪录视为绝对的权威来源。
正是这种“史上最佳”的地位,导致了今年三月的灾难性事件。
2026 年 3 月 5 日,TheNumbers.com 网站消失了。

该网站宕机超过一周,没有任何解释。一周后,它重新上线,但规模大幅缩水。历史图表、单个电影页面,甚至备受喜爱的报告生成器都不见了。
网站上只留下一条笼统的“我们正在重建,请耐心等待”的消息,互联网的反应一如既往——困惑、愤怒和阴谋论四起。Reddit 上甚至有理论认为,这是一场蓄意的“抽地毯”行动,旨在削弱免费网站,将用户推向付费产品。
三个月后,我与 The Numbers 的创始人兼首席执行官布鲁斯·纳什详细交谈了当时的情况。他描述了一段相当不愉快且多事之秋的经历:
我们收到了大量愤怒的邮件,人们质问道:‘你们以前有的那个页面现在去哪儿了?为什么没有了?’
在他的故事中,有许多事情值得每一个运营、依赖或仅仅是欣赏互联网的人警惕。
首先,一些背景信息
1997 年 10 月 17 日,星期五,数学家、前 IBM 软件开发人员布鲁斯·纳什在 Geocities 上发布了一个追踪 300 部电影的网站。
布鲁斯在一篇二十周年纪念文章中描述了这次发布(由于即将说明的原因,这篇文章现在仅存于互联网档案馆中):
我在 Access 数据库中按下一个按钮,将一些 HTML 页面上传到 Geocities,并在好莱坞股票交易所的留言板上发布了一条简短公告,告诉大家我开始分析电影的票房数据,以帮助他们挑选在 HSX 上交易的 MovieStocks。

从这些不起眼的起点开始,布鲁斯和他围绕该网站组建的团队,将 The Numbers 打造成了电影行业最可靠的财务信息来源。
到 2026 年初,该数据库追踪了 78,396 部电影、178,375 条院线上映记录以及 236,176 位人物。
机器人来了
在其存续期间,The Numbers 所面临的挑战发生了巨大变化。在最初的二十五年左右,流量是可管理的,并且大多是有礼貌的。正如布鲁斯所说:
在 AI 出现之前,我们收到的是人类流量,大多是行为良好的搜索引擎爬虫,以及少数为个人项目爬取网站的人。如果有人过于贪婪,我们就能发现并封禁他们。
在过去几年里,世界各地的网站所有者都发现他们的网络流量发生了变化。最初只是人们浏览,后来让位于数量不断增长的机器人。到 2024 年,自动化流量已经超过了人类流量,就在上个月,Cloudflare 宣布机器人已占网页请求的 57.5%。
The Numbers 通过两波明显的浪潮感受到了这种变化。第一波大约始于 2024 年:
随着 AI 训练加入搜索引擎爬虫的行列,我们看到了爬取量的大幅增长。AI 爬虫的行为通常不如搜索引擎规范,这增加了我们为保持网站平稳运行所需的管理任务。
而第二波浪潮更猛烈,破坏性也更大:
大约在 2025 年 12 月,我们看到了另一次流量激增,我将此归因于智能体 AI:一方面是响应提示词而抓取网站的 AI 智能体,另一方面是人们能够编写抓取网站的智能体。
像每一个数据丰富的网站一样,到 2026 年初,The Numbers 正遭受着 AI 机器人以工业规模反复抓取其页面的猛烈冲击。布鲁斯表示,他们只有 10% 的流量来自浏览网站的人类,其余则来自 AI 机器人和自动化流量。

网站试图适应新的机器人
这给网站带来了巨大压力,但布鲁斯和他的团队能够采取措施来减轻最坏的影响。其中最巧妙的方法之一是用机器人自己的语言与它们对话:
网站上有些内容是为大语言模型设计的,目的是让它能告诉别人“这是你获取数据许可的方式”,而不是“这是你抓取网站的方式”。这产生了巨大效果。我们现在收到的许可咨询量大概是原来的十倍。
但缓解措施并不等于摆脱困境。从12月到3月初,团队一直在努力让网站在负载下保持运行。布鲁斯估计:
我们大约90%的时间都花在维持现有网站运行上,只能利用零碎时间开发新的改进系统。
问题因网站年代久远而雪上加霜:它已有三十年历史,大约有16万个源文件,服务于约200万个页面。
随后,在3月5日(周四)凌晨,服务器崩溃了。
团队手忙脚乱地试图弄清楚发生了什么,最初以为是AI流量过大所致。看起来确实是AI的责任……但可能并不完全是他们最初想的那样。
在大量智能体流量的掩盖下,网站日志显示了一些比数据抓取更严重的问题。正如布鲁斯所描述的:
其中一些访问者使用合法URL访问网站,另一些则在寻找后门,很可能是为了在数据出现在网站上之前就获取它,或者操纵呈现给用户的数据。
根据一位从事网络安全工作的朋友的建议,旧服务器被永久关闭了。恢复备份并让这个有三十年历史的网站重新上线,意味着要保护16万个遗留文件,抵御那些已经花了数月时间探测它们的攻击者。
团队在新基础设施上快速搭建了一个网站骨架版本,至少能在他们评估发生了什么以及下一步该怎么做期间,继续提供最新的票房数据。该版本于3月13日(周五)上线。
谁会想要私下访问一个票房网站呢?
乍一看,The Numbers 似乎不是一个明显的目标。它不收集信用卡信息,也没有可以在暗网上出售的诱人客户数据。它是一家小型独立公司,专门发布电影票房收入数据。
有人怎么能指望仅仅通过私下访问他们的网站来赚钱呢?
如果你还没猜到,这件事与预测市场有关。
Polymarket 在电影开画周末每周开设市场,并将 The Numbers 网站列为最终真相来源:
该电影在 The Numbers 页面上“票房”标签下的“每日票房表现”数据,将在三天开画周末的数值最终确定后,用于结算该市场。
任何一个周末市场的金额,按金融市场标准来看都不算大,通常在数万到数十万美元之间,而所有正在运行的票房市场在任何时刻的总金额约为数百万美元。
如果你能每周都比所有人更早看到 The Numbers 的数据,你就会比其他交易者拥有显著优势——在数据发布前略早获知答案,就能让你抢先交易。
在这种情况下,很难确切知道发生了什么。我们知道日志显示该网站被自动探测和抓取了数月,但最终导致网站瘫痪的原因以及是谁干的,仍然是个未解之谜。
但有人利用 AI 在预测市场中获得优势的理论是完全合理的。The Numbers 的经历向我们表明:
我们现在生活在一个电影统计网站值得被黑客攻击的世界,因为预测市场让任何人都能将几乎任何数据转化为金钱。
如今,任何人只需花很少的钱订阅 AI 就能实施网站黑客攻击。
面对大规模智能体 AI 机器人的集群攻击,当前的互联网极其脆弱。
那么,如今的黑客攻击到底有多容易?
2025 年 11 月,Anthropic(Claude 背后的 AI 实验室)发布了一份报告,称这是首次有记录的由 AI 策划的网络间谍活动。一个由国家支持的组织使用其编码工具攻击了大约 30 个组织,其中 AI 完成了 80% 到 90% 的工作,而人类在每个行动中仅介入 4 到 6 个决策点。
Anthropic 自己的结论是:
实施复杂网络攻击的门槛已大幅降低,我们预测这一趋势还将持续。
在更早的一份威胁报告中,Anthropic 说得更加明确:
缺乏技术技能的犯罪分子正在利用人工智能实施复杂的操作,例如开发勒索软件,而这些操作以往需要经过多年的培训才能完成。
与此同时,一款名为 XBOW 的自主人工智能渗透测试工具登上了 HackerOne 美国排行榜的榜首。该排行榜旨在对在真实公司中寻找安全漏洞以获取赏金的人员(此前全部为人类)进行排名,XBOW 在此过程中提交了近 1060 个漏洞。
访问一个拥有 16 万个遗留文件的三十年历史网站,正是人工智能工具已使其探测成本变得低廉的、存在已知缺陷的攻击面。曾经保护小型网站免受除最坚定攻击者之外所有人攻击的专业知识壁垒,在很大程度上已经消失了。
The Numbers 网站现在该怎么办?
布鲁斯和他的团队相对幸运。尽管整个网站一夜之间被攻陷,但他们仍能继续运营。The Numbers 一直免费使用,并且过去几年该网站并未严重依赖广告,因此这次宕机并未摧毁他们赖以生存的收入来源。
他们的核心业务与通过 OpusData 服务销售批量数据、为电影制作人和投资者制作票房分析报告以及发布《商业报告》相关联——所有这些业务均未受到公共网站瘫痪的影响。
但他们确实需要从头开始构建一个全新的网站,来托管那 78,396 部电影、178,375 条发行记录和 236,176 位人物。从备份中恢复网站并非可行之选,正如布鲁斯所指出的:
很明显,我们不能简单地重新上线那台服务器,因为它必然会再次被攻陷,可能就在几分钟之内。
这就是为什么该网站在三月中旬以极简形式恢复上线,并且功能是逐步而非一次性全部回归的原因。
目前,该团队不得不重新思考在 2026 年,一个公共网站究竟意味着什么。布鲁斯的分析是,The Numbers 过去服务于两类受众(人类和搜索引擎),而现在大约服务于六类:人类、搜索引擎、大语言模型训练运行、基于提示词的 AI 流量、智能体 AI 以及预测市场投机者。每一类都有不同的需求和不同的流量特征。正如他所说:
我们从一个运营网站只需关注三件事(内容、广告和 SEO)的世界,变成了每个设计决策都要考虑大约八到十个不同因素的局面。
他表示,目标是服务所有六类受众,为《商业报道》订阅用户提供新的 OpusData 服务和在线功能,更重要的是,帮助网站上的普通人类用户重新获取该网站一直提供的数据,其中部分数据将以全新且改进的形式呈现。
机器人爬取到底能有多糟糕?
说实话,相当糟糕。糟糕到像 Bruce 这样的网站所有者不得不质疑,花费如此多时间和金钱去构建和维护的东西,其价值究竟何在。
保护着全球大量网站的 Cloudflare 发布了一组数据,显示每个 AI 平台每为其爬取的网站回引一位访客,会爬取多少页面。
谷歌每为你回引一位访客,大约爬取 5 个页面。OpenAI 爬取超过 1000 个页面。Anthropic 每回引一位访客,爬取的页面超过 38,000 个。

请注意,该图表采用对数刻度,即底部的每一步都比前一步大十倍,否则这些差异实在太大,我无法在一张图表中展示。
纵观互联网至今的历史,开放网络的原则是:作为允许搜索引擎机器人读取你网站的回报,它们会为你带来读者。但现在,这种交换已不再适用。机器人的数量呈爆炸式增长,而且它们不再回引任何用户。
当这股洪流对准一个小型网站时,它可能会推高带宽费用,甚至可能让整个网站瘫痪。与 Bruce 经历相似的网站包括:
Read the Docs,一个为开源软件托管文档的非营利组织,他们发现一个爬虫在一个月内下载了 73 TB 的压缩 HTML 文件,导致其花费超过 5000 美元的带宽费用。
iFixit,这个维修指南数据库,在一天之内就记录了来自 Anthropic 爬虫的一百万次访问。
一家只有七名员工、销售 3D 扫描产品的公司 Triplegangers,在营业时间被 OpenAI 的爬虫攻击至离线,其 CEO 形容这“基本上就是一次 DDoS 攻击”。代码托管服务 SourceHut 的创始人报告称,他“每周要花 20% 到 100% 的时间”来对抗 AI 爬虫,导致“每周出现数十次短暂宕机”。
Linux 新闻网站 LWN 的编辑描述了来自“literally millions of IP addresses”的爬虫流量,并得出结论:“这是一次分布式拒绝服务攻击”。
当 GNOME 开源项目测量其流量时,发现大约 97% 的流量来自机器人。
一所大学图书馆在 48 小时内封禁了 16,000 个 IP 地址,以维持其目录系统的在线运行。
运营维基百科的维基媒体基金会在 2025 年 4 月报告称,机器人贡献了约 35% 的页面浏览量,但却占据了至少 65% 的最昂贵流量,因为爬虫会批量读取人类读者很少触及的冷门页面。

六个月后,挤压效应的另一半显现出来:维基百科的人类页面浏览量同比下降约 8%,因为人们越来越多地通过 AI 摘要获取维基百科的知识,而不再直接访问维基百科。机器正在以工业规模同时攫取内容和读者。
在公众中进行测试
AI 工具是人类创造过的最强大、也最具破坏性的工具之一。而它们正在现实世界中,由公众实时进行着有效的测试。当年曼哈顿计划想要测试原子技术的威力时,他们可不会每天早上把技术规格发给所有人,然后看哪栋房子被炸毁。
我们迄今为止所构建的世界,对我们所有人都能使用的 AI 模型的强大程度和规模,准备得极其不足。
我不希望这听起来像是一场片面的反 AI 恐惧宣传。AI 以及它能为人类做的事情,有很多值得喜欢的地方。但我们确实需要思考一下,我们正在踏入的是一个怎样的世界。
最先崩溃的是那些为旧互联网构建的东西。开放网络建立在一些假设之上,例如访问者大多是真人、流量大致反映读者数量、以及运营网站的成本与你从中获得的价值相关。如今,这些假设每一条都已过时。
一年前,Cloudflare 推出了按爬取付费模式,允许网站向 AI 爬虫按页面收费。上周,它更进一步,宣布了一种按使用量付费的模式,即当出版商的内容实际出现在 AI 回答中时,他们就能获得报酬;同时宣布,从 9 月 15 日起,其客户的广告支持页面将默认屏蔽未付费的“混合用途”爬虫。
这些措施能否奏效,取决于 AI 公司是选择配合,还是设法绕过。但正如 Bruce 对我所说,总得有人去尝试。
尾声
让我们暂且抛开具体细节,思考一下这里究竟发生了什么。
一个受人喜爱、实用且免费的网站,由一位能干、诚实的人精心运营了近三十年,却被新 AI 经济的两个特征压垮了。
不可持续的机器流量从上方猛烈冲击它,而很可能是一个受金钱驱使的入侵者,在一个入侵网站从未如此容易的世界里,将其彻底击垮。
Bruce 的事业和生计得以幸存,仅仅是因为这个网站并非他的全部业务。
其他人就没这么幸运了。就在上周,一家经营了 37 年的德国纺织公司 ZEGO,在 3 月份的一次网络攻击导致其生产中断六周后,申请了破产。与 The Numbers 不同,他们没有其他业务可以依靠。
网络上充斥着独立的档案库、爱好者的数据库、地方新闻网站、论坛、参考著作。数十年积累的人类努力,运行在陈旧的代码上,由小团队或个人维护,默默地支撑着我们共享知识中远比任何人所承认的要多得多的部分。
The Numbers 正在回归,并且比以往构建得更好。我鼓励你继续使用它,继续支持它,如果你也是众多因某个页面缺失而愤怒地给 Bruce 发邮件的用户之一,那么现在你知道了它缺失的原因,或许可以发一封更友善的邮件。
If you work in or around the film industry, there is a decent chance you have used the work of The Numbers this month, whether you realise it or not.
Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even Guinness World Records.
And it was this GOAT status which caused the catastrophic events of March this year.
On the 5th March 2026, TheNumbers.com website vanished.

The site was down for over a week, without explanation. A week later, it resurfaced at a fraction of its former size. Gone were the historical charts, the individual movie pages, and even the much-loved Report Builder.
With only a generic “we’re rebuilding, please bear with us” message to go on, the internet responded as it always does - with confusion, anger, and conspiracy theories.
One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products.
Three months on, I spoke at length with Bruce Nash, founder and CEO of The Numbers, about what happened. He describes quite an unpleasant and eventful experience:
We got a lot of angry emails from people who are like, 'Where's this page that you used to have and you don't have anymore?'
Within his tale are a number of things that should worry anyone who runs, relies on, or simply appreciates the internet.
First, some background
On Friday 17 October 1997, mathematician and former IBM software developer Bruce Nash launched a Geocities site that tracked 300 films.
Bruce described the launch in a 20th anniversary essay (which now survives only in the Internet Archive, for reasons that will become clear):
I hit a button in an Access database, uploaded some HTML pages to Geocities, and made a brief announcement on the Hollywood Stock Exchange message boards to let people know that I was starting to analyze box office for films to help them pick MovieStocks to trade on HSX.

From those humble beginnings, Bruce and the team he built around the site turned The Numbers into the film industry's most reliable financial source.
At the start of 2026, the database tracked 78,396 movies, 178,375 theatrical release records, and 236,176 people.
The robots arrive
During its lifetime, the challenges The Numbers has faced have changed immensely. For its first quarter century or so, the traffic was manageable and mostly polite. As Bruce puts it:
Pre-AI, we got human traffic, mostly well-behaved search engine crawlers, and a few people crawling the site for personal projects. If someone got too greedy, we could spot them and block them.
Over the past couple of years, website owners the world over have seen their web traffic change. What was initially only people browsing gave way to an ever-increasing number of bots. By 2024, automated traffic had surpassed human traffic, and just last month, Cloudflare announced that bots had reached 57.5% of web page requests.
The Numbers felt this shift in two distinct waves. The first started around 2024:
We saw a big increase in crawls as AI training joined the search engine crawlers. The AI crawlers are generally less well-behaved than the search engines, which increased the management tasks for us to keep the site running smoothly.
And the second wave was stronger and more damaging:
Around December 2025, we saw another big spike in traffic which I attribute to agentic AI: a combination of AI agents that scrape sites in response to prompts, and people being able to write agents that scrape sites.
Like every data-rich site, by early 2026 The Numbers was being hammered hard by AI bots scraping its pages over and over at an industrial scale. Bruce says that only 10% of their traffic is from humans browsing the site, with the rest coming from AI bots and automated traffic.

Websites try to adapt to the new robots
This put enormous strain on the site, but Bruce and his team were able to take measures to mitigate the worst of it. One of the cleverest was talking to the robots in their own language:
There’s stuff on the site which is designed for an LLM to read, so that it can tell somebody ‘here’s how you licence the data’ rather than ‘here’s how you scrape the website’. It’s had a huge effect. We’re now getting probably ten times the volume of licensing enquiries.
But mitigation is not the same as escape. From December through early March, the team struggled to keep the site alive under the load. Bruce estimates that:
Around 90% of our time was spent keeping the existing site running while we spent our spare moments working on a new and improved system.
The problem was compounded by the site’s age: thirty years old, with approximately 160,000 source files serving around 2 million pages.
Then, in the early hours of Thursday 5 March, the servers collapsed.
The team scrambled to understand what had happened, initially assuming it was the sheer weight of AI traffic. It seems AI was to blame... but possibly not only in the way they first thought.
Buried in the flood of agentic traffic, the site’s logs showed something more pointed than scraping. As Bruce describes it:
Some of these used the site using legitimate URLs, others were looking for back doors, most likely so they could get to the data before it appeared on the site, or to manipulate the data presented to users.
On the advice of a friend who works in cybersecurity, the old server stayed off. For good. Restoring the backups and nursing the thirty-year-old site back online would have meant defending 160,000 legacy files against attackers who had spent months probing them.
The team rushed up a skeleton version of the website on new infrastructure, which could at least keep delivering the latest box office figures while they took stock of what had happened and what to do next. It went live on Friday 13 March.
Who would want private access to a box office website?
At first glance, The Numbers may not seem like an obvious target. It doesn’t collect credit card information, and there is no juicy customer data to flip on the dark web. It is a small, independent company that publishes how much money movies make.
How could someone expect to make money purely from having private access to their site?
In case you haven’t guessed it yet, it’s linked to prediction markets.
Polymarket runs weekly markets on opening weekends, and names The Numbers as the ultimate source of truth:
The ‘Daily Box Office Performance’ figures found on the ‘Box Office’ tab on this movie’s The Numbers page will be used to resolve this market once the values for the 3-day opening weekend are final.
The sums on any single weekend market are modest by financial-market standards, typically in the tens to hundreds of thousands of dollars, with a couple of million dollars across live box office markets at any given time.
If you could see The Numbers data before everyone else, every single week, you would have a significant edge over all the other traders - learning the answers slightly ahead of publication would allow you to front-run the trades.
In a situation like this, it is hard to know for certain what happened. We know that the logs showed months of automated probing and scraping of the site, but what finally brought the site down, and who did it, remains an open question.
But the theory that someone used AI to develop an advantage in a prediction market is entirely plausible. The Numbers experience shows us that:
We now live in a world where a movie statistics website is worth hacking because prediction markets empower anyone to turn almost any data into money.
Hacking websites is now something anyone can do with a cheap AI subscription.
The web, as we have it, is incredibly fragile in the face of large-scale swarms of agentic AI bots.
How hard is hacking these days, anyway?
In November 2025, Anthropic (the AI lab behind Claude) published a report on what it called the first documented AI-orchestrated cyber espionage campaign. A state-sponsored group had used its coding tool to attack roughly 30 organisations, with the AI performing 80% to 90% of the work and humans stepping in at only 4 to 6 decision points per campaign.
Anthropic’s own conclusion was:
The barriers to performing sophisticated cyberattacks have dropped substantially, and we predict that they’ll continue to do so.
In an earlier threat report, Anthropic were even clearer:
Criminals with few technical skills are using AI to conduct complex operations, such as developing ransomware, that would previously have required years of training.
Meanwhile, an autonomous AI penetration tester called XBOW reached number one on HackerOne’s US leaderboard, the ranking of the people (formerly all people) who find security holes in real companies for bounties, submitting nearly 1,060 vulnerabilities along the way.
Getting access to a thirty-year-old website with 160,000 legacy files is exactly the kind of known-flaw surface that AI tools have made cheap to probe. The expertise barrier that once protected small sites from all but the most determined attackers has largely evaporated.
What now for The Numbers?
Bruce and his team were relatively lucky. Despite having their entire site knocked out overnight, they were able to keep going. The Numbers has always been free to use, and the site hasn’t relied heavily on advertising for the past few years, so the outage didn’t destroy an income stream they depended on.
Their core business is tied to selling bulk data through the OpusData service, producing comp analysis reports for filmmakers and investors, and publishing the Business Report - all of which were unaffected by the public site going down.
But they do need to build an entirely new website, from scratch, to host those 78,396 movies, 178,375 release records and 236,176 people. Restoring the site from a backup wasn’t an option, as Bruce points out:
It was really clear that we couldn’t just put that server up again, because it would inevitably be brought down again, possibly within minutes.
That is why the site came back bare-bones in mid-March, and why features are returning gradually rather than all at once.
Right now, the team is having to reconsider what a public website even means in 2026. Bruce’s analysis is that The Numbers used to serve two audiences (human beings and search engines) and now serves roughly six: humans, search engines, LLM training runs, prompt-based AI traffic, agentic AI, and prediction market punters. Each has different needs and a different traffic profile. As he puts it:
We’ve gone from a world where running a web site meant focusing on three things (content, ads, and SEO) to about eight to ten different factors that go into every design decision.
The goal, he says, is to support all six audiences, with new OpusData services and online features for Business Report subscribers, and, importantly, to help regular human users of the site regain the data it has always provided, some of it in new and improved form.
How bad could bot scraping really be?
Pretty bad, tbh. Enough that site owners such as Bruce have to question the value of something that will take so much time and money to build and defend.
Cloudflare, which protects a huge share of the world’s websites, publishes data on how many pages each AI platform crawls for every one visitor it sends back to the websites it crawled.
Google crawls about five pages for every visitor it sends you. OpenAI crawls over 1,000. Anthropic crawls over 38,000 pages for every single visitor it refers.

Note that the scale is logarithmic, i.e. each step along the bottom is ten times bigger than the last, because otherwise the differences are quite literally too large for me to include on one chart.
For the history of the internet to date, the principle of the open web was that, in return for letting the search engine robots read your site, they would send you readers. But now, that trade no longer applies. The number of robots has exploded, and they no longer send anyone back.
When this firehose is aimed at a small site, it can inflate the bandwidth bill and possibly even take down an entire site. Sites which can relate to Bruce’s experience include:
Read the Docs, a non-profit that hosts documentation for open-source software, who watched a single crawler download 73 terabytes of zipped HTML in one month, costing it over $5,000 in bandwidth.
iFixit, the repair-guide database, logged a million hits from Anthropic’s crawler in a single day.
Triplegangers, a seven-person company selling 3D scans, was knocked offline during business hours by OpenAI’s bot, in what its CEO described as “basically a DDoS attack”. The founder of code-hosting service SourceHut reported spending “anywhere from 20-100% of my time in any given week” fighting AI crawlers, with “dozens of brief outages per week”.
The editor of Linux news site LWN described crawler traffic from “literally millions of IP addresses” and concluded: “it is a distributed denial-of-service attack”.
When the GNOME open-source project measured its traffic, roughly 97% turned out to be bots.
A university library banned 16,000 IP addresses in 48 hours to keep its catalogue online.
The Wikimedia Foundation, which runs Wikipedia, reported in April 2025 that bots account for about 35% of its pageviews but at least 65% of its most expensive traffic, because crawlers bulk-read obscure pages that human readers rarely touch.

Six months later came the other half of the squeeze, when Wikipedia’s human pageviews fell roughly 8% year on year, as people increasingly get Wikipedia’s knowledge from AI summaries without ever visiting Wikipedia. The machines are taking both the content and the readers at an industrial scale, too.
Testing it in public
AI tools are some of the most powerful and destructive things humans have ever created. And they are being effectively tested by the public in real time in the real world. When the Manhattan Project was trying to work out the power of their atomic tech, they did not do so by sending everyone the specs each morning and seeing which houses blew up.
The world we have built thus far is so incredibly ill-prepared for the power and scale of the AI models we all have access to.
I don’t wish for this to sound like a one-sided anti-AI fear campaign. There is a lot to like about AI and what it can do for the human race. But we do need to consider the world we’re currently stepping into.
What breaks first are the things built for the old internet. The open web was built on assumptions such as that visitors are mostly human, that traffic roughly tracks readership, and that the cost of serving your site is related to the value you get from serving it. Every one of those assumptions is now out of date.
A year ago, Cloudflare launched pay-per-crawl, letting sites charge AI crawlers per page. Last week, it went further, announcing a pay-per-use model in which publishers get paid when their content actually appears in an AI answer, and declaring that, from 15 September, its customers’ ad-supported pages will block unpaid “mixed-use” crawlers by default.
Whether any of this works depends on whether the AI companies play along rather than route around it. But as Bruce put it to me, somebody has to try.
Epilogue
Let’s look beyond the specifics for a moment and consider what happened here.
A beloved, useful, free website, run carefully by a competent, honest person for nearly thirty years, was crushed between two features of the new AI economy.
Unsustainable machine traffic hammered it from above, and in all likelihood a financially motivated intruder, operating in a world where breaking into websites has never been easier, took it down.
Bruce’s business and livelihood survived only because the website was not the whole business.
Others have not been so fortunate. Just last week, ZEGO, a German textile firm that had been in business for 37 years, filed for insolvency after a single cyberattack in March shut down its production for six weeks. Unlike The Numbers, they had no other business to fall back on.
The web is full of independent archives, hobby databases, local news sites, forums, reference works. Decades of accumulated human effort, running on old code, maintained by small teams or single individuals, quietly holding up far more of our shared knowledge than anyone acknowledges.
The Numbers is coming back, better built than before. I would encourage you to keep using it, keep supporting it, and, if you are one of the many people who emailed Bruce in fury about a missing page, perhaps send a kinder one now you know why it was missing.