Stopping the bad bots blitz

Humans are now outnumbered on the internet. According to researchers, bots account for more than half of all internet traffic. The consequences of this development are starting to show.

A bad bot, or a stealth bot as it is also called, is a crawler that intentionally conceals or falsifies its identity or purpose, circumvents access controls or rights reservations, or impersonates human traffic. A new bipartisan proposal in the United States – the Stealth Bot Prohibition Act – starts from a simple principle: automated crawlers should say who they are and why they are visiting websites. It is an idea WAN-IFRA unreservedly endorses and that lawmakers around the world should embrace.

For news publishers, the consequences of today’s bot invasion are increasingly serious.

Malevolent actors use bad bots to steal news content – also behind paywalls – to then put it on what media analysts have called a billion-dollar reseller market, one where AI developers buy from the owners of the bots, not the journalists and publishers. They then use the stolen journalism to create products that directly compete with the publishers’ websites.

In the words of A.G. Sulzberger, “This theft isn’t just happening because publishers are leaving their toys out on the lawn; it’s happening when they are locked up safely in the house.”

Wreaking havoc

Uninvited bad bots can hit websites millions of times in a single day, slowing them down for human readers and potentially leading to extra bandwidth charges or even total outages. In August 2025, one UK technology site was taken offline after being scraped 1.6 million times in a single day. The problem has only escalated since: The Wikimedia Foundation, which hosts Wikipedia, now blocks or throttles roughly a quarter of all automated requests hitting its infrastructure, amounting to billions of requests a day from crawlers that ignore its access policies.

This new flood of traffic is changing the makeup of the internet itself, and in the process, creating a significant financial burden for newsrooms both large and small. Publishers have tried to stop these bots on their own by deploying sophisticated blockers, but it’s almost impossible to stop them when malicious bots are able to disguise themselves and act with impunity. Local news providers, already operating at razor-thin margins, can’t afford the higher technical costs.

The US Stealth Bot Prohibition Act is deliberately straightforward. It would require AI bad bots to disclose their identity and purpose, giving news publishers and other content providers the information they need to protect their work from actors that currently often disguise themselves as humans.

Top executives backing bot action

The proposal has received broad support from news publishers. News Corp Chief Executive Robert Thomson has described stealth bots as “the silent scavengers of the internet”, while Condé Nast CEO Roger Lynch has warned that disguised bots scrape original journalism “with zero accountability.”

The UK is moving on a parallel track. The Automated Online Software (Access and Transparency) Bill, a Private Member’s Bill introduced by MP Damian Hinds and backed by the News Media Association, would require anyone running a bot that systematically copies content from a UK website to disclose who they are, who owns the bot and what the content will be used for. Two legislatures, on opposite sides of the Atlantic, have independently landed on the same basic fix. That convergence, arrived at separately rather than coordinated, is the clearest sign yet that bot transparency is not a niche demand but the foundation any functioning digital market needs.

Transparency is also the essential first step towards a functioning licensing market. Publishers cannot negotiate with companies they cannot identify or protect content when they cannot see who is taking it. The EU’s Copyright in the Digital Single Market Directive gives rightsholders the legal right to reserve “in a machine-readable way” their content from text and data mining, and since 2 August 2026 AI developers can face fines of up to €15 million or 3% of global turnover for ignoring those reservations.

But that right is hollow if the machines knocking on the door are allowed to lie about who they are. An opt-out only works if the crawler reading it identifies itself honestly in the first place, and enforcement is only as strong as the ability to tell truthful crawlers from dishonest ones.

Publishers paying the price

The consequences of the bad bot invasion are real. Publishers bear the cost of producing original reporting and maintaining the digital infrastructure from which it is distributed. Yet bad actors can extract that work, sell access to it and help create AI products that compete directly with the original sources.

The strip-mining of journalism must end, not by pitting publishers against AI companies but by fixing a market that currently cannot function at all. Transparency about who is crawling, and enforceable consequences when they lie about it, are the precondition for any functioning digital marketplace, whether that market rewards publishers, AI developers, or the readers both ultimately serve. Countries should introduce bad bot legislation as a first step toward that market, which is essential to a healthy democracy.

Source link

Similar Posts