I'm curious whether we'll hit a point where some sites will want to be as friendly to bots as possible (hopefully they'll just shift to having sane, rate-limited public APIs so everyone can save time and bandwidth), while the rest want the reverse.
If you're providing a service, and AI agents truly take off to where people say "Claude, go buy me a new shower hose about 5' long, brushed nickel. Use my usual evaluation criteria and don't spend over $20." then a lot of ecommerce stores are very very incentivized to be bot friendly.
And of course every "content" site is the reverse, wanting to not waste their bandwidth and hoping to score exclusive paid content deals with any entity that might want to scrape them.
My cgit instance is incredibly friendly to bots; I go out of my way to support git's fast "smart" HTTP protocol, and the clone URL is linked on every project's root page. I even included the full clone command on 429 error pages.
The end result was that I got millions of requests daily trying to download every HTML page (every variation of action, commit hash, branch, whatever else cgit allows) and basically no traffic to the proper clone route.
Turns out the bots just don't care about your bot-friendly access method and will instead throw more illicitly-proxied residential IPs at the problem.
I've already seen a small amount of transition in that direction. For instance:
Adafruit, last year: They blocked bots very heavy-handedly, including those that were there to gather information at the behest of a present-at-the-keyboard human. This made their products nearly invisible for those who were using ChatGPT instead of Google to find particular things to buy. Even when they were discoverable through forum postings, this added friction meant that I was unable to make a swift comparison betwixt their wares and those of others, and I wasn't buying anything from them because they were making deliberate choices that I felt were hostile to my process.
But that's changed.
Adafruit, at least a week or so ago: The bot was able to get in and find details about the stuff they sell. This lack of friction meant that I was able to make a swift comparison between their wares and those of others. I still bought the particular sensor-chip-on-a-board widget that I wanted elsewhere, but at least Adafruit's offerings were in the running this time.
Very few "user facing" applications will have to deal with real users any longer. From an author's point of view let's admit that it will make things easier.
Yeah, as soon as site operators implement that paid traffic scheme the Cloudflare has been pushing. It's all about being to monetize their compute/content.
well it's funny but when I was writing my bot targeting a particular platform I did a lot of this stuff, because I was paranoid and didn't want to get blocked, so things I did:
detect all element in currently loaded page I want to do things with. element I want to click. scroll to element I want to click. random viewport scrolling action, go slightly over viewport, scroll back up, determine mouth position where I will click - x, y center of clickable item, determined x,y of bounding box of clickable item, choose some point within bounding box where I can click with some randomness for getting to center because obviously humans don't get the exact center every time do they. set random mouse move speed. mouse choppiness function (determine a number of x.y coordinates on the way towards the x.y you will actually be clicking move towards those each in turn with randomness of mouse speed, so that nobody can say huh, this mouse is moving towards the place it will click with an unnaturally direct and unnaturally smooth movement.) Get to where I will click. wait random time to click. click.
If I do not continually get new elements I want to click based on my criteria just by this method now I need to scroll to see if I can detect new.
random number of times I will scroll. random scroll speed. random scroll distance.
start scroll. hit scroll distance stop. random up down scroll.
My automation goal was not to make my bot able to do the work of hundreds of people. It was to do the work of a few people.
However I eventually gave it up since the target platform was continually deciding I, as a human, was somehow behaving as a bot, it followed that it was somewhat irrational to try to make a bot that behaved as close to a human as possible.
(note there was all sorts of other erratic behavior boiled into my bot, for example how it wrote text, since it had to do keypresses, and randomized typing speeds, and making errors, and back spacing to fix errors)
on edit: before you start telling me about how it's not randomized enough or whatever, I am of course giving a curtailed high level description of the kinds of noise I added in to my bot behavior, because I just really wanted to spin up a bot that could spend a few hours doing stupid work while I did other things, and then write up a report and then later I could come in and do some more clever things.
it should be noted that later I was told by someone who should know that I had over-thought and over-architected my bot and nobody checked for even 10% of the things I was doing to make my bot behave the way a human did.
I was experimenting with Google Ads traffic to one of the websites I cared about, and even just looking at the existence of mouse move and scroll events, let alone the patterns, made me understand how many bots, or users who just (mis)clicked the ad and never did any action on my website, were counted as paid clicks in Google Ads. Analyzing for patterns like in this article would've made even more sense.
I alway wondered why so many people (not saying you're one of them) get mad at being charged for misclicks and bots.
It all averages out.
You bid $3 a click because your conversion ratio is x, and paying $3 is worth it. Even if 90% of the traffic is bots, it works out.
If Google filtered out all the bots, the conversion ratio for you and your competitors would go up, and your competitors would be willing to pay $27 per click and you'd outbid at $30, being at the exact same result at the end of the month spent for the exact same results.
For me, $3 a click would've worked for some projects, but $30 would not work for sure. Knowing that it's $30 and not $3 before the start of some campaigns, I would've saved some money. I instead paid that money to Google for this lesson.
This is another one of those horse-blinders style arguments that considers only for-profit/institutional/etc contexts and so doesn't realize their very method of measuring something through javascript program execution is biased or that it's creepy.
I'm a human, I don't execute javascript programs and by default my browser does not either. Anyone using this will label me a bot.
Hello! Author of the article here. I am working on this because I had one of my websites (which served downloads related to the game Minecraft) flooded with hundreds of thousands of requests coming from scraper bots from all around the world that even tried to download the files.
This made me hit the bandwith cap of my hosting provider and had to pay extra per TB used that month from my own pocket, since the website was not making a profit at all.
I never thought of increasing profits for any institution, but to stop the waste of bandwith many independent websites suffer.
if by "bot" we mean "not economically interesting", aren't websites within their right to filter our hacker-nerds who probably won't click on ads and buy useless objects?
the ratio of this comment vs. parent is absurd. tfa's prior mentions "anti-cheat" in games. while cheaters do suck, a lot of draconian trash has crept in under the guise of that banner.
i have been finding my locked down experience has been becoming degraded lately fwiw. i am not economically "interesting", and the consequences of valuing privacy in 2026 is limiting information access. shame really.
You're absolutely right. The modern meaning of "bot" is "not economically exploitable". I wish they'd just say that rather than using the plausible deniability for what they doing re: "bots".
I'm curious whether we'll hit a point where some sites will want to be as friendly to bots as possible (hopefully they'll just shift to having sane, rate-limited public APIs so everyone can save time and bandwidth), while the rest want the reverse.
If you're providing a service, and AI agents truly take off to where people say "Claude, go buy me a new shower hose about 5' long, brushed nickel. Use my usual evaluation criteria and don't spend over $20." then a lot of ecommerce stores are very very incentivized to be bot friendly.
And of course every "content" site is the reverse, wanting to not waste their bandwidth and hoping to score exclusive paid content deals with any entity that might want to scrape them.
My cgit instance is incredibly friendly to bots; I go out of my way to support git's fast "smart" HTTP protocol, and the clone URL is linked on every project's root page. I even included the full clone command on 429 error pages.
The end result was that I got millions of requests daily trying to download every HTML page (every variation of action, commit hash, branch, whatever else cgit allows) and basically no traffic to the proper clone route.
Turns out the bots just don't care about your bot-friendly access method and will instead throw more illicitly-proxied residential IPs at the problem.
I've already seen a small amount of transition in that direction. For instance:
Adafruit, last year: They blocked bots very heavy-handedly, including those that were there to gather information at the behest of a present-at-the-keyboard human. This made their products nearly invisible for those who were using ChatGPT instead of Google to find particular things to buy. Even when they were discoverable through forum postings, this added friction meant that I was unable to make a swift comparison betwixt their wares and those of others, and I wasn't buying anything from them because they were making deliberate choices that I felt were hostile to my process.
But that's changed.
Adafruit, at least a week or so ago: The bot was able to get in and find details about the stuff they sell. This lack of friction meant that I was able to make a swift comparison between their wares and those of others. I still bought the particular sensor-chip-on-a-board widget that I wanted elsewhere, but at least Adafruit's offerings were in the running this time.
[dead]
Very few "user facing" applications will have to deal with real users any longer. From an author's point of view let's admit that it will make things easier.
the best economic use-case for the semantic web (text/turtle, text/n3); businesses providing manifests of their stocks and services
plus, AI would also skip search result ads, and, if those have to legally be displayed as such, then there goes the market for them, I'd think
Yeah, as soon as site operators implement that paid traffic scheme the Cloudflare has been pushing. It's all about being to monetize their compute/content.
well it's funny but when I was writing my bot targeting a particular platform I did a lot of this stuff, because I was paranoid and didn't want to get blocked, so things I did:
detect all element in currently loaded page I want to do things with. element I want to click. scroll to element I want to click. random viewport scrolling action, go slightly over viewport, scroll back up, determine mouth position where I will click - x, y center of clickable item, determined x,y of bounding box of clickable item, choose some point within bounding box where I can click with some randomness for getting to center because obviously humans don't get the exact center every time do they. set random mouse move speed. mouse choppiness function (determine a number of x.y coordinates on the way towards the x.y you will actually be clicking move towards those each in turn with randomness of mouse speed, so that nobody can say huh, this mouse is moving towards the place it will click with an unnaturally direct and unnaturally smooth movement.) Get to where I will click. wait random time to click. click.
If I do not continually get new elements I want to click based on my criteria just by this method now I need to scroll to see if I can detect new.
random number of times I will scroll. random scroll speed. random scroll distance.
start scroll. hit scroll distance stop. random up down scroll.
My automation goal was not to make my bot able to do the work of hundreds of people. It was to do the work of a few people.
However I eventually gave it up since the target platform was continually deciding I, as a human, was somehow behaving as a bot, it followed that it was somewhat irrational to try to make a bot that behaved as close to a human as possible.
(note there was all sorts of other erratic behavior boiled into my bot, for example how it wrote text, since it had to do keypresses, and randomized typing speeds, and making errors, and back spacing to fix errors)
on edit: before you start telling me about how it's not randomized enough or whatever, I am of course giving a curtailed high level description of the kinds of noise I added in to my bot behavior, because I just really wanted to spin up a bot that could spend a few hours doing stupid work while I did other things, and then write up a report and then later I could come in and do some more clever things.
it should be noted that later I was told by someone who should know that I had over-thought and over-architected my bot and nobody checked for even 10% of the things I was doing to make my bot behave the way a human did.
That's great!
I was experimenting with Google Ads traffic to one of the websites I cared about, and even just looking at the existence of mouse move and scroll events, let alone the patterns, made me understand how many bots, or users who just (mis)clicked the ad and never did any action on my website, were counted as paid clicks in Google Ads. Analyzing for patterns like in this article would've made even more sense.
I alway wondered why so many people (not saying you're one of them) get mad at being charged for misclicks and bots.
It all averages out.
You bid $3 a click because your conversion ratio is x, and paying $3 is worth it. Even if 90% of the traffic is bots, it works out. If Google filtered out all the bots, the conversion ratio for you and your competitors would go up, and your competitors would be willing to pay $27 per click and you'd outbid at $30, being at the exact same result at the end of the month spent for the exact same results.
For me, $3 a click would've worked for some projects, but $30 would not work for sure. Knowing that it's $30 and not $3 before the start of some campaigns, I would've saved some money. I instead paid that money to Google for this lesson.
The problem is when your competitors get the idea to spam your campaign with fake traffic, which is a real thing that happens.
'Clicks' like that have bolstered stock prices for decades now.
This is another one of those horse-blinders style arguments that considers only for-profit/institutional/etc contexts and so doesn't realize their very method of measuring something through javascript program execution is biased or that it's creepy.
I'm a human, I don't execute javascript programs and by default my browser does not either. Anyone using this will label me a bot.
Hello! Author of the article here. I am working on this because I had one of my websites (which served downloads related to the game Minecraft) flooded with hundreds of thousands of requests coming from scraper bots from all around the world that even tried to download the files.
This made me hit the bandwith cap of my hosting provider and had to pay extra per TB used that month from my own pocket, since the website was not making a profit at all.
I never thought of increasing profits for any institution, but to stop the waste of bandwith many independent websites suffer.
if by "bot" we mean "not economically interesting", aren't websites within their right to filter our hacker-nerds who probably won't click on ads and buy useless objects?
the ratio of this comment vs. parent is absurd. tfa's prior mentions "anti-cheat" in games. while cheaters do suck, a lot of draconian trash has crept in under the guise of that banner.
i have been finding my locked down experience has been becoming degraded lately fwiw. i am not economically "interesting", and the consequences of valuing privacy in 2026 is limiting information access. shame really.
You're absolutely right. The modern meaning of "bot" is "not economically exploitable". I wish they'd just say that rather than using the plausible deniability for what they doing re: "bots".
[dead]
[flagged]