All articles Security

Bot traffic is most of your traffic

On a small site with no audience, roughly half of what arrives is automated. Here is what that looks like and what to do.

Photo: Alan Levine from United States (CC BY 2.0) / Wikimedia Commons

Put a new site online with no links to it and no announcement. Within hours something will request /wp-login.php. Within a day something will try /.env.

None of that traffic is interested in you. It is untargeted scanning, and it is the background radiation of the internet. It is also, on a quiet site, frequently the majority of requests.

What it actually looks like

The pattern is remarkably consistent. In rough order of volume:

  • Path scanning. /wp-admin, /phpmyadmin, /.git/config, /.env, /backup.sql. Someone is looking for a misconfigured server, and it does not matter to them whose.
  • Credential stuffing. Lists of leaked passwords tried against any login form that answers.
  • Content scraping. Crawlers that are not search engines, taking your pages for someone else's dataset.
  • Vulnerability probes. Requests carrying payloads for whichever framework CVE is current that month.

Almost all of it is a bare HTTP client. It does not run JavaScript. It does not hold cookies. It requests one path and moves on.

Why it matters beyond noise

Three reasons, in increasing order of expense:

  1. Your analytics are wrong. You are making product decisions on numbers that include machines. A traffic spike that is a scanner looks exactly like a traffic spike that is interest.
  2. You are paying for it. Bandwidth, function invocations, database queries, and — if you are metered on API calls — allowance. We count refused calls against your quota deliberately, because a loop hammering you with invalid keys is exactly the traffic you most want to notice.
  3. Some of it is looking for credentials. The /.env request is not idle curiosity. It is somebody hoping you deployed with the file in the web root, and it costs them nothing to ask a million sites whether you did.

What a check can and cannot do

A browser check asks the client to do something a browser does trivially and a script does not do by default: execute JavaScript, hold a nonce, wait a beat, answer. Most automated traffic fails this immediately, because most of it never runs a page at all.

What it does not do is stop a determined adversary driving a real browser. Nothing does, cheaply. Anyone who wants to script Chrome and solve your challenge can, and no product on the market changes that — whatever the marketing says.

The goal is not to be impenetrable. The goal is to make the cheap attacks stop working, which removes almost all of the volume, so that whatever is left is small enough to look at.

The tuning that matters

Three settings do most of the work:

  • Hold duration. How long the challenge waits before accepting an answer. A script that answers in four milliseconds did not wait. Two to three seconds is invisible to a person and fatal to naive automation.
  • Search engines. Leave them through, or you will deindex yourself. This is the setting people forget and then wonder why traffic fell off a cliff a fortnight later.
  • Device classes. Admitting desktop and mobile but not a headless class removes a whole category without touching anyone real.

The mistake everyone makes

Turning the check on and not installing the snippet.

The setting configures the check. The snippet is what calls it. With the setting on and no snippet, nothing is asked about anything, and the feature does precisely nothing while appearing to be enabled.

We now say so explicitly on the key's page — it shows how many requests have actually been checked, and tells you plainly when the answer is none — because it was the single most common support question we had, and every one of those conversations was somebody reasonably concluding the product was broken.

Start building

Lock a key to your domain in about five minutes.

Get started