← Back to blog

Half my traffic is bots

August 15, 2026


Last month 48% of the traffic to my Kimbundu dictionary was bots. Crawlers and scrapers, not people. I checked the data expecting to be annoyed. I am not.

kimbundu.org got 14,375 pageviews in those four weeks. 6,900 came from machines. One crawler walked the whole alphabetical index, 425 pages, in about an hour. Another 2,200 "visitors" hit one page each and left.

Dig into those 6,900 machine pageviews and they split into three groups.

The first is search engines. Googlebot, Bingbot, and their friends. They index the site so people can find it, and they are the reason around two thousand real people reach the dictionary through search every month. I want more of these.

The second is the interesting group. Large language model crawlers. GPTBot, ClaudeBot, CommonCrawl. They do not announce themselves, so they look like noise. The crawler that walked 425 pages of the index in an hour was almost certainly one of them. It was not reading the dictionary. It was ingesting it.

The third group is neither. A botnet of roughly 2,200 one-page hits spread across thirty countries that have no particular connection to Kimbundu. Brazil, India, Argentina, Mexico, one page each, then gone. That is not an indexer and not an AI lab. It is a misconfigured crawler or someone harvesting pages to resell. Pure noise.

Most people would block all of it. I am not going to.

Kimbundu is a Bantu language of Angola. On SIL's Digital Language Support scale it sits at "Still", the lowest of five levels. The scale measures how well a language works in digital systems, and Kimbundu barely registers. There is a dictionary published in the 1940s and not much else digitised.

When a language is that poorly digitised, the next generation of AI does not know it exists. It will not have the data, so it will guess. My job is to change that.

The most structured Kimbundu data on the web is now mine. Over 10,700 entries, digitised from that 1940s dictionary and updated to the modern ILN orthography. When a large language model trains, it learns Kimbundu from whatever it can find. Some of those crawler copies will end up in training data. That is distribution I do not have to pay for.

It is working in real time too. A month ago one or two people a day reached the site through ChatGPT. Last week it was 30 to 50. That is not training. That is ChatGPT citing the dictionary and sending people to it. A different pipeline, same effect. The data is getting used.

This project will never make me rich. I am building something people will come and take from, on purpose.

The mission is simple: make sure Kimbundu is a language the next generation of AI knows, not one it hallucinates.

So when the next crawler shows up, it is welcome to every page. That is the point.


More on why this matters: Building a Digital Kimbundu Dictionary in the Age of AI.