The first anti-AI tarpit. Rewritten in Python. Traps LLM crawlers in an infinite maze of fake pages and Markov babble. - NEPENTHESWEB/nepenthes-py
Not sure about “first” theres a metric ton of these out there. Hell I made one with a honeypot and fail2ban years ago.
Yeah there was yet another whose name I can’t recall now. Some purple banner on the website…
This tool has two commits that have
Made-with: Cursortags in the commit message.I used to use a tool that did the exact same thing in around 2000-2002. It was called wpoison and in those days we used it to trap web-crawlers that were harvesting email addresses to add to spam lists.
It’s sad that we’ve had to come to this.
That the AI tech bros think they’re such hot shit that they shouldn’t have to play nicely with everyone else, and that the courts of the world seem to be largely backing them up.
They’re being allowed to destroy and plunder human culture to feed their machines, things you would get thrown into the slammer for doing.
So do whatever you have to do to preserve your corner of the internet. Whether that’s battening down the hatches and defending yourself with Cloudflare/Anubis (for the love of god don’t use reCapcha as you’re just giving Google data to train Gemini), or attacking them with something like Nepethes, then do it.
The more effort they have to spend scraping, or time spent cleaning up their models, is more investor money thrown into the black-hole - and they’ve only got so much of it to spare.
120%. If I had the funds to buy the hardware, I’d have a whole room dedicated to running 500 instances of these alone. xD
This sounds like a bad idea. It’s what codeberg does, and it false positives feeding me gibberish content.
How do humans get out of the tarpit when it false positives?
Thank you, added to https://noailist.org/anti-ai-projects/
check the other comments here
Ironic that this is posted on github.
Yeah… :/
How would such a Turing tarpit work that delivers good-looking, plausible, but very subtly non-functional code? So subtly that it takes an above-average developer to fix the result?
It simply wouldn’t. To poison models that way, you need a huge amount of such subtly wrong code. Preparing intentionally bad code at sych scale is impractical.
It’s easier to show clear garbage. That won’t poison the model, but it gives them more work, and denies them access to good code, so they’ll train on slop.
They’re doing a good job of sabotaging themselves, they don’t need our help - not more than simply not serving the crawlers.









