The first anti-AI tarpit. Rewritten in Python. Traps LLM crawlers in an infinite maze of fake pages and Markov babble. - NEPENTHESWEB/nepenthes-py

  • mesa@piefed.social
    link
    fedilink
    English
    arrow-up
    3
    ·
    1 天前

    Not sure about “first” theres a metric ton of these out there. Hell I made one with a honeypot and fail2ban years ago.

  • eleijeep@piefed.social
    link
    fedilink
    English
    arrow-up
    22
    ·
    2 天前

    This tool has two commits that have Made-with: Cursor tags in the commit message.

    I used to use a tool that did the exact same thing in around 2000-2002. It was called wpoison and in those days we used it to trap web-crawlers that were harvesting email addresses to add to spam lists.

  • Th4tGuyII@fedia.io
    link
    fedilink
    arrow-up
    12
    ·
    2 天前

    It’s sad that we’ve had to come to this.

    That the AI tech bros think they’re such hot shit that they shouldn’t have to play nicely with everyone else, and that the courts of the world seem to be largely backing them up.

    They’re being allowed to destroy and plunder human culture to feed their machines, things you would get thrown into the slammer for doing.

    So do whatever you have to do to preserve your corner of the internet. Whether that’s battening down the hatches and defending yourself with Cloudflare/Anubis (for the love of god don’t use reCapcha as you’re just giving Google data to train Gemini), or attacking them with something like Nepethes, then do it.

    The more effort they have to spend scraping, or time spent cleaning up their models, is more investor money thrown into the black-hole - and they’ve only got so much of it to spare.

  • moldy_rice@piefed.keyboardvagabond.com
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 天前

    This sounds like a bad idea. It’s what codeberg does, and it false positives feeding me gibberish content.

    How do humans get out of the tarpit when it false positives?

  • HaraldvonBlauzahn@feddit.org
    link
    fedilink
    arrow-up
    2
    ·
    2 天前

    How would such a Turing tarpit work that delivers good-looking, plausible, but very subtly non-functional code? So subtly that it takes an above-average developer to fix the result?

    • algernon@lemmy.ml
      link
      fedilink
      arrow-up
      1
      ·
      51 分钟前

      It simply wouldn’t. To poison models that way, you need a huge amount of such subtly wrong code. Preparing intentionally bad code at sych scale is impractical.

      It’s easier to show clear garbage. That won’t poison the model, but it gives them more work, and denies them access to good code, so they’ll train on slop.

      They’re doing a good job of sabotaging themselves, they don’t need our help - not more than simply not serving the crawlers.