Other accounts:

@subignition (dead?)
@subignition

  • 6 Posts
  • 1.1K Comments
Joined 3 年前
cake
Cake day: 2023年11月1日

help-circle




  • Thanks, but I’m not sure what value you think you’re adding by actually going and finding that - if you had understood my post, you would know that my critique of the lack of specifics in the OP was a reason to doubt it was written by a human, not anything to do with whether or not the details of the post were credible or not.

    edit: the point is, including the filter enough was enough to make the post contents credible, but if it had specifically followed up with something like “and this is retrieved at src/lib/policies/system.ts:100” it would have seemed like it was written by a human and not an LLM.




  • Hi, not the person you replied to, and (edit: not*) agreeing with any of this blocklist shit, AND my ire is firmly not shifted away from Tesseract, but just wanted to break down the characteristics of the OP that suggest (to me) it was composed mostly with an LLM and fine tuned after. (Though to be clear, that’s a pretty unimportant characteristic of the post considering the actual content)

    1. The division of the text into small “sections” of five or fewer short paragraphs separated by headers is probably the most obvious hallmark.

    2. Kind of difficult to put into words, but abrupt / awkward / overly poetic metaphor and… “overconfident”? “sensationalized”? phrasing both give me an LLM vibe. examples:

    It’s a policy decision dressed as an API error, and it’s what had db0 chasing a version mismatch that never existed.

    I went through the code to see how that was implemented. The hardcoded list turns out to be the small half of the system.

    The spam work is what makes the rest unauditable: “it’s a spam list” answers every individual question and none of the whole.

    It’s downloaded at runtime from a file nobody has ever looked at.

    spam filtering and political editorial got welded into one undocumented, remotely-updatable blob, shipped hidden

    That’s the licence working as designed. But forking isn’t disclosure, and the admins who need this are precisely the ones with no reason to go looking.

    1. Overuse of the rule of three. This one is probably the most tenuous, because genuine human writing does make frequent use of the rule of three. But when LLMs do it, it’s usually at a high enough density where it starts sounding like it’s trying too hard, or it’s clickbaity.

    Admins can’t see it, can’t configure it, and aren’t told it’s there.

    That’s the live policy, base64-wrapped gzip, 111KB of JSON when it unpacks.

    There are Mastodon and Friendica accounts in there: people who have never used Lemmy, hidden by a Lemmy frontend, with no possible way of finding out.

    355 of the 552 usernames are plain alphabetic, twelve characters or under, median length eight.

    For the hidden users, communities and keywords, you get no message at all.

    Self-host it and you cannot disable this, nothing in your config admits it exists, and the contents can change without you pulling a commit.

    If you run Tesseract, you are relaying a 111KB moderation policy you have never read, under your instance’s name, to users who don’t know it exists.

    1. Tonally inconsistent with the intended audience. I realize Lemmy/Fediverse users are more technically savvy than usual, but even then… consider the claim “I went through the code to see how that was implemented.” early in the post. A lot of plausible technical jargon is used, and there are plenty examples of the actual filters and filter patterns themselves, but the post makes claims about what is happening without really being specific at all about how.

    Could be explained by me just not being familiar enough with the technical side, but saying stuff like “There’s a second filter policy fetched over HTTP every time the app loads” without breaking it down any further seems a bit suspicious to me? It doesn’t actually call out any function, file, or… i dunno, a logical starting point, for a competent reader to dig into the source themselves. Reads like it’s written for either an expert who doesn’t need any details, or for an audience with zero skepticism whatsoever.

    1. And least importantly, because my focus is on the presentation of the post moreso than the accuracy, IF rimu’s read of the code is correctly calling out a factual error in the OP (and let me be clear that I’m not enough of a programmer to evaluate whether it is), that is pretty strong evidence of an LLM getting something wrong.

    Again… this wall of text isn’t meant as a refutation of anything you’re saying, or a reason to dismiss the contents of the OP uncritically. Seems pretty clear after reading through db0’s posts about it that this is seriously fucked up. I hope I’ve been clear enough about that.









  • I mean, there’s a lot of ways bodies are dirty, but that just means you have to recognize the reasonable limits to cleanliness and learn to deal with what remains. (“shit happens”, after all)

    As someone who struggled a lot with black-and-white thinking when I was younger, I can definitely see how someone could develop some warped views in the absence of decent education / role models.