• 2 Posts
  • 210 Comments
Joined 3 年前
cake
Cake day: 2023年6月11日

help-circle
  • Wikipedia has a search feature. Not many people use it, but Wikipedia is huge, so enough people do. Back then, a couple dozen full text searches a second. And maybe 100 quick prefix searches.

    Anyway, one thing I made for it was a way to search the wiki text with regular expressions. It’s useful for maintainers.

    Wikitext is the funny markup language Wikipedia uses. If you know html or md it’s like that, but it can kind of “call” other pages. Whatever. Doesn’t matter. It’s text. And people have strong opinions about it.

    You can see it by clicking “edit” (sometimes looks like a pencil) and then clicking “edit without logging in” then “source editing”. Those lovely boxes at the top for a person are made with something like

    {{Infobox person
    | name               = Douglas Adams
    | birth_place        = [[Cambridge]], England
    }}
    

    Imagine you have 1000 articles that use birthplace instead of birth_place. Both work, but it’s driving you and all your editor friends crazy there are two ways to do it. This is believable. Trust me.

    You want to find every page that has an {{Infobox person followed by birthplace. Then fix them. The nerds who came before us invented a language to ask that question called “regular expressions”. The one for this looks like \{\{Infobox\wperson.+birthplace. Maybe. Nerds will know regular expressions are bad for this. But nerds will also know that regular expressions being bad has never stopped anyone from using them anyway.

    One way to run these regular expressions is to convert them into an https://en.wikipedia.org/wiki/Nondeterministic_finite_automaton . You build these in memory and they are not big. But you can’t really run them directly. They are pretty and fairly easy to read once you get used to them. But you can’t easily ask “does this match this text”. At least, not with the tools I had.

    I could only run https://en.wikipedia.org/wiki/Deterministic_finite_automaton . It’s deterministic! Much nicer. And you can go from a nondeterministic one to a deterministic one. Easy. Our forenerds solved the problem. You use a https://en.wikipedia.org/wiki/Powerset_construction .

    The trouble is, it can make very very very big deterministic finite automata. Like, if there are 3 states in the nfa you can get and 8 state dfa. 4 is 16. 5 is 32. In the worst case. Usually you get much better. But a fairly big but not super frightening regular expression can turn into a big nfa. And the dfa would need more states then there are grains of sand on earth. Too big for computer.

    So, user asked for this regular expression. The search servers tried to convert it, and filled up their memory and died. They ran Out Of Memory. OOM. The usual thing is you copy the full memory so you can look later and start the server again.

    But that can take time. And sometimes people are confused and haven’t set up the restart to be automatic. And they can’t find you. So search stays broken for a while.









  • I live in the US South and “folks” is a normal word for “people” here. I’m intentional not looking it up and trying to go from memory to get my dialect. I don’t use it for “humanity”, but will for any subset. “The folks who live in that house.” “White folk.” “City folk.” I dunno. There’s a shit ton of dialects in the US South East and they differ by geography and income and race and all kinds of stuff. And maybe they aren’t dialects. I’ve heard “register”. I’m not a language knower.

    Should you use the term “black folks”? Ask a couple black folks if it makes em feel bad. Do what they say. Don’t listen to this white boy.

    It’s built into my dialect so I’d use it. But if folks told me it made them feel bad if stop.

    I have no idea when I pluralize it and when I don’t.