I really hope they die soon, this is unbearable…

  • Thorry@feddit.org
    link
    fedilink
    English
    arrow-up
    51
    ·
    20 hours ago

    Yeah I had the same thing. All of a sudden the load on my server was super high and I thought there was a huge issue. So I looked at the logs and saw an AI crawler absolutely slamming my server. I blocked it, so it only got 403 responses but it kept on slamming. So I blocked the IPs it was coming from in iptables, that helped a lot. My little server got about 10000 times the normal traffic.

    I sorta get they want to index stuff, but why absolutely slam my server to death? Fucking assholes.

    • Ephera@lemmy.ml
      link
      fedilink
      English
      arrow-up
      15
      ·
      15 hours ago

      My best guess is that they don’t just index things, but rather download straight from the internet when they need fresh training data. They can’t really cache the whole internet after all…

      • Techlos@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        11
        ·
        15 hours ago

        Bingo, modern datasets are a list of URL’s with metadata rather than the files themselves. Every new team/individual wanting to work with the dataset becomes another DDoS participant.

      • Spice Hoarder@lemmy.zip
        link
        fedilink
        English
        arrow-up
        6
        ·
        13 hours ago

        The sad thing is that they could cache the whole internet if there was a checksum protocol.

        Now that I’m thinking about it, I actually hate the idea that there are several companies out there with graph databases of the entire internet.