| ▲ | johnorourke 5 hours ago |
| Anubis[1] is a superb fix for sites not behind Cloudflare/Fastly/Bunny etc. We had millions of bot requests, on a site serving all countries so we couldn't block by country, with fake user-agents so we couldn't block using that. It uses 'proof of work' to detect real browser software. [1] https://anubis.techaro.lol/ |
|
| ▲ | basilikum 4 hours ago | parent | next [-] |
| It doesn't detect "real browsers". It simply adds some friction and is niche enough that AI scrapers have not bothered to bypass it yet. |
| |
| ▲ | leros 4 hours ago | parent | next [-] | | Yeah this is one of those things that hurts some real users and AI scrapers have figured out how to bypass. Not worth it IMO. | | |
| ▲ | mostlysimilar an hour ago | parent [-] | | How do you figure it hurts real users? The amount of compute/energy used on the proof of work is pretty minimal. You're using more when you watch a YouTube video or browse a JS-heavy web app. Of course a sophisticated scraper can "figure out" how to bypass. It isn't trying to be foolproof, it's adding an extra cost to deter massive amounts of bot traffic. I put it in front of my hobby project because I can't afford to serve hundreds of thousands of bot requests from residential proxies all across the world, and I didn't want to route all of my traffic through a third party company like Cloudflare. I've been happy with Anubis. |
| |
| ▲ | axus 2 hours ago | parent | prev | next [-] | | There is enough visceral hatred for the Anubis branding, I am surprised an AI skill for bypassing it hasn't been broadcast yet. | |
| ▲ | randomblock1 an hour ago | parent | prev [-] | | Most of the friction is just JS overhead for the computations, a compiled solver is like 1000x faster. If Anubis ever gets popular enough that scrapers care, it would be trivial to defeat. And last I checked you could bypass it by just modifying the user agent |
|
|
| ▲ | RattlesnakeJake 5 hours ago | parent | prev | next [-] |
| I wish they'd ease up on the whole "don't change the logo without paying us" thing. The furry anime character is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures. |
| |
| ▲ | vablings 4 hours ago | parent | next [-] | | It's actually a genius idea. If you are someone who the professional presentation of not having an anime girl on the loading page is required, then you can afford to fork over the cash to fund development. | | |
| ▲ | zuzululu 39 minutes ago | parent [-] | | I guess the disconnect here is a bunch of HN'ers believing professional companies and websites want to attach their branding to a sexualized anime character and that they are willing to pay to remove it. Which one then wonders why they would install it in the first place. |
| |
| ▲ | altairprime 4 hours ago | parent | prev | next [-] | | > is a turnoff for anyone with a brand or personal image that doesn't mesh with those subcultures. Most people either don’t know or don’t care about “those subcultures”. I bet a lot of older people think it’s a cartoon figure of Betty Boop (nurse) and miss the furry bit since it appears and disappears quickly. Most people also don’t have a brand. So it seems like you’re describing a concern that only affects a tiny fraction of people: - Not interested in paying for custom branding, so obviously not a corporation or influencer - Dislikes cartoons - Aware of, and hostile towards, “furry” subculture That has to be an exceedingly small fraction of potential users of Anubis, and given how much businesses and branders will pay to custom-brand something, I’d counsel them to stay the course. Sure, a few never-payers will never pay, but they wouldn’t have anyways, so they can cope with Nurse Betty or look elsewhere for a competing free product. If you think about this in physical market square terms — in other words, a bazaar — it seems horrendously rude to complain about a shop logo sticker on a free product handed out to anyone that walks up and asks for it. If you want it white-labeled so you can write your own name/logo on it, you pay for the privilege of displacing their name with yours. But you don’t stand there and loudly complain that their shop mascot has dog ears while holding a freebie bag of product, without losing the respect of everyone who hears you doing so. | |
| ▲ | teddyh 4 hours ago | parent | prev | next [-] | | If you have a brand or personal image that you are investing in, you can surely afford to invest in paying for a branded version of Anubis. | |
| ▲ | miladyincontrol 3 hours ago | parent | prev | next [-] | | I'd argue its rather functioning exactly as intended with regards to obtaining paid users. Also this view seems a bit elderly. For most towards the end of the millennial curve and younger, anime is no longer subculture, its just general culture at this point. Although I would agree it isnt necessarily what you'd want for every platform and web presence. | |
| ▲ | pibaker 22 minutes ago | parent | prev | next [-] | | How hard is it to maintain a fork that changes nothing except the logo? If you can't be bothered to maintain a trivial fork, then why should the author of anubis be bothered to serve your branding needs? It's not like you have a service contract or anything do you? | |
| ▲ | gfaster 4 hours ago | parent | prev | next [-] | | I think that's partly why they do it? If you care about that, you should pay? | | | |
| ▲ | ivanjermakov 2 hours ago | parent | prev | next [-] | | Pay to change logo??? Just fork Anubis and remove one div... It's not SaaS first software, you're expected to deploy it on your web server. | |
| ▲ | xena 3 hours ago | parent | prev | next [-] | | I wish my rent would stop going up year over year. | |
| ▲ | inigyou 4 hours ago | parent | prev | next [-] | | They're not stopping you. They're asking you politely not to. | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | ryan_n 4 hours ago | parent | prev | next [-] | | You just discovered their business model congrats. | |
| ▲ | RobotToaster 4 hours ago | parent | prev | next [-] | | haproxy-protection is an alternative. | |
| ▲ | pc86 4 hours ago | parent | prev | next [-] | | It's MIT licensed, you are free to do whatever you want to it, including removing the logo. They're basically just saying "we'd prefer you didn't do this, but we're not preventing you from doing it." | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | bakugo 3 hours ago | parent | prev | next [-] | | It's open source, just ask your AI agent to change the logo. | | |
| ▲ | xena 3 hours ago | parent [-] | | Just keep in mind that changing the logo is a great way to move yourself down the priority list for bug reports and support. | | |
| ▲ | bakugo 2 hours ago | parent [-] | | Someone willing to take the 2 minutes to ask Claude to change the logo can probably also ask it to fix any bugs they find, or add any new features they might want. |
|
| |
| ▲ | brazukadev 3 hours ago | parent | prev | next [-] | | that's really smart of them. Just pay to use the service you need. | |
| ▲ | peter_stokes 4 hours ago | parent | prev | next [-] | | you really revved the weebs with this one | | |
| ▲ | duskdozer 3 hours ago | parent [-] | | Nah, I don't especially like the logo myself when I come across those sites, but I do like that "brand" people who want to make money off it and not pay dislike it even more. |
| |
| ▲ | righthand 4 hours ago | parent | prev | next [-] | | You wish they'd give your preferable imagery for free and not make you feel bad for using their free software for personal gain. Your subculture is irrelevant and not special. | |
| ▲ | 1bpp 4 hours ago | parent | prev [-] | | If you 'don't mesh' with that then you are not worth protecting anyway :) | | |
|
|
| ▲ | D2OQZG8l5BI1S06 3 hours ago | parent | prev | next [-] |
| Because of this wasteful crap the internet is so slow nowadays... Try opening gcc bug tracker on your phone: https://gcc.gnu.org/bugzilla/ |
| |
| ▲ | esperent 3 hours ago | parent | next [-] | | This is like getting angry at cookie banners instead of all the companies tracking and selling your data. You're complaining about the symptom (needing to have these checks) not the cause (if they don't, 99% of their traffic will be bots, the site will slow to a crawl and be unusable anyway). In any case, I saw the dumb anime girl for about 2s then the site loaded. Not a big deal. | |
| ▲ | cuu508 2 hours ago | parent | prev | next [-] | | The Anubis challenge took ~5 seconds. What do you propose instead? | |
| ▲ | xena 3 hours ago | parent | prev [-] | | The GCC bug tracker uses the meta-refresh challenge, which does not require JavaScript. Due to the fact that the server makes sure the client has waited at least 75% as long as it should, the HTML has to add one second to the meta-refresh wait. Patches welcome. Meta refresh granularity is in single digit seconds. |
|
|
| ▲ | wbl 4 hours ago | parent | prev | next [-] |
| Anubis sucks because CPU is cheap for scrapers and hard for humans. |
| |
|
| ▲ | marklar423 4 hours ago | parent | prev | next [-] |
| I'm assuming a bot running a headless browser instance can still get past it? It's still valuable to raise the cost of scraping of course. I don't think anything can really stop a determined scraper from impersonating a human. I wonder though if a system similar to Anubis but mining some crypto would make bots _welcome_ - since they're paying for their traffic. |
| |
| ▲ | drum55 4 hours ago | parent [-] | | People tried this in 2013 or so, there's no point to it. Doing proof of work in javascript in a browser is so crushingly, pointlessly slow that there's no value at all. Some browsers also intentionally detect attempts to do proof of work and attempt to block it entirely. | | |
| ▲ | RobotToaster 4 hours ago | parent [-] | | > Some browsers also intentionally detect attempts to do proof of work and attempt to block it entirely. Then they'd be blocking themselves from the website. | | |
|
|
|
| ▲ | czk 4 hours ago | parent | prev | next [-] |
| make your only legit users mine fake crypto to access your site, only costs them 5% battery on an android device |
| |
| ▲ | econ 4 hours ago | parent [-] | | They think scraping = money so you can just ask them to scrape (copy paste websites into a text area) and it will feel like payment. |
|
|
| ▲ | drum55 5 hours ago | parent | prev [-] |
| Which is trivially bypassed by an actual implementation of the proof of work in non-javascript, rendering it absolutely useless. The website is approximately 3800x times slower than native code, and hundreds of thousands of times slower than the CUDA kernel claude wrote. The "proof of work" is just non existent at that point, they're solved in milliseconds for what would take the browser version 10 minutes or more, it's security by obscurity being dressed up as something more. pow_server http://127.0.0.1:8080 backend avx512-x16
──────────────────────────────────────────────────────────
uptime 00:03:12
solver ● BUSY difficulty 9, 0.3s
queue [####################............] 5/8 peak 12
──────────────────────────────────────────────────────────
accepted 1240 solved 1180
503 shed 48 504 timeout 2 4xx/5xx 10
──────────────────────────────────────────────────────────
last difficulty 5 nonce 645376 in 9 ms (101.6MH/s, avx512-x16)
hashes 3.90GH total avg 65.3MH/s
Ctrl-C to stop
Claude even made a nice little API server for it after implementing midstate compression, AVX multi way hashing, and a CUDA kernel. This doesn't stop the literal LLM it's trying to block from solving the challenges, it's really annoying that everybody is using it and claiming that it's something that's usable in the real world as a result of it using proof of work. It's obscure, and obscure is fine so long as nobody is pretending that it is secure. |
| |
| ▲ | harshreality 4 hours ago | parent | next [-] | | It is not absolutely useless, empirically, which you'd discover if you had a website getting hammered by bots and experimented with anubis as a countermeasure. While dedicated scrapers/attackers could work around it, and they could do so much more efficiently than the client-side js, almost none of them do. Unless you like paying additional hosting resource fees to serve bots, it's a worthwhile option, and is less annoying to typical human visitors than cloudflare's interactive captcha/challenge which is what most people use. The main author is aware that the algorithm is far from ideal for this purpose. See https://news.ycombinator.com/item?id=48869064 . If more bots start to answer the primitive challenge anubis uses now, that'll hasten implementation of a different algorithm. Don't let the perfect be the enemy of the good enough. For now, the algorithm or challenge scheme almost doesn't matter. Since it's much smaller-scale than cloudflare's challenges, that's probably why very few scrapers and botnets bother to solve anubis's trivial sha2 pow. Targeted attacks may not be repelled at all. That's not the point. | | |
| ▲ | kro 4 hours ago | parent [-] | | It does not even require the PoW thing Anubis does. I've setup a simple logic that just: Checks for existence of a specific static cookie, if it does not exist, output a small page that sets the cookie via JS and reloads. Sadly this kills Noscript, but it would be possible to add a form in <noscript> that when submitted sets the cookie serverside. Is this trivial to bypass? Yes. It still keeps out 95% of unwanted bots.
Reality is most do not target you specifically they just want to mass-scrape with low effort. Running headless browsers is way more expensive for their op I've extended this with a FCRDNS checked exclusion for Googlebot. Another quite effective measure I figured out was checking the existence of Sec-Fetch-Dest header if the User-Agent claims to be a modern browser. If you don't want to close down too much. Also, I only apply these rules to routes that are not cheap and cached. | | |
| ▲ | harshreality 3 hours ago | parent [-] | | That's not far from what anubis does for clients that are determined to have light souls. It doesn't always send a PoW challenge. For a webapp that sets a long-lived cookie, that cookie could be used to bypass anubis completely, or lower the weight in anubis so that it doesn't send its pow challenge unless there are major red flags. If bots start to abuse that exception, it can be removed. |
|
| |
| ▲ | inigyou 4 hours ago | parent | prev | next [-] | | But the people you're defending against don't do that. They also don't load CSS but for some reason the security theater PoW won the mindshare. | | |
| ▲ | econ 4 hours ago | parent | next [-] | | I once discover you can put escaped XML or json in css content. The purpose was to have static data sets that work cross domain. No headers to configure no letting strangers run all you can eat malicious js on your site. | |
| ▲ | antonvs 2 hours ago | parent | prev [-] | | Security theater is what gets the economic rewards. |
| |
| ▲ | rokkamokka 3 hours ago | parent | prev | next [-] | | Like any lock, it's mainly to deter less determined adversaries (which account for the vast majority) | |
| ▲ | Galanwe 4 hours ago | parent | prev | next [-] | | The point of PoW access is not that its hard to bypass, it's that you cannot bypass it at scale. | | |
| ▲ | gruez 4 hours ago | parent | next [-] | | >it's that you cannot bypass it at scale. Define "scale". For any reasonable wait that you're willing to impose on your users, any PoW scheme heavily favors attackers. They have unlimited time and can be scraping even while they're asleep. Your visitors on the other hand don't have that luxury. You might argue that's not the point and it's only to stop dumb scrapers that are effectively ddosing your site, but if it's just dumb scrapers, you could've stopped them less onerous measures like tls or javascript fingerprinting. | |
| ▲ | petu 4 hours ago | parent | prev [-] | | If algorithm used is static and GPU-friendly, then what stops bypass at scale? | | |
| ▲ | drum55 4 hours ago | parent [-] | | It's more or less designed for it, it's SHA256 with a break in the middle for midstate compression to be effective, and the difficulty system is based on a misunderstanding of how bitcoin PoW works ("number of zeros" is never, ever a consideration in bitcoin, it's a match to a floating point target). sha256(challenge + ascii(nonce)) means that the first compression round of the function can be cached and the second compression round is just the nonce plus the cache. This is the same trick used in Bitcoin mining and would have been avoidable by putting the nonce first, so immediately any non-naive code has to do half the proof of work as the vanilla solver. |
|
| |
| ▲ | gum_wobble 4 hours ago | parent | prev [-] | | how so, can you link to any sources? | | |
| ▲ | drum55 4 hours ago | parent [-] | | The prompt used for Opus 4.8 was: write a implementation of the anubis proof of work in native c code, optimized for speed above all else. use every trick available to make the proof of work as efficient and fast as possible, including modern processor tricks on the x86 platform. your code should avoid using external libraries where possible, include tests, and be readable and concise. a reference for what needs to be met is in this repository. https://github.com/TecharoHQ/anubis
Then let’s develop this more. turn this solver into a local HTTP server that can be given work in the request, and it returns solved work. make an end to end tester that sends test work to the solver and waits for a valid response. add support for solving with a GPU using cuda.
Then it was done more or less, it happily made a local server that supports solving the challenges given to it in bulk with priority based queue and can tolerate potentially tens of thousands of requests a second with no issue. The CPU time spent solving the challenges is less than the SSL setup for the connections. The GPU version does in excess of 20GH/s (but with high latency) though I didn't really test it, I'm not using this for anything but proving a point that the LLM itself can write the bypass tools and run them happily. | | |
| ▲ | harshreality 4 hours ago | parent [-] | | In a thread last month about scrapers, the author mentioned working on a switch to hashx.[1] In addition, nothing prevents anubis from sending a wasm solver instead of js, reducing the gap between a custom native solver and a js solver. [1] https://news.ycombinator.com/item?id=48869064 | | |
| ▲ | xena 3 hours ago | parent | next [-] | | I'm almost ready to ship the wasm feature in the next version of Anubis after the one that's about to come out. The big blocker is that testing against dozens of googles chrome to ensure functionality on abandoned smart TV oses takes a nontrivial amount of time. As an example of the level of debugging and the like required: https://github.com/TecharoHQ/anubis/pull/1684/changes/67621f... My office gets very warm when chromesweep runs. This is something that is complicated enough that even though LLM tools can help, it's not a magic bullet. It's just complicated in general. | |
| ▲ | kijin 3 hours ago | parent | prev [-] | | They should make it mine actual coins for the site owner. The more bots try to access the site, the more profitable it will be! |
|
|
|
|