Remix.run Logo
mghackerlady 18 hours ago

I mean, at this point Kagi could probably start their own, and there are plenty of smaller indexes (Mojeek, Yep, and Marginalia come to mind)

curiouscats 16 hours ago | parent | next [-]

Or create a new partnership with various like minded organizations and create a new index that each can use in their own way (perhaps those you mention plus DuckDuckGo and others that my brain isn't bringing to mind right now [maybe Wikipedia, EFF, internet archive, maybe Nebius would help...]).

I know this wouldn't be easy but maybe such an arrangement would be possible where there is some shared index and then the partners can each use it, extend it... in their own way.

mghackerlady 16 hours ago | parent [-]

I could see a collaborative open source web index run by the WMF, Internet Archive, EFF, GNU, DuckDuckGo, and Kagi being too big to simply ignore like a previous commenter suggested might happen to a new index. I could also see Qwant being interested

dewey 18 hours ago | parent | prev [-]

Bigger search engines have a lot of custom integrations that are not just crawling the clear web. You have a lot of benefits from collecting data for a longer time in this market.

I've often switched between Google and DDG but always came back to Google as in direct comparisons I was always not finding the right results in DDG. Kagi is the first time that I have not switched back a single time as nothing changed negatively in the quality and quantity of the results.

It's a noble goal to have your own search index, but it's duplicating a lot of work that others with much more resources already do well.

zargon 16 hours ago | parent [-]

In practice it is no longer actually possible to create a new search index with anywhere near the breadth of Google, regardless of the amount of resources you have available. Google (and to a lesser extant the couple of other historic engines) are entrenched enough to be allowed access by webmasters in robots.txt. And in the new era of mass abusive scraping by LLM companies, there's no longer the option to ignore robots.txt and scrape anyway, as sites do anything and everything they can to block any and all scrapers besides those already firmly entrenched.