| ▲ | derefr 2 days ago | |
It's obvious that the companies exposing search features for their own service data don't actually want you to be able to search the data with precision. They 1. think they know better than you, and so want to show you stuff besides "what you think you want", because their metrics say you might click it anyway; and 2. they want big long listings for you to scroll through, so they have more opportunities to inject paid promotional elements into those listings. It's just another kind of enshittification. That being said... the first wave of web search engines weren't built by the companies hosting the content. Search engines as tools became popular because, even in 1998, browsing and navigating sites (esp. corporate sites) to surface content, was already becoming an increasingly adversarial experience. Search engines were built by companies spidering other sites' content. This was content that was often — if you were navigating along these sites' happy paths — found five to ten links deep through a confusing warren of subtle and unintuitive click targets. This content was the original "deep web." And search engines made it shallow... often against the spidered sites' consent. Companies at the time much preferred "directory" sites that would only ever link to their landing pages. To companies of that era, sites linking directly to specific URLs of your site would be like if the Yellow Pages listed specific directory-extensions of your company's phone number! But this first battle in the "war against deep-linking" was a losing one from the moment it began, since early websites were almost inherently bot-accessible. The second battle, in the early 2000s, where sites were built as [non-deep-linkable] Flash apps, took longer to determine, only finally resolving (again in favor of an open web, and thus third-party search) when laws and regulatory compliances both started to force companies back toward UA accessibility. But for going on a decade now, we've been embroiled in a seemingly-indefinite third battle in the "war against deep-linking" — this time with companies tucking every possible type of user-generated content behind a login wall of some walled-garden everything [web]apps. Permalink URLs for individual content-items still exist in these webapps; but you get redirected to a login interstitial if you visit them. Where's the adversarial spirit of the first search-engine companies today? Where are the companies trying to index the modern "deep web" of login-walled content? Why can't any company give me a search box that searches "into" YouTube video transcripts, "into" Facebook Marketplace, "into" public Slack and Discord and Telegram communities (probably requiring archiving, ala how Deja News/Google Groups surfaced Usenet), etc? Sure, this sort of thing is intensely adversarial (but see https://en.wikipedia.org/wiki/HiQ_Labs_v._LinkedIn — while it might be against a site's ToS, but it's not against the law.) And sure, this sort of thing requires per-service connectors (but any thought of this business model therefore being "unscalable" comes from a pre-coding-agent era.) But isn't it also, very clearly, "what people want"? | ||