Remix.run Logo
▲ simonw 6 hours ago

My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.

If I'm running an agent system but I'm not allowed to store the responses - or provide a "share transcript" button - that's a pretty significant limitation.

The answer to that question is inevitably buried deep in the terms. Here's the relevant section I found for Ceramic, in their list of things you can't do:

> (n) collect, aggregate, store, or compile Output, including search results, relevance scores, or rankings, for the purpose of creating or contributing to any database, dataset, index, or corpus, whether or not such database, dataset, index, or corpus is used for a purpose that competes with Ceramic; (o) resell, syndicate, or otherwise make Output available to any third party on a standalone basis or as a separately accessible component of another product or service; provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query, and is not independently accessible, extractable, or downloadable by end users or third parties; or (p) retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use, unless expressly permitted in an applicable Order Form.

https://www.ceramic.ai/terms-of-service

Am I alone in caring about this?

▲infogulch 4 hours ago | parent | next [-]

It seemed like this part gives you the exception you wanted:

> provided that you may display Output to your authorized end users within your own application so long as such Output is integrated into your application's functionality, is incident to the end user’s real-time query ...

but it continues:

> ... and is not independently accessible, extractable, or downloadable by end users or third parties

How can you prevent end users from extracting it if its visible? Why even have the exception if you just throw it out with an impossible to meet restriction like this?

▲Loquebantur 3 hours ago | parent | next [-]

So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

The weird attitude in the Internet Tech company scene is akin to Gold Rush scenarios.

Who are the native people?

▲slowpoison an hour ago | parent | next [-]

Don't think it's unreasonable. They crawled it and made it available in an easy to digest form. You can't take their copy and start distributing it infinitely while paying them for one-time read. Write your crawler, and crawl the web on your own if you want... and give it away!

▲Loquebantur an hour ago | parent [-]

You glaze over the fact, they stole the data to begin with and never recuperate their sources. That's not "reasonable".

They don't make it available "easily" either, they place all kinds of hurdles on it, including you having to pay for stolen goods.

You then go on, equating wildly different situations: on one hand, a multi-billion dollar company, easily able to set up such a scheme. On the other (me representing) the public, regularly anything but destitute.

So no, you're the unreasonable one here. So are they.

▲sroussey 29 minutes ago | parent [-]

Search engines that obey robots.txt are stealing?

If someone does not want on a search engine index and they say not to index, and then get indexed anyhow is stealing.

But a “please read my content and make available to your users” then claim doing so is stealing seems a bit out there.

What am I missing?

▲atmosx 3 hours ago | parent | prev | next [-]

> So they crawled the web, stole the information to populate their own database and then pretend it was "illegal" for others to steal it back?

These days, this seems to be "modus operandi". The bet is who can get closer to the administration to suddenly enforce the un-enforce-able. For your own safety. You wouldn't steal a car now, would you?

▲Loquebantur 2 hours ago | parent [-]

To acquiesce to manufactured "social truths", to resign complacently to them being "modus operandi" means to place yourself outside of the society you live in.

You make yourself subject to a social reality you presume outside of your control. But you enabled them, if by nothing else, by your silent acceptance.

Moral and ethical judgements cannot be left to the very same people they are supposed to restrain in the first place.

▲debo_ an hour ago | parent [-]

heavy metal guitar chords commence

▲Loquebantur an hour ago | parent [-]

You engage in ridiculing your own demise. Why so sure you'll enjoy it?

▲DANmode an hour ago | parent | prev | next [-]

“You're looking at 'em...”

- Tony Soprano

▲measurablefunc 3 hours ago | parent | prev [-]

The data is the moat, the compute is a cost center.

▲foota 4 hours ago | parent | prev [-]

Not to mention: "retain, cache, or store Output beyond what is reasonably necessary to display such Output to your authorized end users in the ordinary and real-time course of use" which would seem to preclude storing it in a long lived session.

▲simonw 4 hours ago | parent [-]

The phrase "reasonably necessary" is infuriatingly vague.

▲cj 5 hours ago | parent | prev | next [-]

My general stance on things like this is to think about the intent -- why does the company have that in their TOS. Use that as a proxy for assessing the likelihood of the company enforcing the terms against you.

▲derac 3 hours ago | parent [-]

This is only valid up to the level of risk you can tolerate for them pulling the rug out from under you.

▲btown 2 hours ago | parent [-]

Which is why, as much as I love Cloudflare, I wouldn't route things like this through their billing. The risk that a company's entire infrastructure goes down, perhaps even by a fraud/risk flag by an incorrectly-configured AI (including, say, if the search API providers are back-sharing their own potentially-broken abuse flagging metadata with Cloudflare), is far too great.

▲sanderjd 6 hours ago | parent | prev | next [-]

You are not alone, I agree that this is a frustrating limitation.

▲ 2 hours ago | parent | prev | next [-]
[deleted]
▲phoghed 3 hours ago | parent | prev [-]

It’s a shit tier web scraping startup, just violate their terms, who cares.

▲writtenone 3 hours ago | parent [-]

This is the right way to think about it. If there's any fear of getting caught, just use a reputable VPN or one of the hundreds of residential proxy providers.

▲kokanee 2 hours ago | parent [-]

The whole point of this product from Cloudflare is to let LLMs "top up" their corpus of knowledge with up-to-the-minute search results after they have been trained on the contents of the entire Internet. The idea that courts would enforce intellectual property rights on little upstarts trying to make LLM wrappers without enforcing any TOS affecting the massive training scrapers is ridiculous. Probably true, but logically unjustifiable.