Remix.run Logo
▲ enraged_camel 2 hours ago

>> And in the case of these open weight models: I can run it on my own infra and not give any data to anyone.

It's worth noting that the overwhelming majority of people who use Chinese models don't do this. Yes, it is nice to have the option, and there are US-based inference providers that claim to not send your data to China and maybe indeed don't, but in the grand scheme of things, we need to remember the adage that became popular during the social media era: if something is free (or, in this case, close to free), you are the product.

▲CharlieDigital an hour ago | parent | next [-]

Even if folks are not running their own inference infra, there are still services like Fireworks, AWS Bedrock, and others that are running the open models. I suspect anyone doing serious work with it is likely using a US hosted provider and I'd guess that by volume, US use of Chinese open models is using a US hosted platform (enterprise).

▲jacquesm an hour ago | parent | prev [-]

I actually do do this. I'm not sure who the 'overwhelming majority' is and where you got the data (link would be appreciated) but everybody that I know that runs these is doing so on their own infra.

▲enraged_camel an hour ago | parent [-]

Context is useful. The parent said: "it's a cheaper product that's almost as good or better in some cases"

The only open models that are "almost as good or better in some cases" require massive amounts of RAM. I posit that most people cannot afford a decked out Mac Studio, and therefore run the smaller "flash" variants on more normal devices. The issue is that those are nowhere near frontier-level in terms of capability.

▲CharlieDigital an hour ago | parent [-]

"Running your own infra" also includes managed infra like Bedrock, Foundry, etc.

Not just your local machines.

Enterprises are where you see this adoption. Legal, finance, tax; sensitive context where the data must be contractually opaque to external parties.

▲jacquesm 40 minutes ago | parent [-]

Precisely. I have figured out a nice recipe that is quite affordable, 288G of VRAM for a little under 20K, it takes some fiddling though, but once it works it is really neat.

PCIe is incredibly powerful tech.