| ▲ | walrus01 8 hours ago |
| There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month. |
|
| ▲ | mdasen 2 hours ago | parent | next [-] |
| That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher. Average electric rates by region: New England 28.1 cents
Mid Atlantic 25.1 cents
East North Central 20.8 cents
West North Central 14.8 cents
South Atlantic 16.1 cents
East South Central 15.5 cents
Mountain 14.6 cents
Pacific Contiguous 26.1 cents
Pacific Noncontiguous 42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... |
| |
| ▲ | AgentMatt 2 hours ago | parent [-] | | They are most likely not based in the US, but converting to USD to make comparison easier. | | |
|
|
| ▲ | ljlolel an hour ago | parent | prev | next [-] |
| can send safely context if there’s confidential computing ala my site https://trustedrouter.com/ |
| |
| ▲ | skinfaxi an hour ago | parent [-] | | How do you prove you are running exclusively on Nitro enclave instances or GCP confidential spaces? |
|
|
| ▲ | tiahura 10 minutes ago | parent | prev | next [-] |
| There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. Are there? At the highest levels of defense and law, AWS and Azure are used. Having tried selling some of these entities on doing things in-house, there seems to be little interest. |
|
| ▲ | fooker 7 hours ago | parent | prev | next [-] |
| Great, so the other member of the set matters for you more than cost. Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second? |
| |
| ▲ | sm-silversight 3 hours ago | parent | next [-] | | >Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second? Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation. | | |
| ▲ | walrus01 3 hours ago | parent [-] | | Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product. |
| |
| ▲ | walrus01 7 hours ago | parent | prev | next [-] | | Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output. On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code. | |
| ▲ | irishcoffee 44 minutes ago | parent | prev [-] | | The whole mentality of thinking one knows better than another about what they need causes infinitely more problems than it solves. |
|
|
| ▲ | light_hue_1 6 hours ago | parent | prev [-] |
| As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds. |
| |
| ▲ | walrus01 6 hours ago | parent | next [-] | | Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place or c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds? Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing. | | |
| ▲ | 30 minutes ago | parent | next [-] | | [deleted] | |
| ▲ | jmalicki 4 hours ago | parent | prev [-] | | A) is very doable with e.g. Amazon Bedrock. They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights. What kind of privacy needs do you really have beyond that? | | |
| ▲ | walrus01 4 hours ago | parent [-] | | It is not my use case but given recent political developments in international relations caused by the executive branch of the US government, off the top of my head, I could think of a lot of European or Canadian firms for which that would not be an option. No matter what they might promise about European sovereignty. For a good 'ol patriotic US domestic company? Sure. | | |
| ▲ | Taunt4 3 hours ago | parent [-] | | Yes, its the US cloud act risk EU companies run up against on hyperscalers like MS/AWS. Even for EU companies running open weights on EU stacks LLM inference on the GPU must process plaintext and I can't find any EU provider with NVIDIA H100/H200/Blackwell CC mode plus SEV-SNP or TDX, where you can cryptographically verify the workload ran somewhere the operator cannot inspect. Personal compute is therefore the only option if you want personal autonomy privacy for IP &c. Maybe another option is to use cloud compute rented to fine tune a personal model that suits your own needs that would help bring the cost down, I don't know enough about this area to know if it kills the "intelligence" of those domains due to limited ?cross-verification within the LLM. |
|
|
| |
| ▲ | solarengineer 6 hours ago | parent | prev | next [-] | | There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud. | |
| ▲ | trollbridge 3 hours ago | parent | prev | next [-] | | Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers? I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems. | |
| ▲ | frognumber 5 hours ago | parent | prev [-] | | Having worked in / adjacent several such industries, a lot of the question depends on scale. A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this. Worst-case example: Bootstrapped startup working in military. It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc. * Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. | | |
| ▲ | amluto 5 hours ago | parent [-] | | > Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised. I would tend to trust Amazon to avoid being compromised. |
|
|