| ▲ | walrus01 7 hours ago |
| It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes. Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware. Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB. I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8. |
|
| ▲ | fooker 7 hours ago | parent | next [-] |
| > Even if the output is like 5-6 tok/s, that might be usable for some purposes. You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second. I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn't be surprised if this is an empty set. |
| |
| ▲ | walrus01 7 hours ago | parent | next [-] | | There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month. | | |
| ▲ | mdasen an hour ago | parent | next [-] | | That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents. There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher. Average electric rates by region: New England 28.1 cents
Mid Atlantic 25.1 cents
East North Central 20.8 cents
West North Central 14.8 cents
South Atlantic 16.1 cents
East South Central 15.5 cents
Mountain 14.6 cents
Pacific Contiguous 26.1 cents
Pacific Noncontiguous 42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph... | | |
| ▲ | AgentMatt an hour ago | parent [-] | | They are most likely not based in the US, but converting to USD to make comparison easier. | | |
| |
| ▲ | ljlolel 30 minutes ago | parent | prev | next [-] | | can send safely context if there’s confidential computing ala my site https://trustedrouter.com/ | |
| ▲ | fooker 6 hours ago | parent | prev | next [-] | | Great, so the other member of the set matters for you more than cost. Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second? | | |
| ▲ | sm-silversight 2 hours ago | parent | next [-] | | >Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second? Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation. | | |
| ▲ | walrus01 2 hours ago | parent [-] | | Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product. |
| |
| ▲ | walrus01 6 hours ago | parent | prev [-] | | Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output. On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code. |
| |
| ▲ | light_hue_1 5 hours ago | parent | prev [-] | | As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds. | | |
| ▲ | walrus01 5 hours ago | parent | next [-] | | Do you mean by trading dollars for the privacy you need as: a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place or b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place or c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds? Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing. | | |
| ▲ | jmalicki 4 hours ago | parent [-] | | A) is very doable with e.g. Amazon Bedrock. They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights. What kind of privacy needs do you really have beyond that? | | |
| ▲ | walrus01 4 hours ago | parent [-] | | It is not my use case but given recent political developments in international relations caused by the executive branch of the US government, off the top of my head, I could think of a lot of European or Canadian firms for which that would not be an option. No matter what they might promise about European sovereignty. For a good 'ol patriotic US domestic company? Sure. | | |
| ▲ | Taunt4 2 hours ago | parent [-] | | Yes, its the US cloud act risk EU companies run up against on hyperscalers like MS/AWS. Even for EU companies running open weights on EU stacks LLM inference on the GPU must process plaintext and I can't find any EU provider with NVIDIA H100/H200/Blackwell CC mode plus SEV-SNP or TDX, where you can cryptographically verify the workload ran somewhere the operator cannot inspect. Personal compute is therefore the only option if you want personal autonomy privacy for IP &c. Maybe another option is to use cloud compute rented to fine tune a personal model that suits your own needs that would help bring the cost down, I don't know enough about this area to know if it kills the "intelligence" of those domains due to limited ?cross-verification within the LLM. |
|
|
| |
| ▲ | solarengineer 5 hours ago | parent | prev | next [-] | | There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud. | |
| ▲ | trollbridge 2 hours ago | parent | prev | next [-] | | Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers? I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems. | |
| ▲ | frognumber 4 hours ago | parent | prev [-] | | Having worked in / adjacent several such industries, a lot of the question depends on scale. A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this. Worst-case example: Bootstrapped startup working in military. It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc. * Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. | | |
| ▲ | amluto 4 hours ago | parent [-] | | > Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security. If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised. I would tend to trust Amazon to avoid being compromised. |
|
|
| |
| ▲ | btown 16 minutes ago | parent | prev | next [-] | | One aspect of this is cyberattack proliferation by way of "Hey boss, I saw this TikTok that says if you let me invest [a tiny piece of the neighborhood's profit|our militia's budget] into some RAM, I could get a fully autonomous cyber operation up and running that pays for itself via ransomware etc. within weeks. You like it, we upgrade to something that can work even faster. We don't need the hacker guy from Swordfish with fifty monitors, we just need my cousin who likes building gaming PCs." That's a world that I don't think we're ready for. | | |
| ▲ | 9x39 3 minutes ago | parent [-] | | A similar world is already here. Young men 14-?? already compromise and attempt to extort organizations daily, sometimes cluelessly from western nations, often not. It doesn’t have to be gangs when the home country doesn’t care / isn’t technologically or culturally developed. Already seeing AI-written payloads and frameworks in the wild. I think it’ll turn out that AI won’t build you a maintainable ERP but it can create C2 networks, exploit POCs or even 0-days potentially, and let kids make their own ransomware tooling. Then we’re dealing not with a handful of cybercrime tool makers but a generational problem. |
| |
| ▲ | cobbzilla 4 hours ago | parent | prev [-] | | I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower. I don’t know if that’s 100x more than I’d pay (opex-wise) with an nvidia setup, but I can say the one-time capex is a great deal cheaper. Avoiding VRAM and DDR5 (fast DDR4 should be OK) are the biggest cost savers. ECC RAM is worth the extra price. General datacenter-quality hardware has less price sensitivity, and plenty of bang for your buck. | | |
| ▲ | walrus01 4 hours ago | parent | next [-] | | Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour. Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the RAM bus speed. An older 2U rack server that is 600W load (typically a 1+1 power supply server when plugged into two kill-a-watt would show 300W on each, equal load balancing) when maxed out with stress-ng might be only 450W total running inference. If you have 600kWh used in a month by running something 24x7 and your power is $0.15 a kWh, that's more like $90/mo (not counting cooling or any ancillary costs for the environment where it's in). | | | |
| ▲ | trollbridge 2 hours ago | parent | prev | next [-] | | If you actually were running this thing at 80% or 100% load, then the first thing you'd want to is get a better PDU and then connect your servers to that (48V DC). | |
| ▲ | sebmellen 4 hours ago | parent | prev [-] | | (Context: Parent comment was edited after I wrote this comment) Where in the world are you finding that much RAM in a racked server for $200/month? | | |
| ▲ | walrus01 4 hours ago | parent | next [-] | | I think he means electrical bill at his estimated wattage load of the server and his known kWh cost, not rented server/hosting cost. | | | |
| ▲ | cobbzilla 4 hours ago | parent | prev [-] | | At my house. I have 5Gbps fiber and could pay for 10 or 25 if I need it. | | |
| ▲ | sebmellen 4 hours ago | parent [-] | | Gotcha. But to be clear, you’re talking only about energy usage, correct? | | |
| ▲ | cobbzilla 3 hours ago | parent [-] | | Yes, what other opex is there? It will have good ventilation, I’m not worried about cooling. |
|
|
|
|
|
|
| ▲ | embedding-shape 7 hours ago | parent | prev | next [-] |
| I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case. Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that. |
| |
| ▲ | walrus01 7 hours ago | parent [-] | | I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully. | | |
| ▲ | embedding-shape 6 hours ago | parent | next [-] | | Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer. And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch? I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :) | |
| ▲ | justincormack 3 hours ago | parent | prev [-] | | You can rent one in the cloud to try it |
|
|
|
| ▲ | barbacoa 2 hours ago | parent | prev | next [-] |
| They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs. So full CPU local AI inference may become viable option in coming years. |
| |
| ▲ | lallysingh 20 minutes ago | parent [-] | | This is essentially guaranteed. There are lots of useful smaller models that we should be able to run locally. Over time they'll be more and more capable and require less API usage. |
|
|
| ▲ | magicalhippo 6 hours ago | parent | prev | next [-] |
| > running it on a no GPU, but tons of RAM server Or from SSD using something like Colibri[1]. Not going to be quick, but at least runable. [1]: https://github.com/JustVugg/colibri |
| |
| ▲ | walrus01 6 hours ago | parent [-] | | It's a great concept but I think it would cross the line from 'very slow' to 'so slow it's unusable' at this size. Even if we say you have an NVME SSD that does 7GB/s reads, that's dramatically slower than being able to hold the whole thing in DRAM. Like the difference between 1.3 tok/s in RAM vs 0.1 tok/s with a colibri-like method. edit: the results I have seen from people trying colibri with fast consumer grade PCI-E 4.0 NVME SSD are 0.1 tok/s on models that are <700B in size, things that are well under 800GB on disk. With something that's 3T in size it'll probably be a lot slower than hat. | | |
| ▲ | zozbot234 4 hours ago | parent | next [-] | | For single stream inference of a MoE model, the size of active sparse parameters will matter a lot more than total parameters. This is generally around half of the reported active parameter count - the other half being a dense subset that can be easily cached in VRAM even on fairly modest consumer setups. So the achievable performance may be quite a bit better than a naïve assessment might suggest. | |
| ▲ | magicalhippo 6 hours ago | parent | prev | next [-] | | It claims to support using multiple devices RAID-0 style, which should boost performance, but yea probably not very useful for most. But still fun you can run it at home. | |
| ▲ | LtdJorge 4 hours ago | parent | prev [-] | | On a server machine you can have more than 100GB/s of NVMe if you parallelize (RAID 0 and the like). But it's still gonna be noticeably slower. | | |
| ▲ | walrus01 4 hours ago | parent [-] | | 1536GB of DDR4 ECC server RAM is somewhere between $4000-6000 USD used right now, by the time you put in parallel enough NVME SSD to approach good speeds, you'd be approaching that (and also likely running out of PCI-E bus lanes directly attached to the same motherboard to reasonably do so). |
|
|
|
|
| ▲ | johndough 5 hours ago | parent | prev | next [-] |
| > But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB. The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3 Also, their previous models were native INT4, so it would be weird if they went larger now. |
|
| ▲ | 5 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | Havoc 4 hours ago | parent | prev | next [-] |
| > Even if the output is like 5-6 tok/s On a 3T model I’d imagine you’d be closer to 0.05 tks |
| |
| ▲ | amluto 4 hours ago | parent [-] | | Presumably it’s MoE and only needs to read a small fraction of the weights per token. Bonus points if you can get decent speculative decoding without becoming ALU-limited. | | |
| ▲ | zozbot234 3 hours ago | parent [-] | | Speculative decoding is not really worthwhile for sparsely-loaded models. You end up paying in both memory bandwith and compute (loading experts based on wrongly-predicted tokens) which leaves you worse off overall. It becomes viable (even for sparse MoE) once you're batching so widely that you end up having to load most of your total weights anyway. | | |
| ▲ | amluto 2 hours ago | parent [-] | | > Speculative decoding is not really worthwhile for sparsely-loaded models. If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens, without too much quality loss. Another fun idea might be to try to build a model where the router chooses the expert 1-3 tokens in advance. |
|
|
|
|
| ▲ | sandworm101 7 hours ago | parent | prev [-] |
| Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day. |
| |
| ▲ | walrus01 7 hours ago | parent [-] | | The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn). But the memory bus speed is fully committed when generating tokens or thinking. | | |
| ▲ | sandworm101 3 hours ago | parent [-] | | Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins. | | |
|
|