Remix.run Logo
999900000999 4 hours ago

And hire 2 or 3 dev ops to keep it running ?

That another 400 to 700k.

It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.

wongarsu 4 hours ago | parent | next [-]

Where do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job

arjie 4 hours ago | parent | next [-]

Apparently it’s going to take the 3 of us to do this, mate. Going to get so much reading done.

LeonM 4 hours ago | parent | prev [-]

What you get is not what you cost.

40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.

margalabargala 3 hours ago | parent [-]

Yeah but you don't need two such people, or even one, dedicated to this single rack.

A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.

999900000999 3 hours ago | parent [-]

Who is going to upgrade the models ?

Who is going to fix it when the api does something weird ?

Who is going to proactively make sure it’s not overheating?

Chat GPT has enterprise contracts for a reason.

margalabargala 2 hours ago | parent | next [-]

The same person managing the company's email accounts and whatnot.

I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.

niltecedu 34 minutes ago | parent | prev | next [-]

My company is not even the same ballpark, but even we already have people for it. And doesn't include the fact that you can get a colo location and just a MSP or a contractor to do it for you

stymaar 2 hours ago | parent | prev | next [-]

All of the answers to your questions above are in gp's comment already:

> A company of the size that this is worthwhile for, probably has dedicated devops on staff already

3 hours ago | parent | prev [-]
[deleted]
shrubble 2 hours ago | parent | prev | next [-]

As mentioned, it is a "large enough company" already; they have full time sysadmins running things.

Adding another rack beside the VMWare cluster, managing any storage/networking issues, etc. will be incremental costs; they already have a pager (probably not a pager not anymore just an app on their phone) like rotation schedule etc.

russell_h 4 hours ago | parent | prev | next [-]

> However, I don’t trust hosted LLMs for anything that needs to be private.

Why not? Do you trust AWS with things that need to be private?

Kevcmk 4 hours ago | parent [-]

More than I trust frontier labs. AWS doesn't need to recoup 9 digits USD of capex

senderista 4 hours ago | parent [-]

So you can just use Bedrock?

w0m an hour ago | parent | prev | next [-]

> I don’t trust hosted LLMs for anything that needs to be private

I'd update this to

'I don’t LLMs for anything that needs to be private'

What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.

JacobAsmuth an hour ago | parent | next [-]

Why do you feel that creating the LLM from scratch is sufficient to trust it? Are you suggesting that you personally would read all 15 trillion tokens (plus every single agentic trade used in RL, along with its relative advantage in the batch) and personally guarantee that gradient descent would train a model which would not exfiltrate your corporate data?

Or that perhaps you have a perfect alignment algorithm which you are unwilling to share with the broader research community (evil)?

mdp2021 an hour ago | parent | prev [-]

> What's to prevent the LLM from

A NN per se is a file... The executable that runs it can "act"...

lumost 4 hours ago | parent | prev | next [-]

There will be cloud/SaaS vendors who have lower cost of labor/capital due to automation and financing terms.

Having these models in the open caps the inference margin.

GodelNumbering 4 hours ago | parent | prev | next [-]

> And hire 2 or 3 dev ops to keep it running

Not a devops but I'd say one full time is already too many.

dboreham 4 hours ago | parent [-]

Yes but zero is not enough and where do you get a fraction of a competent dev op from?

layer8 4 hours ago | parent [-]

From the other dev-op work you’re doing.

ayewo 3 hours ago | parent | next [-]

Understood but sharing your existing devops resources with this will soon become a bottleneck especially when any major downtime will keep several engineers (and long-running agents) blocked from any meaningful work until availability improves.

layer8 3 hours ago | parent [-]

Not my experience, from an SMB that maintains its own hardware and services. You have a certain contingent of competent engineers who distribute their work across projects, and it generally works out fine. Or course you plan with some redundancy and fall-back plans in your systems.

ayewo an hour ago | parent [-]

The top poster mentioned LLM spend of millions/month to justify the estimated capex of $6m to self-host Kimi on own infra.

Add to this number another $1.5m/yr in opex, so not sure I’d call such an enterprise wealthy enough to spend those kinds of sums on LLMs an “SMB”.

JacobAsmuth an hour ago | parent | prev [-]

Now you're thinking like management!

slicktux 4 hours ago | parent | prev | next [-]

Just like that new jobs created by AI! Localized model maintainer/technician.

toomuchtodo 4 hours ago | parent | prev | next [-]

You’ll slap some training on existing technologists/infra/sysadmin folks and perhaps have a support contract for the edge cases (hardware troubleshooting and advanced replacement).

(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)

JimmaDaRustla 43 minutes ago | parent | prev | next [-]

Why would a singe system require 3 full-time dev ops?

clint 4 hours ago | parent | prev [-]

Just let it manage itself, what could go wrong! :)