| |
| ▲ | fweimer 2 hours ago | parent | next [-] | | Centralized inference can easily increase batch size, leading to huge efficiency gains in the usual scenario where most users have just one or very few session. Using local resources efficiently requires some way to increase the batch size. I'm not sure if we are there yet. | | |
| ▲ | janalsncm an hour ago | parent [-] | | I think the point is that if people are able to run inference on their laptops batch size efficiency won’t matter. And before that, businesses will be able to get decent results with dedicated inference hardware. |
| |
| ▲ | ipdashc 3 hours ago | parent | prev | next [-] | | I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well? Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries. Plus, you know, the whole economies of scale thing. Local LLM has a lot of privacy and independence benefits, but I'm not really seeing the world where it becomes more energy- or cost-efficient to buy your own hardware (and use it 1% of the time) versus sharing a giant machine, or even the same machine, in a datacenter (where it has a much higher utilization factor). | | |
| ▲ | mattmaroon 3 hours ago | parent | next [-] | | All white collar work will be done by AI soon, there won’t be scaling down, just scaling up. I do food trucks and festivals and I’ve got AI doing so much of my non-meatspace work now that I’m buying $10-$20 a day in tokens. At that rate, hardware starts to look cheap, and I have less of a use case for it than most white collar workers. | | |
| ▲ | ipdashc 3 hours ago | parent | next [-] | | Genuine question, what are you spending that on? The $20/month ChatGPT/Codex subscription has largely been enough for me as an IT worker. | | |
| ▲ | mattmaroon 2 hours ago | parent [-] | | There. I have the business plan with two seats and I use them both and blow through it pretty fast. I think it’s because much of what I have it do involves using a browser. For instance I have it pull various permits from cities and there’s no API for that. |
| |
| ▲ | redrove 2 hours ago | parent | prev [-] | | I would also be interested in what you use AI for. Is it marketing? bureaucracy? | | |
| ▲ | mattmaroon 2 hours ago | parent [-] | | Yes. Bureaucracy for sure when I throw festivals. ChatGPT pulls permits for me. (I think that uses a lot of tokens because of the browser control?) Manages the admin side along with some tools I built in Lovable via mcp servers. Marketing definitely. My food truck side is relatively high volume and it manages my kitchen and warehouse side, basically generating all of the instructions my employees follow, managing and updating my PoSes, creating signage assets for specials, etc. Reels/posts production and managing ad spend. Here’s a fun one. I switched payroll providers after several years and suddenly my unemployment insurance rate went from 0.8% to 12.75% which is borderline debilitating to me. I knew something was wrong but not what and I work a lot of hours and calling the state takes forever and is usually unhelpful. ChatGPT figured out that it was a penalty rate and dug in for me. Turns out because I’m seasonal and have no payroll for one quarter of every year, Gusto did not file a quarterly wage report, so even though I owed nothing I was delinquent. Gotta love government, it’s the only place where you can be delinquent for $0. ChatGPT filed the report and requested a retroactive re-rate, which they granted. I’m sure I would have figured this out eventually but it would have taken hours, or $5 in tokens. | | |
| ▲ | ipdashc 4 minutes ago | parent | next [-] | | Makes sense. Thanks for the post! | |
| ▲ | redrove 2 hours ago | parent | prev [-] | | Thanks for the very detailed answer! Would you say ChatGPT does everything you need today or do you still see some gaps? i.e. things you think it should be able to do but currently doesn’t, or things that still take too much effort on your part to setup chatgpt to do it. | | |
| ▲ | mattmaroon 2 hours ago | parent | next [-] | | Also, I have certainly gone down some unproductive rabbit holes figuring out what to do with it. Time spent on it is an investment, just like automating anything. You spend hours upfront to save them on an ongoing basis. But, I think I’m on the black on it already after just a few months of heavy use. For instance I just tell it to book my dumpsters, bathrooms, sanitation crew, and security for X event. It goes and pulls event details, looks through my email to see who I get those from, and emails them relevant details, with no prompting. Another good example: we launched a really fancy hot cocoa concept last fall that was a hit and I wanted to try to go to all of the local pumpkin patches in October and Christmas tree farms after Thanksgiving to serve when they have big crowds. I asked it to contact all of the ones in my area and it sent out 70 emails and I booked several spots. It made me a nice database so I can see who followed up and who I need to reach out to again, etc. and it just gets those from my inbox. Hours of my time saved with simple prompts. So while some things take awhile to pay off, some are instant. | |
| ▲ | mattmaroon 2 hours ago | parent | prev [-] | | Huge gaps. I expect the tooling to get better for non-programmers. Codex and Cowork are great, but you still feel like you’re trying to hammer the square peg through the circle hole often when using it for non-programming tasks. The AI is good enough to do a lot of tasks but the tooling just isn’t caught up to it yet. I’d say it’s freed up ten hours a week of my time. And that’ll only improve. | | |
| ▲ | rene-veerman an hour ago | parent [-] | | how'd you get it to run?
not even upgrading ollama could get it installed on my end :( |
|
|
|
|
| |
| ▲ | SahAssar 2 hours ago | parent | prev | next [-] | | I'm pretty sure a small-to-medium org has plenty of things that could be queued/scheduled to run when there is downtime. It requires some planning and thought though and I don't think most orgs are there yet. | | |
| ▲ | LinXitoW 19 minutes ago | parent [-] | | Yeah, but again, why would you? When it comes to open weight models, there's a good amount of competition, so you already get a really good price, without any optimization or anything. Seriously, you can do years of Deepseek inference for the hardware to run just 1 or 2 requests against a slower, dumbed down model on your own hardware. It makes no sense to buy hardware right now, when the price is completely disconnected from any material reality. It's much better to use the cloud providers VC funding by using their cheap as F offering. Either AI becomes less useful, or hardware costs come down. Either way, you'll be in a better position in 3 years than you are today. |
| |
| ▲ | toyg an hour ago | parent | prev [-] | | $2k? LOL. That's not even the GPU budget, these days. The economies are currently out of whack because of underproduction of components and memory, so LLM providers have a few years of runway to entrench. Plus, the whole capex Vs opex thing that helped AWS will help here too, for sure. At some point, though, things will change. More production will come online, and providers will have to end the current speculative subsidizing and jack up prices. It's a bit like the dot-com era: the initial rush to land-grab web portals and e-commerce sites eventually died, once enough skills and infrastructure came online, and the bubble burst. |
| |
| ▲ | NhanH 4 hours ago | parent | prev | next [-] | | The whole cloud story lies on two aspects: - Hyperscaling “we are going to serve billions of people in our applications”, which is becoming increasing unlikely as regional tech companies become more dominant than than the global one (this one is as much about geopolitics as technology) - Operations is hard, in which case non-frontier models should be increasingly capable. Devops for small-ish deployment is one of the few cases where it is hard to clam you need deep expertise and AI can’t do it. Previously, the claim is that you need people specialized in ops, which is expensive. Now… My prediction is that not just cloud LLM, but cloud business general will have to change. Not yet in the next 5 years, but probably 8-20 years-ish | | |
| ▲ | redrove 2 hours ago | parent [-] | | I don’t disagree but saying “cloud will change in the next 8-20y-ish” is a bit of a non-argument, you’re not really stating any thesis to speak of; Change is a given over that time frame. | | |
| ▲ | NhanH 2 hours ago | parent [-] | | Ah it was not explicit enough. They will die, for some definitions of death -- I don't think they disappear, but they should be a niche, rather than the dominant doctrine. | | |
| ▲ | redrove 2 hours ago | parent [-] | | I’m not so sure I agree, given the overwhelming concentration of capital and regulatory capture they have, I just don’t see them going anywhere; becoming more niche rather than even more of a standard is “going away” to a certain extent as far as I can see. |
|
|
| |
| ▲ | manmal an hour ago | parent | prev | next [-] | | Yeah. Local agent sessions are not backed up in the cloud. And uptime is better with local models. | |
| ▲ | SV_BubbleTime 5 hours ago | parent | prev [-] | | [dead] |
|