| ▲ | vablings 3 hours ago |
| That's pretty stupid. Most people who are incurring significant costs are just tokenmaxxing rather than being efficient with usage. You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment. I feel like people who are later to the AI game just like to "oneshot" and sink a bunch of usage into generating garbage |
|
| ▲ | Tsarbomb 3 hours ago | parent | next [-] |
| There really is a skill to using it effectively. I've tried coaching some of the devs on my team. Some get it, some don't. Our company has been tracking token usage and models used vs output (tickets, story points, PRs, deploys, etc...). A dev got chewed out, even after I warned him, because he spent over $2k in a single month almost exclusively on Opus while his actual productivity in terms of what he delivered was abysmal. |
| |
| ▲ | weinzierl 2 hours ago | parent | next [-] | | I get it but it goes against the grain for me. Isn't it ironic that we have to waste our precious and expensive human brain cycles to think about how to use AI cheaply so that it is not more expensive than us? In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do. | | |
| ▲ | remus an hour ago | parent | next [-] | | > In other words I want to spend 100% of my mental capacity in the problem domain for the things AI cannot do for me, like steering, grounding, verification and not for things AI could do. It sounds like the parent is less talking about this, and more people burning tokens while not getting useful work done. | |
| ▲ | Gigachad 31 minutes ago | parent | prev | next [-] | | This is much like how devs got grilled for creating expensive test VMs on AWS. Someone has to pay for all this at the end of the day. | |
| ▲ | ForHackernews 2 hours ago | parent | prev [-] | | I dunno, using tools and resources effectively is arguably the essence of good engineering. |
| |
| ▲ | n4r9 3 hours ago | parent | prev | next [-] | | What do sorry points mean anymore. | | |
| ▲ | senko 3 hours ago | parent | next [-] | | I love the typo. | | |
| ▲ | btown 2 hours ago | parent [-] | | Claude, vibe code me an entire startup, the actual product doesn't matter, but it should all be based on the incredible pun "turn 'sorry' points into story points." /goal get accepted into Y Combinator, you have an unlimited token budget, be bold. EDIT: no, do not just make a product that gives away your unlimited token budget to users for free! | | |
| |
| ▲ | SOLAR_FIELDS 3 hours ago | parent | prev | next [-] | | Did they ever have meaning? It's always been a nebulous feels term | | |
| ▲ | Terr_ 3 hours ago | parent [-] | | Attempting a serious but not-a-certified-whatever answer: "Points" do have meaning when properly used as a kind of moving-average tool for forecasting within a particular context. Problems arise when people try to perma-peg them to particular tasks, or (worse) man-hours or (much worse) man-hours across teams. Even just encouraging the humans to answer in terms of hours/days taints the accuracy of the forecast by introducing a kind of bias. | | |
| ▲ | fdsajfkldsfklds 2 hours ago | parent [-] | | For forecasting what, if not man-hours? | | |
| ▲ | t-writescode 2 hours ago | parent [-] | | Effort. Which is a very nebulous term, I agree. So, what you do is you recognize every ticket has a somewhat variable “actual effort”; and, if you’ve been honest in approximate effort pointing, you’ll know your team (or your own) velocity. From there you can run Monte Carlo simulations - say a few hundred thousand, and get a pretty good estimate of actual time spent. I’ve seen it work before with shocking accuracy. | | |
| ▲ | Terr_ an hour ago | parent [-] | | To play with the math analogies, imagine a black-box function: estimate(human_estimator, task_description, world_state) -> numeric_effort Assume that for various practical reasons, we've decided it's one of the best functions out there. How do we use it effectively, especially when it has noise, and drifts over time with unseen changes to the human_estimator and the hideously complex world_state? A popular option is to run it multiple times with different person/task combinations, putting a projected number on to each task. Afterwards, the tasks finished in sampling period ("sprint") become a quantifiable total for that period ("velocity"). Do the same process again with the next set of tasks, and you can figure out which ones are likely to fit if the velocity doesn't change much. If you know the velocity will change due to losing staff or vacation days... well, we apply a multiplier and hope for the best. Trying to "fix" the meaning of points is maladaptive, because they reflect many changing things which are outside our control and can't be independently measured. |
|
|
|
| |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
| |
| ▲ | onehair an hour ago | parent | prev | next [-] | | rookie numbers. in one of the top companies, i know someone who tokenmaxed so hard they ended up spending $50000 | |
| ▲ | proxyscore 2 hours ago | parent | prev | next [-] | | So you are going to blame this one that dev? Define productivity, and while at it, quality, maintainability , modularity and so forth. | |
| ▲ | PunchyHamster an hour ago | parent | prev | next [-] | | It's funny, the thing that makes effective prompt also makes effective documentation/communication. It's bizzare to see people that made clown issues (not enough detail etc.) suddenly start writing detailed prompts just because it is AI that will do the task and not the human on the other side. | |
| ▲ | cyanydeez 3 hours ago | parent | prev [-] | | there's a manifold to what "effective" means. The problem is once you get into the vibe flow, it's really difficult to eject yourself into the other realms of vscode or IDE or whatever it is you normal do because the vibing provides no anchor to what you're doing. Even if these models are smart enough to reorient themselves, they get entirely stuck in a desert and now you're asking someone to just pull up stakes and digg them out even thought they only watched them get there and the UI provides so much speed that no human can comprehend how they got there in the first place. It's like asking a pilot to take over in an emergency situation when they're not tasked with any of the every day requirements of the job. The orgs are relying on borrowed time of experienced professionals, and that's going to erode away and what replaces it is mostly people who understand how to navigate context but not use any of the _classic_ tools. It's a real conundrum and won't be easily surfaced but for a decade. | | |
| ▲ | mainmailman 2 hours ago | parent [-] | | I’m trying really hard to keep my skills up but it doesn’t feel productive when I’m using it to write code. It feels like I’m slowing down the AI to the point that it’s not as effective as just letting it go. But I don’t get all the learning that comes from that time along the way. Have you found ways to stay sharp while using it? Or are you relying on other projects outside of work to keep your skills fresh? |
|
|
|
| ▲ | cogman10 2 hours ago | parent | prev | next [-] |
| That's what happens when token usage becomes a performance metric. As has been done at my company. |
|
| ▲ | onehair an hour ago | parent | prev | next [-] |
| in my company there are a few who keep sharing screenshots of reaching limits on 3 separate subscriptions, 2 of them their personal on top of the company subscription |
| |
| ▲ | esseph 12 minutes ago | parent [-] | | Wonder when subscription-hopping attacks will become more often (jumping from a personal model to injecting instructions into the business account and exfiling data) |
|
|
| ▲ | hatthew 2 hours ago | parent | prev | next [-] |
| I feel like it's only within the past few months that opus got to the point where guiding the model is faster than doing things myself. I tried out sonnet recently and it was not a net positive to my work. I feel like anything that I'd trust haiku to handle isn't worth doing in the first place. For context, I'm doing a range of tasks, everything from one-shotting adhoc scripts to having 4 hour 10M+ token conversations debugging things. |
|
| ▲ | usaar333 3 hours ago | parent | prev | next [-] |
| > You can get 99% of jobs and work done with Haiku/Luna in a collaberating working enviroment. Optimally? Opus will pay for itself if you save just 10% of your time |
| |
| ▲ | qznc 30 minutes ago | parent | next [-] | | Only if all money is equal. Budgets in big enterprises work differently. | |
| ▲ | geodel 3 hours ago | parent | prev [-] | | True. I always Opus to pay for itself if it wants to get used by me. | | |
|
|
| ▲ | funnym0nk3y 2 hours ago | parent | prev | next [-] |
| Sorry, but that is nonsense. Compared to opus haiku doesn't cut it most of the time. |
| |
| ▲ | usef- an hour ago | parent [-] | | I think they mean the new Haiku, which is mildly above Luna now . If you have a plan written by a smarter model (so the hard parts are solved) they can be great at implementation. |
|
|
| ▲ | perching_aix 3 hours ago | parent | prev [-] |
| What on earth do you even do with these models? Or does a "collaborating work environment" mean that everything is basically spoonfed to them? Or do you only ever use ghost suggestions? I genuinely cannot even fathom. Just how do you even get into a state where tasks are so clear and cookie cutter? These things are abhorrent. Not only are they not useful, it's an outright form of psychological torture to try and use them. They almost fight you. Luna doesn't even respond to steers properly! You try steering it and it immediately gets distracted and then just stops. I can imagine coercing Sonnet into doing some of my tasks okay, but Haiku? Especially 4.5? Really? |
| |
| ▲ | dpkirchner 2 hours ago | parent [-] | | I think you might be overestimating the sort of projects most of us have worked on throughout our careers -- we haven't been doing much groundbreaking work. LLMs can easily and successfully write most code. | | |
| ▲ | steve_adams_86 29 minutes ago | parent | next [-] | | I equate most LLM work to squeezing a glue bottle It's just glue code It's not complicated. Someone just has to be there to squeeze the bottle | |
| ▲ | perching_aix 2 hours ago | parent | prev [-] | | It's possible it's my role distorting my perception, cause technically I don't write software, I work an SRE role. None of my items come pre-chewed or paced, it's all good luck and god bless. I'm desperately trying to classify and standardize my work items and delegate them to less capable models, because my usage is clearly unsustainable and this same sentiment as above keeps being pushed on me too. But all my tasks are genuinely fairly arbitrary, so there's no real way around the agent actually being able to reason about business and technical context proper. It's not even that they're hard, it's just that they're dynamic. I can get Luna to do things like walk our observability stack and perform a healthcheck, then defer to a stronger model if anything looks super off, but if I'm being entirely honest, this could basically be just a script. Which Opus 5.5 will immediately write for itself if it doesn't yet exist, run that, and then off it goes depending. But Luna will never actually do an investigation proper. Heck, it can't even read our dashboards most of the time, tripping up on Grafana minutia. It feels like that surgeon vs surgeon comparison, where you're made to decide based on their surgery success rate, and the better succeeding surgeon simply reward hacks the number by only operating on less dicey cases. Except there's no objective way to make this classification here, so jackasses like the above get to play with my insecurities with full obnoxious confidence, while I'm left desperately trying to slim my usage and failing to do so between two moments of crippling self doubt and blockers. | | |
| ▲ | AIblemblio 2 hours ago | parent [-] | | Our ai basic analysis for SRE / k8s based platform is haiku and its surprisngly good. I wouldn't even tried it, i would still just go with even Opus (we don't have that many alerts) but it really surpsied me. When i ran into usage limits a few days ago i switched most to Sonnet and again was surprised how good it is now. | | |
| ▲ | perching_aix an hour ago | parent [-] | | I wish we had such a platform (and was properly adopted). Maybe then the necessary context would be properly organized, and these lesser models could be effective here as well, especially if combined with harnessing integrations too. I still have a hard time accepting that Haiku/Luna tier models can be effective there even then, but I'll just have to take your word for it I suppose. |
|
|
|
|