Remix.run Logo
▲ bluegatty 2 hours ago

This is a failure of the AI foundries; if we have to use totally different prompting techniques for every model, this wont work.

AI is rapidly saturating it's ability to be useful and these products need to start to mature.

It's not 'fun' to manage 50 different broken MCPs and their variety of ways in which they are broken.

It was 'fun' at the start, now it's just 'broken technology'.

Astra and Opus 5.5 are the 'starting point' for the next era of AI where we expect robust tooling.

▲ACCount39 an hour ago | parent | next [-]

All LLMs understand natural language. All LLMs understand examples. That's honestly more compatibility than you get nearly anywhere, in anything.

The reason why advanced prompting is a moving target is that a lot of prompting is "use extra instructions to compensate for specific ways in which the target LLM is weak or prone to errors". And guess what? LLMs get better over time - obsoleting your advanced prompting.

"Tune a prompt to death for the specific task and the specific model" gets you better performance in the moment, but "trust LLM to be smart" ages a lot more gracefully.

▲bluegatty an hour ago | parent [-]

"but "trust LLM to be smart" ages a lot more gracefully."

That it doesn't even work now.

The word 'smart' there is actually doing a lot of heavy lifting, it's entirely contextualized.

▲TomGarden 2 hours ago | parent | prev | next [-]

Agreed.

I'm genuinely worried about all our short term investment in mitigating the failure modes of models that may only be SOTA for a few months.

It's very possible people being 'late' adopting AI may end up with a leg up, not only because they spent more time polishing personal skills during this time, but also because they don't bring all the baggage of 'AI competence' that is becoming irrelevant at breakneck speed.

▲skybrian an hour ago | parent | next [-]

Suppose you use LLMs in a more straightforward way, like asking coding agents to make specific changes rather than attempting to build a software factory?

That could be seen either as early adoption that’s overfitted to current capabilities or as late adoption of LLM’s more advanced capabilities.

▲1980phipsi an hour ago | parent | prev | next [-]

The people who are adopting LLMs later are also slower to adopt new technology in general. The samples of people who adopt early and adopt late have different characteristics.

▲pbronez an hour ago | parent | prev | next [-]

This is “second mover advantage.” There are several dynamics that can make it better to wait and move later. Framed in terms of firms, moving late is advantaged when:

- the product category is long lived

- switching costs are low for buyers

- there are objective standards of quality

- product imitation costs are low

https://insight.kellogg.northwestern.edu/article/the_second_...

Let’s consider those criteria for an individual competing in the labor market with AI. The category should be long lived, AI is here to stay. Switching costs (here, hiring/firing by employers/clients) are low. Objective quality standards fails; technical labor is notoriously difficult to quantify. Imitation costs (can you copy someone else’s good ideas) are moderate but decreasing. That’s where model and tooling improvement shows up.

Based on this analysis, I agree that late movers are well positioned IF the market leaders continue to improve models and tooling to integrate best practices that were previously individual skills.

Early movers should exploit the lack of objective standards. Use your experience with the first generation of tools as marketing to win and retain clients. Continue to invest in soft skills like communication.

▲maipen an hour ago | parent | prev [-]

What personal skills are you referring to?

▲saretup 2 hours ago | parent | prev | next [-]

> It was 'fun' at the start, now it's just 'broken technology'.

It was even more 'broken' at the start. We overcame some of the issues by 'prompt engineering', which is needed less in the newer, smarter models.

▲bluegatty an hour ago | parent [-]

Of course - what I mean to say is that we did not perceive it as broken.

The first combustion engine was a miracle. It only becomes 'broken' when we evaluate in some kind of applicable context.

▲lunchbucket 20 minutes ago | parent | prev | next [-]

You'll find similar documentation anytime a language or framework or other systems software ships a new major version. It doesn't seem like the way to prompt Opus has changed all that much. Certainly not enough to require a "totally different prompting technique."

▲rinconrex an hour ago | parent | prev | next [-]

Counterpoint, the differentiation is maturity. If all models are simply interchangeable commodities, what's the payoff for Anthropic or OpenAI?

Vastly different ways of interacting with each provider is another story, but really we are pretty spoiled here. Slightly different prompting techniques is not really a big deal. If anything it shows the user has some nuance and appreciation for what each model provides.

Fow what it's worth, I am super happy with Opus 5.5. Less verbose than 5 and just gets work done. The progress has been astounding, and if I have to coax it out a bit differently on Opus 5.5 vs Astra 6, I am happy to pay that small price.

▲cindyllm 2 hours ago | parent | prev [-]

[dead]