Remix.run Logo
mindwok 7 hours ago

Gotta be honest, almost every "how to use AI" resource seems pointless to me. I'm either going to ask the AI how to do it, or if it's about using the AI then we can just bake it into the harness or wait for Anthropic/OpenAI to do it for me because they're always trivial.

All of these resources on agentic workflows, managing agent memory, harness engineering, etc. appear to just be theatre to me.

cloakandswagger 6 hours ago | parent | next [-]

Remember in 2023 when people thought "prompt engineering" would be the new software engineering and invested tons of time into learning CoT, ReAct, thread-of-thoughts, etc?

Those were mostly obviated by reasoning models and harness updates by 2024.

It seems pointless to invest energy into the latest/greatest AI technique or framework when they're going to either be absorbed or replaced on a 3 month cycle.

embedding-shape 5 hours ago | parent | next [-]

Isn't it clear that some people are better at working with/prompting LLMs than other people? Or is the idea that what you write to them and how you use them doesn't matter, it's all up to the model/harness? To me this seems clear, so then clearly this is a skill, which typically is called "prompt engineering". Specifically CoT or the other things you mention wasn't referred to as "prompt engineering" as far as I know, that skill is more about how you communicate with the LLMs and how you use them, rather than what specific processes/workflows/technologies you use.

NichoPaolucci 4 hours ago | parent | next [-]

I actually think that good prompting MOSTLY comes from good writing skills in general. Being able to more clearly state things to an agent, knowing what pieces of context are entirely unnecessary and which are important, having a larger vocabulary helps too.

Of course, there are other areas that can improve model output (Direction rather than open-ended assistance requests, using keywords + plugins that help, the "your output should include: " style prompting).

A few of us run almost the same exact setup at my shop (Base Claude Code w/ SuperPowers + a context repository) and the models are somewhat unhelpful to some, and give meaningful output to others. The only correlation I notice is that their prompts are no-good. Not from a meta "prompt" engineering standpoint, but from a general English 101 standpoint.

"dudde no i wanted the function to return 3 things. not like that. do it again"

VS something like

"Modify the "renderThreeVars()" function signature to accept another variable called "z" and add it to the return statement at line 64."

Obvious exaggeration, but you get the point.

Centigonal 26 minutes ago | parent | next [-]

Is Superpowers any good? My coworkers who've used it seem to think that its main purpose is to consume a lot of tokens.

user43928 3 hours ago | parent | prev | next [-]

Why not open-ended assistance requests?

I ask it all the time about whether X is feasible, how we can get started on Y, and to investigate issue Z.

It is working great for me in a >100k LOC project.

Perhaps this works less well with weaker models. I suspect the people who say Qwen 3.6 27B is working well, are using prompts like "modify the renderThreeVars() function in rendering.py".

2 hours ago | parent [-]
[deleted]
cousinbryce an hour ago | parent | prev | next [-]

Well said. I’ve been trying to put my finger on this for a while. The interaction plane for an LLM is __all of human language__

vablings 2 hours ago | parent | prev [-]

Speaking a prompt with a long ramble is very powerful.

Centigonal 25 minutes ago | parent [-]

I will often do "speak a long rambling set of ideas and ask the LLM to summarize it into a prompt or spec -> manually refine -> drop refined prompt into fresh conversation"

cloakandswagger 4 hours ago | parent | prev | next [-]

As part of a previous job I needed to audit internal AI usage from a largely non-technical employee population.

The prompts were, predictably, really bad. Broken English, sentence fragments, vague requests, lack of context. Yet somehow, the users always got the answer they were looking for. It might have taken a few extra turns with questions from the model, but the end result was the same.

It's humbling, but a flowery, carefully crafted prompt is at best slightly more efficient than a "CAN A DOG BE EATIN SUN FLOWER SEED?" peasant prompt.

NichoPaolucci 4 hours ago | parent | next [-]

You call that a peasant prompt, but it's actually almost perfect. Couple notes, but it's 95% of the way there. "Can a dog eat sunflower seed" is probably the perfect version, just 1 extraneous word in this version.

Unless the user wanted to know if a cat could eat sunflower seed or something.

cloakandswagger 2 hours ago | parent | next [-]

To be pedantic, these types of prompts work best when in first-person/roleplaying. So the "perfect" prompt here would be something like, "I'm a dog and I just ate 20 grams of salted sunflower seeds with the shell on. Because I'm a dog I sometimes eat things without thinking about it. I'm worried about the short and long-term physiological consequences of what I've just done..."

Xirdus 2 hours ago | parent [-]

I copied your prompt verbatim to ChatGPT and to Google.

ChatGPT kept the charade for all of one sentence. Then it dropped to talking about "your dog" the rest of the way. It even starts the final paragraph with "if, instead, you mean you (a human) ate them...", and finishes with the question "is this about an actual dog or yourself?"

Google did consistently refer to me as a dog, but its entire focus was on the steps "your human" should take, no advice for the dog itself.

In both cases, it looks like the first person roleplay was entirely inconsequential for the usefulness of the output. I think the current gen AIs have outgrown this trick and you can safely forget about it.

cloakandswagger an hour ago | parent [-]

You have to steer it back with "No I'm a dog"

selimthegrim 3 hours ago | parent | prev [-]

I am tempted to try get me a lawyer dog

watwut an hour ago | parent | prev [-]

> Broken English, sentence fragments,

Why is would that be "bad prompt"? It is machine inpit, if machine can interpret fragment all the better.

cyclopeanutopia 4 hours ago | parent | prev | next [-]

Nah, you might be confusing prompt engineering with having domain knowledge. :)

embedding-shape 3 hours ago | parent [-]

I think some people who are better at "prompting" even without domain knowledge could be better at getting LLM agents to produce good results than people with good domain knowledge but without the skills to prompt well. Just a hypothesis though, would be fun to try it out for real sometime :)

ryandvm 3 hours ago | parent | prev | next [-]

I feel the same way about "prompt engineering" as I feel regarding the term "parkour" - you know, running and jumping on stuff.

Are people really putting on their resumes that they are capable of reading and writing and appropriately defining and limiting context? That's all prompt engineering is - it's being able to communicate effectively and elucidate your objectives.

Congratulations to all you English majors out there, you're about to make $350K/year.

embedding-shape 29 minutes ago | parent | next [-]

Personally I don't, but why not? People aren't embarrassed to put their language skills ("be able to communicate in this specific language" - not special, it's just another language?), their leadership skills ("effective business communication" - big deal) or that they are a people-person ("can talk with others" - most people can do this) on their resume.

Xirdus 2 hours ago | parent | prev | next [-]

> Are people really putting on their resumes that they are capable of reading and writing and appropriately defining and limiting context?

People put whatever buzzwords will get them through initial screening. My resume contains tons of banal shit like agile, automated testing, Linux, AI (since 2018), and design patterns.

cwmoore 3 hours ago | parent | prev [-]

Parkour is for the more energetic peripatetic.

omega3 5 hours ago | parent | prev | next [-]

What's clear is that there is a lot of hype around LLM and people who were previously valued for their IC are now in the business of shilling.

CuriouslyC 5 hours ago | parent | prev | next [-]

RL has basically killed prompt engineering. You still need to provide the right context and process, but how you communicate with them beyond that is no longer so important.

sjh9714 4 hours ago | parent [-]

[flagged]

troupo 4 hours ago | parent | prev | next [-]

> Or is the idea that what you write to them and how you use them doesn't matter, it's all up to the model/harness?

Yes, it is. Source: had models inplement complex things from scratch and bullshit regardless of whether it was a one-line prompt or a detailed "SOTA witchcraft magic spells that are guaranteed to work"

saberience 5 hours ago | parent | prev [-]

It's not really a skill. The models are at this point smarter than you are, so the idea that you can prompt them "better" is laughable really when discussing frontier models.

It's like imagining you could "prompt" Richard Feynman to be smarter at Physics.

That is, for 99.9% of engineers, if you want the model to do a code review of your project, the best solution is to just ask Fable, "Hey Fable, do a code review of this project." Throwing in extra text like "think like a senior engineer", "ensure you focus on DRY principles, KISS, self documenting code, etc", doesn't make a difference.

These sorts of tricks used to work with dumber models, but now, like I said before, it's like thinking you can prompt Linus Torvalds into writing better C than he already can do.

defrost 5 hours ago | parent | next [-]

FWiW

> it's like thinking you can prompt Linus Torvalds into writing better C++ than he already can do.

  Linus Torvalds, the inventor (and beloved dictator) of Linux, has always been quite harsh about C++ and why he rejects it for Linux kernel development. He’s not just been very vocal about it, but also brought up some arguments against the use of C++ that are worth reviewing in detail.

~ https://medium.com/@jankammerath/linus-torvalds-critique-of-...
saberience 3 hours ago | parent [-]

I mean, the point stands.

The models are beyond expert level in many areas at this point.

Do you really believe that adding extra junk to your prompt is going to make the model write code better than it does already?

Again, imagine going to Terrence Tao and "prompting" him to get better at Maths, do you think you can do it? What prompt would you give to him to make him produce better maths. Unless you're already a world-leading Mathematician I think you would find it hard.

cityofdelusion an hour ago | parent | next [-]

Models are not rational thinking brains, so the point is moot.

Extra prompt text isn’t to make the model smarter, it’s to pull attention towards what you want. “Do code review” is far different from “Make sure changes align with existing architecture” or “follow these enterprise standards XYZ”. The model just acts as an average of its training data, which may or may not align with your goals.

alwillis 2 hours ago | parent | prev [-]

> I mean, the point stands. The models are beyond expert level in many areas at this point.

It's not that the models aren't smart or whatever; we know they're extremely capable.

There's always going to be value in being able to clearly communicate what you want the model to do, especially when the context is lacking.

Yiin 3 hours ago | parent | prev | next [-]

I find this very much untrue in my practice and my testing, where fable prompted to do a review found surface level issues, while prompting along the lines of "assume it's wrong, prove it's correct" found much more in depth and real issues.

gavmor 2 hours ago | parent | prev | next [-]

"DRY" and "KISS" are, possibly, the least interesting principles to which a software designer might adhere.

4 hours ago | parent | prev | next [-]
[deleted]
sjh9714 4 hours ago | parent | prev [-]

[flagged]

bingemaker 4 hours ago | parent | prev | next [-]

I still believe in writing good prompts or good instructions. Bad prompts can sometimes blow up the bill. A poorly written spec can waste a lot of tokens

thunky 4 hours ago | parent | next [-]

Sure but lets not pretend this is engineering.

lukan 4 hours ago | parent | prev [-]

Definitely. But the special knowledge how to talk to a certain model is usually not worth it.

Being able to write clear, always is.

cubefox 39 minutes ago | parent | prev [-]

> Remember in 2023 when people thought "prompt engineering" would be the new software engineering and invested tons of time into learning CoT, ReAct, thread-of-thoughts, etc?

Prompt engineering was a GPT-3 era term, which couldn't understand instructions. Then ChatGPT came out in late 2022, which made actual prompt engineering superfluous.

deepfriedbits 7 hours ago | parent | prev | next [-]

Not only that, but all of this tooling around models has such a short shelf life as the models themselves grow in capabilities, they absorb the tooling. We've already seen it over and over again.

mexicocitinluez 3 hours ago | parent [-]

Amen on both your and OP's comments.

These tools work pretty well out-of-the-box. I'm sure I could squeeze out better token usage or streamline some tool calls, but it's not something I really want to focus on. Just like I don't want to endlessly configure my IDE, I don't really have patience with spending time on anything besides actually building something.

tedggh 4 hours ago | parent | prev | next [-]

There was an article by Vercel on how ineffective tool calling is compared to just a single md file with clearly defined instructions. They showed a clever way on how to use compressed indexes. My experience with tool calling was similar to what Vercel described, and I spent hours trying to perfect it. Since reading Vercel’s findings, Claude.md and Agents.md, maybe a Project.md it’s all I use.

39 minutes ago | parent | prev | next [-]
[deleted]
knollimar 6 hours ago | parent | prev | next [-]

>I'm either going to ask the AI how to do it

LLMs seem terrible at using LLMs in harnesses. Have you seen how they rot their context with the stuff they put in .md files if you let them?

You'd have to have the LLMs search, and thus these resources could be for them more than you

esperent 5 hours ago | parent [-]

Yes, but I have had moderate luck with creating an "agent-instructor" skill that has strict instructions around keeping language strong, unambiguous, concise, and always presenting me with exact diffs to review before writing anything.

Another thing in it is a strict line count. Any increase in line count requires my approval. That last one is important because it plays well with two biases: models don't tend to create long lines so they won't try to cheat that way, and they're strongly inclined to keep churning out lines so I take that away from them.

fwip 4 hours ago | parent [-]

Did you tell Claude "make an agent instructor skill, whatever you think is good, go for it," or did you use the knowledge you had gained about how AI works and how to write good instructions for it?

LogicFailsMe an hour ago | parent | prev | next [-]

Once you understand enough to spawn independent agents for discrete tasks, it's hard to justify investing brain space into some harness that will likely be obsolete next week, if not tomorrow. And by the time the harness game converges, I suspect most of the high-end models will behave like them! Intrinsically.

rowanseymour 2 hours ago | parent | prev | next [-]

Right now I feel like I'm getting so much value out of working daily with Fable, but I'd be embarrassed to share my sessions because they're pretty much how I would talk to a colleague. Spelling mistakes and half-baked thoughts included. Gone are the formal sounding prompts and specs I was proudly writing 6 months ago. But it's getting the results.

bad_username 4 hours ago | parent | prev | next [-]

> I'm either going to ask the AI how to do it

First you have to know that "it" exists and is possible. A cookbook like this introduces readers to concepts and features they didn't know were there in the first place.

oggreen 6 hours ago | parent | prev | next [-]

Thanks for saying the quiet part out loud. Everytime I do a demo it seems like a waste of time compared to just actually building.

Unless I'm doing something super complicated even taking time to set things up like subagents, etc. seems like a waste compared to just building.

The only things that really seem beneficial (for Claude Code) seems to be learning to set up loops, memory and finding relevant MCP servers.

edot 4 hours ago | parent | prev | next [-]

Exactly. I think we all have to come to this conclusion ourselves, because the message we’ve been getting from the LLM companies is “sweet user, you DO add value to the LLM, you gave it that custom subagent, remember?”, and then they go and roll that idea into the next version. Once they do that a few times you go “eh why would I bother, this will be a default soon. I’ll just prompt it like normal”.

Just use the vanilla settings.

dec0dedab0de 3 hours ago | parent | prev | next [-]

I suspect certain workflows like langchain and others like it will retain usefulness into the future. Having deterministic steps before and after the llm is the way to go for anything that might be potentially harmful. Which I guess is what the harness does, but why be limited to a generic harness when we can use it to make specific ones for our needs.

CuriouslyC 5 hours ago | parent | prev | next [-]

It was more important with earlier models, as they were less RL'd to golden paths, so the context could steer them more.

Now process engineering is more important than prompt engineering.

john___matrix 7 hours ago | parent | prev | next [-]

This is how I've gone about it with my Shopify site.

I appreciate there may be more efficient ways to do things if you're an actual developer or engineer but I'm not and the tools I've been able to build so far have been both fun to make and add value to our workflow as a small 2 person ecommerce business.

I do at least have a background in web as a designer for many years and having worked with developers I can at least spec and understand how a product might work which is one thing Claude/AI isn't hugely helpful with and often it's the testing and QA phase as always where the problems and shortcomings expose themselves.

coldtrait 2 hours ago | parent | prev | next [-]

Yea i do the same. I can't be bothered to watch videos and courses on how to use claude better when I can simply ask it myself.

tclancy 5 hours ago | parent | prev | next [-]

Ok, thanks, because I looked at one example and was like, “I am supposed to take someone’s OpenAPI doc and translate it back to English for the model?” And if they’re implying I should have an AI do that, why don’t they just build that step into the models rather than having someone prompt an AI to write these half-ass docs?

NooneAtAll3 6 hours ago | parent | prev | next [-]

if you know how to ask Ai about how to use Ai, you already know how to use Ai

the point of guides is to provide assurance for people unfamiliar to the process in the first place

mindwok 6 hours ago | parent [-]

If the bar is knowing how to type a question into a box, I'm confident almost everyone is better off starting with that then reading a "cookbook" that starts with installing python packages.

hiccuphippo 3 hours ago | parent [-]

I've see people write "broooo pleaseeee!" into the box. Something about the universe providing better fools.

jagenabler2 3 hours ago | parent [-]

I do this often, I manage to get the results I’m looking for

throwatdem12311 6 hours ago | parent | prev | next [-]

If these things are so smart why do I have to coax it to be useful with all these magic spells scrawled on markdown parchment.

giancarlostoro 3 hours ago | parent | prev | next [-]

Not only that, but they change on a whim with new ideas on how to do things every few weeks.

OtherShrezzing 5 hours ago | parent | prev | next [-]

I think the main thing of interest in the linked site is the dates. You can quickly get a view of what was possible and when.

Xalutiono 5 hours ago | parent | prev | next [-]

You don't learn about progress if you don't take part in it.

Before /goal was ralph-wiggum it was def an interesting learning experience, it was interesting to see how Claude became a lot better doing this itself but it still took 6 Month and more.

You can wait and sit it out and suddenly you get fired if you miss the point when to start spending more time and energy on topics like this.

jie-yang 6 hours ago | parent | prev [-]

[flagged]