| |
| ▲ | jacobgold 2 days ago | parent [-] | | Drawing the conclusion that "humans fail" and "models fail", so they must be similar, is very wrong. You could have humans calculate 2+2 all day and get a surprisingly high error rate. That reveals a flaw in how humans operate. LLMs fail for entirely different reasons. Their mistakes don't imply they're human-like at all. It's not about the error rate. | | |
| ▲ | ACCount37 2 days ago | parent [-] | | You're saying that a class of mistakes points out a "major (possibly fundamental) flaw". I'm pointing out some very similar classes of mistakes in humans - well known, well documented and widely exploited. They just keep paying the "IRS" in gift cards, buying lottery tickets and getting the captain's age wrong. If you're using the existence of flaws in LLMs to deny the claim of intelligence to them, then why do "generally intelligent" humans exhibit some impressively similar-looking flaws? And, if we're talking about that conspicuous similarity - do they actually fail "for entirely different reasons"? Or do you just want the reasons to be "entirely different" - and not the same reasons viewed at a different angle? Because the similarities between humans falling for trick questions or scams, and LLMs falling for adversarial questions or prompt injections don't look coincidental to me at all. One of the oldest patterns in scamming is overwhelming and confusing the victim. Numerous prompt injection methods seek to overwhelm and confuse an LLM - if an LLM can't keep track of things, can't grasp what's going on, it's far more likely to lose track of what's a prompt and what's data, overlook past instructions or go past its behavioral guardrails. And humans who fall for trick questions like "1kg of feathers" or "captain's age" due to shallow attention and naive pattern matching? They fail in surprisingly similar ways to how LLMs fail on SimpleBench tasks that are filled with overwhelming adversarial distractors. Many "trick questions" are tricky to humans and LLMs alike - to the point that it's unlikely to be coincidental. | | |
| ▲ | jacobgold 2 days ago | parent [-] | | > If you're using the existence of flaws in LLMs to deny the claim of intelligence to them... That's not the point at all. It's the fact that they fail in ways completely unlike humans. You also have the burden of proof reversed. Its on you to prove these LLM agents are human-like intelligences if that's your claim. No one can prove this because it's false. | | |
| ▲ | ACCount37 2 days ago | parent [-] | | You are the one claiming that "they fail in ways completely unlike humans" insistently. Now go cough up some proof. I'll wait. | | |
| ▲ | Jensson 2 days ago | parent | next [-] | | If they didn't you wouldn't need the operator, you'd have replaced all your programmers with no drawbacks by now. As long as we keep hiring humans that is all the evidence you need that these AI fails in ways humans don't. | | |
| ▲ | gghackernewsgg 2 days ago | parent [-] | | Junior developers fail in different ways than senior developers too; that's why seniors oversee juniors. But this doesn't necessarily mean that the senior's and junior's intelligences differ in kind Your argument "AI needs supervision, therefore it fails in different ways than its operator does" holds. Your argument "AI fails in different ways than its operator, therefore the AI's intelligence is different in kind" doesn't hold. |
| |
| ▲ | jacobgold 2 days ago | parent | prev [-] | | This isn't even controversial. The proof is available to anyone who uses these systems: They hallucinate tool state, drift from the objective while seeming to comply, switch languages randomly (Cyrillic or Japanese characters in output), confuse tasks they've planned for completed ones, and of course follow prompt injections embedded in files or web pages. | | |
| ▲ | vidarh 2 days ago | parent | next [-] | | I switch languages "randomly" all the time when I think about something in another one of the languages I know. Some word will trigger it and before I know it I will continue in the other language. In fact just the other day I commented on it to my fiancee after I randomly switched to French because we were discussing a trip and I mentioned a French location and pronounced it in French, and suddenly I was in "French mode" entirely unintentionally and it took a sentence before I realised. That you think this is unique to LLM's suggests you simply don't know the diversity of human thought as well as perhaps you think you do. That's fine - none of us have a very complete view of that. | | |
| ▲ | jacobgold a day ago | parent [-] | | Yes, people speak multiple languages and switch between them. But it's a superficial analogy to the behavior of LLMs which do something different and for different reasons. | | |
| ▲ | vidarh a day ago | parent [-] | | You're misrepresenting what I wrote. I specifically pointed out that I switch languages without intent to do so. When you suggest that is a "superficial analogy" after you were the one pointing out LLMs switching language as something that sets them apart, you're seriously reaching. I can often pinpoint afterward what was likely the trigger: E.g. I used a word that is the same in two languages, and continue in the second; I pronounced a word in its native language for whatever reason, and continued in that language; my "context" suddenly included another language because someone else spoke the other languages within earshot of me. What makes you think this is materially different from an LLM switching language because its probability distribution gives a word in a different language because it fits in context? In the examples I gave, each even made a word in the language I switched to more probable as a reasonable continuation, just as with an LLM. I'm not claiming the mechanisms are identical, or even similar, but the behaviour most certainly is more similar than "a superficial analogy" would imply. | | |
| ▲ | jacobgold a day ago | parent [-] | | That's a fair point. I agree that the behavior can look similar even when the mechanism is different. But in practice, all the analogies I've seen are in fact superficial, including this one. The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor. | | |
| ▲ | Kim_Bruning a day ago | parent | next [-] | | It's called "code-switching" or "code-mixing", and bilinguals do it all the time. When an immigrant kid does it, you don't call a doctor, you call it adorable. By the way, it's not switching topic. You just pick the concept closest to what you mean from your combined vocabulary. If you're not paying close attention, you might switch language though (until the next concept you need is from the other language again, at which point you switch back) And you're aware the paper "Attention is all you need" came out of machine translation research at Google, right? You hold an internal semantic representation and map in and out from arbitrary natural languages. I think the (bi-, tri-, multi-)lingual approach is the only proper way to translate, and this is a hill I will fight on! Google may have gotten more than they bargained for on that particular translation experiment; though they failed to capitalize on it initially, with OpenAI running with the ball. | |
| ▲ | vidarh a day ago | parent | prev | next [-] | | Most of the time when I see an LLM abruptly switching languages it usually continues with something directly related to whatever triggered it. I'm sure there are other failure modes where it may change topics too, just like humans also regularly digress when triggered by certain words etc. | | | |
| ▲ | TeMPOraL a day ago | parent | prev [-] | | What makes you think the similarity is the superficial part, and not the difference in mechanism? I'd argue it's the latter. > The LLM that abruptly switches languages will also likely switch to a wildly unrelated topic. If a human behaved that way, you'd call a doctor. Haven't met many kids, I see. Or even normie adults talking. I know plenty that tend to jump from topic to topic once they get into a stride talking, and they're not the ones diagnosed with ADHD. |
|
|
|
| |
| ▲ | TeMPOraL 2 days ago | parent | prev | next [-] | | So just like me, including the prompt injections if you count "nerd sniping" as such? (And in particular, switching languages on the fly is normal for people who speak more than one well, it's something you learn not to do for the sake of people less comfortable with the languages involved.) | | |
| ▲ | jacobgold a day ago | parent [-] | | > So just like me, including the prompt injections if you count "nerd sniping" as such? Who would count that as prompt injection? It's a superficial analogy. If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise. | | |
| ▲ | vidarh a day ago | parent | next [-] | | Not nearly all prompt injections are by any means that absolute unless starting from the exact same state. Many of them will also work only probabilistically unless you turn temperature to 0 for exactly that reason. And at the same time, whole books have been written about how reliably we can induce certain behaviours from humans. E.g. the Blue-seven phenomenon [1] - I've personally experienced that second hand and it was how I learned about it by searching for it subsequently because I suspected it was a known thing, having read about cold reading before. A co-worker came back from lunch and recited a story about a cold reader that had run a routine on him exploiting the blue-seven phenomenon, and I knew before the story finished that the answer would be "blue" and "seven". See also Cialdini's book "Influence" which is full of examples of just how predictable peoples reactions are to a whole lot of things. That there isn't a perfect overlap does not mean there aren't plenty of similar "hacks" that causes us to respond in very predictable ways. [1] https://en.wikipedia.org/wiki/Blue%E2%80%93seven_phenomenon | | |
| ▲ | jacobgold a day ago | parent [-] | | Anyone is free to draw whatever analogies they want, but they either make sense or they don't. Outside of philosophical discussions or science fiction, comparing influencing humans to prompt injection is silly. | | |
| ▲ | vidarh a day ago | parent | next [-] | | This is very much a philosophical question, and the only reason you're calling it silly instead of giving an actual argument is that it doesn't support your views. | |
| ▲ | TeMPOraL a day ago | parent | prev [-] | | I wouldn't call it silly, because see the two as directly equivalent and fundamentally the same thing. |
|
| |
| ▲ | TeMPOraL a day ago | parent | prev [-] | | > If you were vulnerable to prompt injection, I could order you to do absolutely anything you're capable of doing and you would be helpless to do otherwise. Yes. Depending on specificity and timescales involved, we call that "reading comprehension" or "social engineering" or "peer pressure" or "motivating literature" or "advertising" or "propaganda" or "religion". |
|
| |
| ▲ | saberience a day ago | parent | prev [-] | | Humans make all these mistakes too, in fact, humans make more mistakes than AI does in coding at the moment. |
|
|
|
|
|
|