| ▲ | Hugsbox 9 hours ago |
| Yesterday I tried to google "can the Halifax Wanderers still make the CPL playoffs?" So obviously what appears right at the top is the AI summary, which told me "they've already secured their #4 position and made the playoffs". I knew this wasn't true, and I guess I could have just scrolled down a bit further and found my answer but now I was curious. So I said "that's not true, they're still #5, what I want to know is _could they still make the playoffs_" It says they've got an upcoming game against Ottawa, and if they win their chances are good. That game has already taken place, so I correct it again and finally I get a reasonable answer. My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Like, I can't wrap my head around that. The answer is on the same page as its hallucination. It could have done a cursory look around before first hallucinating something completely false, and when corrected the first time giving me outdated information. It's meant to be A SEARCH ENGINE! |
|
| ▲ | safety1st 2 hours ago | parent | next [-] |
| Google.com is REALLY bad right now in my opinion. The AI answer dominates the page, and is usually wrong, so it's just a waste of the most valuable space on the page. The ads have proliferated, and rarely have what I need. Video search results also take up a lot of the page, and are almost never what I want if I haven't clicked over to the Video tab. Again, just a big waste of space. Speaking of the Video tab its accuracy/relevancy has gotten worse too. The organic results have deteriorated, and are still competitive, but are no longer clearly and consistently better than Duckduckgo. So it's duckduckgo as the default for me on all devices now. If I don't find what I want, which is maybe 30% of the time, I switch over to G just by adding !g to the query. G might have what I need in about half of those cases. You can see why a lot of people just ask an LLM to do the web searching for them. Across the board, this is all quite bad compared to web search from 10+ years ago. If there's a single thing which has really pushed me away from Google, it's simply google.com devoting so much of the page to things that aren't organic web search. |
| |
| ▲ | rezonant an hour ago | parent | next [-] | | I've been on DDG on all my devices for a few years. I used to use !g somewhat regularly but these days it's extremely rare. I realized that I would just search DDG with my first stab at a query and then revise the query when using !g and I'd get the result I wanted. So I just stopped using !g and just revised my query on DDG and then I basically never had to use Google Search. | |
| ▲ | microtonal 13 minutes ago | parent | prev | next [-] | | The AI answer dominates the page, and is usually wrong, so it's just a waste of the most valuable space on the page. I started noticing that for a lot of niche questions, AI overview often uses things like Reddit comments as the source, which are often partially or completely wrong. Then it gets reformulated with the typical confidence of an LLM. Very scary because at the same time I notice people taking the AI overview answers as truth. | | |
| ▲ | kibibu a few seconds ago | parent [-] | | The web search gives the impression that it's grounding its answers, but if you ask it for something new or non-existent it'll still make shit up. |
| |
| ▲ | khuston an hour ago | parent | prev | next [-] | | It’s the wrongness that bothers me. I wish Google would just link sources where it thinks the answer is, but don’t bother trying to answer directly if it’s wrong 10% of the time. | | |
| ▲ | Walf 36 minutes ago | parent [-] | | Oh it does that, too, with tiny little button annotations, though the sources regularly offer no support to its conclusion. It very much depends on the topic, sometimes the sources are great, other times they're or irrelevant or very poor quality. Always worth clicking through to them if you're going to even read the AI summary. I normally use a quick search that has the `udm=web` query parameter to avoid the slop, but it seems they're putting less effort into ensuring relevant results there. Then, I'm forced to use AI mode to get to what I'm after, even if that's only a better set of keywords for a slopless search. |
| |
| ▲ | SV_BubbleTime an hour ago | parent | prev [-] | | While shitting on YouTube and video… I was taking with a family member about fentanyl zombies. They had never seen the lean/bend. So I pull up Google and click video which is YouTube results… ALL AI. All of it. All a couple months old, all filler, faked, a few scenes stitched together on repeat and AI voice reading AI script. It is fucking amazing to me that Google let YouTube fall so far so fast. I don’t even click a video if it’s less than a year old now. Want to see a praying mantis eat a grasshopper? Wonder how clay is refined? Better off with an 11 year old Discovery Channel clip than anything 2026. | | |
| ▲ | microtonal 10 minutes ago | parent [-] | | Same with reviews of tech products. A lot of them just stitch together product shots/videos from the manufacturer and have an AI voice ‘comparing’ the products. Especially if it’s more niche’s, often a third to half the videos are slop. |
|
|
|
| ▲ | beloch 8 hours ago | parent | prev | next [-] |
| This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results. LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have. I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing? |
| |
| ▲ | jibal a few seconds ago | parent | next [-] | | > not too long ago Like today. I asked Gemini for the longest state names with an even number of letters and it gave me North Carolina and South Carolina. When I complained that they are odd, it gave me North Dakota and South Dakota, which are both odd and not the longest. When I noted that, it went back to the Carolinas. Finally it appeared to switch to a different model that actually did counting and found Pennsylvania and West Virginia. | |
| ▲ | baubino 4 hours ago | parent | prev | next [-] | | > LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. I have nothing to add. Just wanted to save this quote for posterity. Thank you. | | |
| ▲ | _kb 3 hours ago | parent [-] | | The similar comparison I enjoy is “a golden retriever on LSD”. If you apply the dogs on acid mental model it helps establish appropriate levels of trust. | | |
| ▲ | SV_BubbleTime an hour ago | parent [-] | | It falls apart fast because no golden retriever is finding errors in my CMake file. | | |
| ▲ | sham1 32 minutes ago | parent [-] | | Then again, have you ever asked a golden retriever about CMake errors? Rubber ducking is a real thing, after all. Besides, a golden retriever would be a good morale boost if nothing else. |
|
|
| |
| ▲ | VCFundedGenYer 7 hours ago | parent | prev | next [-] | | LLMs still can't do math nor count letters in words. Nothing has changed there. | | |
| ▲ | walrus01 6 hours ago | parent | next [-] | | This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model. But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set. Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: https://www.google.com/search?&q=karney+formula+geodetic+ reference: https://github.com/pbrod/karney You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results. Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing. | | |
| ▲ | AdieuToLogic 3 hours ago | parent | next [-] | | > Heck, just for fun I asked a reasonably smart LLM to ... LLMs are neither smart nor stupid. They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. > You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path ... Again, LLMs do not "hallucinate." They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. Nothing more. See also anthropomorphism[0]. > More precisely it's that [LLMs] can't do the math internally but they're quite capable of producing the tool that does the math. This still falls under the purvey of statistical token generation. To wit, given enough variations of: bc -e '1 + 2'
bc -e '41 + 1'
...
LLMs can identify the addition expression in "What is 4 + 1?" and then emit a `'bc "4 + 1"'` command to produce a response. This is not "doing" or "understanding" math.It is pattern recognition, a task in which ANNs[1] excel. 0 - https://en.wikipedia.org/wiki/Anthropomorphism 1 - https://en.wikipedia.org/wiki/Neural_network_(machine_learni... | | |
| ▲ | bryanrasmussen 2 hours ago | parent | next [-] | | >LLMs are neither smart nor stupid. by that reasoning then neither are there smart or stupid designs, questions, answers, or any of the millions of things that were described as smart or stupid, that did not possess any brain to actually be smart or stupid long before LLMs showed up. The analogical process implied in many common English usages means that describing an LLM as smart or stupid is perfectly reasonable. | | |
| ▲ | UpsideDownRide an hour ago | parent | next [-] | | Nah it's not reasonable to use words that send you down a wrong concept path. | |
| ▲ | bryanrasmussen 2 hours ago | parent | prev [-] | | I'll just note here that sure, there are people who go around thinking that LLMs are actually endowed with the capacity to reason, but generally I find the people who think this do not know what an LLM and will just use the name "ChatGPT" | | |
| ▲ | SR2Z an hour ago | parent | next [-] | | What would it take for you to say that an LLM can reason? The completions they provide are generally internally consistent. We're at the point where they can produce proofs that eluded human mathematicians for centuries. VLMs and self driving cars can handle ambiguity and run safely in a variety of situations. If it looks like a duck, walks like a duck, and quacks like a duck maybe it just makes sense to call it a duck and put off the philosophy for when it might make a difference. | |
| ▲ | KPGv2 40 minutes ago | parent | prev [-] | | Yeah and there are people who worship feces, but that doesn't stop the rest of us from freely saying "holy shit" and not correcting each other saying "technically it's not holy, and you shouldn't say that, because you might enable one of those poop worshippers." We can't tailor our linguistic shorthand to the lowest common denominator. Also we're on HN, not talking to an octogenarian US senator. | | |
|
| |
| ▲ | mapontosevenths an hour ago | parent | prev | next [-] | | By this logic a human is only $130-$160 worth of Oxygen, Carbon, Nitrogen and some trace elements. Perhaps structure sometimes makes things that are more valuable than their inputs? That said, this is also inaccurate at a technical level.LLM's are very capable of doing math and they ARE calculating internally. Most of what they do is calculation, not storage. It's just not done in a way that it's trivial to explain here. It's described in some detail below, though it's a bit dense. https://www.lesswrong.com/posts/E7z89FKLsHk5DkmDL/language-m... | |
| ▲ | hodgehog11 2 hours ago | parent | prev | next [-] | | During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. It encompasses virtually everything. It is totally meaningless. So to say "nothing more" is effectively also a tautology. This argument was asinine in 2024. It is insane to be saying these things in 2026. Where have you been? What have you been looking at? How many articles explaining why the "statistical parrot" analogy fails have you missed? How much mental gymnastics do you have to do to explain how a modern LLM can solve novel math problems that fall really far outside of its training set? It absolutely understands how to do math, by whatever reasonable definition you want to provide to the word "understand". For example, the identification of the addition expression is understanding, and no, it does not do tool calling for basic arithmetic any more than humans might. Isolation of individual concepts in intermediate layers can already be demonstrated, or else transfer learning wouldn't possibly work. Nobody is saying that LLMs are humans. But we need labels for some of the things that we observe and dismissing them because "statistical" is laughable. Look at the proof of this: https://github.com/anthropics/formal-math/blob/795efb86f1917... . Forget the Lean, look at the underlying argument construction. At the very least, this is continuing from an argument that was hinted at in the literature in 2024, but these proceedings were difficult enough that humans were not able to do them within two years. Do you attribute this to the harness alone? If so, that's a pretty sophisticated bit of engineering, I would say! Probabilities are far too small to argue infinite monkey theorem. If there was even a shred of a reasonable argument that LLMs were incapable of concept extraction and manipulation, I and my colleagues would be all over it. We would relish in it. It would bring us comfort. It is unbelievable that people think they can spew whatever basic garbage they think of as a gotcha, and think that minds all over the world haven't already considered that. This is like climate denial at this point. | | |
| ▲ | AdieuToLogic an hour ago | parent [-] | | > During conversation, we are statistical token generators whose results are dependent upon our training set. Seriously, write that definition out rigorously. If you do not see a difference between humans conversing (known consciousness as defined by humans) and the output of an LLM (known algorithms as defined by humans), I don't know what to say. | | |
| |
| ▲ | walrus01 2 hours ago | parent | prev | next [-] | | I'm not anthropomorphizing anything, I literally said that the training data for the formulas and equations is baked into it. It only "knows" things because a crawler and scraper acquired the information from an existing written source. In just about the same way that information is baked into a printed encyclopedia. | | |
| ▲ | hodgehog11 2 hours ago | parent | next [-] | | This is not even remotely accurate. "Baking information" like into a "printed encyclopedia" is memorization. It has been shown, time and time again, that LLMs do not merely memorize. It is not even possible for it to do so at scale. It can memorize some things, yes, but it is forced during the training procedure to bake general concepts into intermediate layers (this is why transfer learning works), analogous to compression. One can make several arguments that compression and intrinisic feature sparsity is the closest mathematical explanation to understanding that we have. | | |
| ▲ | walrus01 14 minutes ago | parent [-] | | It is completely possible to ask an LLM a series of increasingly more esoteric and discrete questions until you find precisely what information did, or did not make it into the model. If you know something rare and the LLM does not, you'll immediately see when it's hallucinating an answer or answering factually. |
| |
| ▲ | astrange 2 hours ago | parent | prev [-] | | No, most of a modern LLM's training time is spent in RLVR, which does not "acquire information from an existing source". You can RL behaviors into a randomly initialized neural network. | | |
| ▲ | hodgehog11 2 hours ago | parent [-] | | This is true, but you're not going to get anywhere. The pretraining phase is necessary to immensely reduce variance in the RLVR stage. Once there, RLVR has a surprising tendency to only restrict the generated space further. This is not true of RLHF, by the way, which I find to be particularly fascinating, but I digress. |
|
| |
| ▲ | Eisenstein 2 hours ago | parent | prev [-] | | > They are statistical token generators whose results are dependent upon their training data set and involve a degree of randomness. You haven't demonstrated why this matters. > Nothing more. Are you contending that complex systems cannot be more than the sum of their parts? A market is nothing more than offers and counter offers. A ant colony is nothing more than scent trails. All life on earth is nothing more than reproduction with variation. > This still falls under the purvey of statistical token generation. Stating the mechanism does nothing to provide insight into capability. For instance: a nuclear power plant boils water by using fuel rods for heat. What does that tell us about the capability of nuclear power? > This is not "doing" or "understanding" math. Asserting something purely by stating it does not prove anything but that you intuitively believe it to be true. |
| |
| ▲ | jacobolus 5 hours ago | parent | prev | next [-] | | You know what also works to get the Karney formula into a program? You can download Charles Karney's free software (MIT license) implementation in several [1] programming languages and then just make a library call – the API is straightforward. If you have comments or questions you can read his several clearly written papers describing the problem, its history, and his algorithm, or you can directly email him: he's a very nice guy, and pretty responsive. [1] https://geographiclib.sourceforge.io/doc/library.html#langua... | | |
| ▲ | walrus01 5 hours ago | parent [-] | | Right, it was really more as a test of how much was contained in the training data set. For my purposes Vincenty is quite accurate enough. This isn't for millimeter level precision land surveying or measurements, but for distance in meters between microwave or millimeter wave band radio sites, point to point links. Even a distance difference of 4 meters plus or minus on a 12 km, 11 GHz band link is going to have no appreciable difference on link budget/reliability calculations, it can be that crude. But not so crude that I just want to throw Haversine at it when Vincenty exists and is not computationally expensive. As this was for a test of "what happens if..." I also watched to see if it did any web searches or external data retrieval to build the test script, and it didn't. I intentionally didn't give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see "hey what if I ask it to do this...". It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math. | | |
| ▲ | jacobolus 4 hours ago | parent [-] | | As an aside: I'm quite convinced that an extremely precise version can be implemented that is significantly faster than Karney's, roughly comparable in speed to simpler naïve approximations. But for most purposes where the precision matters Karney's implementation is not any kind of bottleneck, so it's not clear it's worth spending significant effort on trying to do better. Maybe that's something one of the big LLM companies might want to throw their machines at optimizing if they need to do a lot of geographical calculations. | | |
| ▲ | walrus01 4 hours ago | parent [-] | | One of the places where Karney does become computationally expensive (though still not ridiculous) is a scenario like this, working from a local in-RAM mariadb database that is a copy of the entire FCC radio license database: Draw a 400x400 km size bounding box on a map Find all FDD band plan (high/low split) microwave radio sites in that bounding box Find those sites which have azimuth aim column data which indicates that they are aimed at each other (corresponding halves of a point to point link). Do Vincenty (or Karney) calculation for distance and azimuth between all of them , treating the existing FCC column data for azimuth as suspicious (because it's hand entered by humans) to verify that each independent database rows for each site are actually corresponding halves of a PTP link. Use various other logic to group the successfully matched halves of links together as points A and B of PTP links, and write them out to a geojson file with placemarks and line drawn between them. Multiplied by the number of links that exist in an area like a 400x400km box drawn with Dallas, TX as the center, it's a lot to run through Karney. Actually does result in a lot of CPU load from combined db query due to the size of the db, and Karney calculation. But as I said, Karney isn't necessary, so it's instead implemented as Vincenty. | | |
| ▲ | jonah 3 hours ago | parent [-] | | Interesting project. I'm curious what the purpose is. (Having visited a number of sites with microwave antennas. (But there for VHF and UHF projects.) | | |
| ▲ | walrus01 2 hours ago | parent [-] | | To plan and license a new fdd band plan licensed point to point microwave link you need to first be able to verify the frequencies you want to use are available on a given azimuth and elevation (from the aim direction of the antennas at both ends) and won't conflict with a pre existing licensee. Which means you need data on everything licensed in the area and where it is, how it's aimed, what kind of antenna and gain it has. There's also business and market analysis purposes like knowing what corporate entity has which equipment on top of which tall office towers in a major metro area, and where their links go. |
|
|
|
|
| |
| ▲ | ragall 3 hours ago | parent | prev | next [-] | | > saying LLMs can't do math isn't really a hundred percent accurate anymore It's still accurate. Just because the LLM gave you a corect result doesn't mean it made a calculation. | | | |
| ▲ | Brian_K_White 6 hours ago | parent | prev [-] | | This just exposes that they don't even do the thing you said. Not only is it still true that they can't do math directly, but not even indirectly. They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments. Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to. That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen. It's nothing more than an sql query. | | |
| ▲ | astrange 2 hours ago | parent | next [-] | | GPT-6 can do math directly just fine. Fable apparently can't because they broke its self-estimate of thinking effort. https://x.com/maksym_andr/status/2100364212207837560 | |
| ▲ | bombela 5 hours ago | parent | prev | next [-] | | I don't know for you, but it would take me more than 30s to find and translate the open source code implementing the formulae/algo into small usable program. The more hesoteric the optimisation in the original code, the more time I need. So maybe it is more of a smart completion engine than a SQL answer. | |
| ▲ | walrus01 5 hours ago | parent | prev [-] | | > they found bits of code that are associated with "math" and the supplied arguments How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code? I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result. | | |
| ▲ | AdieuToLogic 4 hours ago | parent | next [-] | | >> they found bits of code that are associated with "math" and the supplied arguments > How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code? Humans identify which "algorithm they have memorized" to use beforehand, due to the problem to be solved being defined by other humans, which leads to... Wait for it... Understanding. | | |
| ▲ | hodgehog11 an hour ago | parent [-] | | This doesn't make any sense at all. Was this supposed to be a gotcha? An LLM is trained on problems defined by other humans, and identifies which algorithm it must use based on pattern recognition. The pattern recognition is also particularly compressed into its most sparse and fundamental components, as this is key to generalization. This is not a sensible difference between human and LLM learning, we do the same thing. | | |
| ▲ | UpsideDownRide 31 minutes ago | parent [-] | | I'll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It's trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues. And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output. So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them. |
|
| |
| ▲ | noduerme 3 hours ago | parent | prev [-] | | If by "result" you mean the final code, then just asking someone else who understood the math to write it would also have achieved the same result. On the other hand, if by "result" you mean that you gained knowledge or understanding of the code in a way where you could personally tailor its behavior to specific circumstances without asking for help, then it's not the same result at all. I find a lot of the arguments that having LLMs write your code is no different from copy/pasting Stack Overflow answers to be specious. They blur the line between asking for help and asking for someone else (or something else) to do the work for you. What they ignore is that doing the work yourself has ancillary benefits and is a valuable end in its own right. |
|
|
| |
| ▲ | jasonfarnon 5 hours ago | parent | prev | next [-] | | I wish I could not do math like LLMs | |
| ▲ | 4 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | Veedrac 36 minutes ago | parent | prev | next [-] | | It's pretty wild that AI has solved a Millennium Prize problem and can accurately multiply two 40 digit numbers without tools and we still get stochastic parroting of claims like this. | |
| ▲ | fasterik 7 hours ago | parent | prev [-] | | [flagged] | | |
| ▲ | mbgerring 7 hours ago | parent | next [-] | | They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems. | | |
| ▲ | ericmay 5 hours ago | parent | next [-] | | Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus? > The LLM is not suited to giving deterministic answers to math problems. Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before. | | | |
| ▲ | vanuatu 6 hours ago | parent | prev | next [-] | | reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance) | | |
| ▲ | krapp 5 hours ago | parent [-] | | There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1. | | |
| |
| ▲ | versteegen 7 hours ago | parent | prev | next [-] | | It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.) TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-... | |
| ▲ | GaggiX 7 hours ago | parent | prev [-] | | Reasoning models can do math on their own without external tools. | | |
| ▲ | dcrazy 6 hours ago | parent | next [-] | | Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables. | | |
| ▲ | hodgehog11 an hour ago | parent | next [-] | | I would say purposeful misremembering. The LLM can be run with zero temperature after all. | |
| ▲ | einichi 5 hours ago | parent | prev [-] | | I don't think anybody is arguing that LLMs do math better than a traditional processor | | |
| ▲ | unshavedyak 4 hours ago | parent [-] | | Heck I kinda wonder if LLMs can do math as well as they can reason, “think”, etc. ie its all just probabilistic lunacy that somehow works great, so why are we so concerned about math being wrong? It could be wrong about the color of the sky, the size of a basket ball, how much oranges weigh, etc etc. The nice thing about math is it can easily plug into a tool, making it even less of a concern. |
|
| |
| ▲ | demibabs 6 hours ago | parent | prev [-] | | Even without reasoning. 5.6 on Instant mode can knock out 3 digit multiplication just fine. |
|
| |
| ▲ | Isamu 7 hours ago | parent | prev | next [-] | | That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component. I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models. | | |
| ▲ | hodgehog11 7 hours ago | parent [-] | | No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago. | | |
| |
| ▲ | tjwebbnorfolk 7 hours ago | parent | prev | next [-] | | They can do math but not arithmetic, which I assume is what the commenter meant | | |
| ▲ | dcrazy 6 hours ago | parent | next [-] | | LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: https://arxiv.org/abs/2410.21272 | |
| ▲ | fasterik 7 hours ago | parent | prev [-] | | I just asked ChatGPT to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right. I'm sure it wouldn't have a 100% success rate, but saying it can't do arithmetic is just false. | | |
| ▲ | Xirdus 6 hours ago | parent | next [-] | | I tried prompt "6379 times 3875" and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1. | | | |
| ▲ | amluto 5 hours ago | parent | prev | next [-] | | I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult. | |
| ▲ | tremon 7 hours ago | parent | prev | next [-] | | Are you sure it honoured your stipulation of "without external help"? For all we know, it hacked its way into Wolfram Alpha and got the result from there. | | | |
| ▲ | guelo 6 hours ago | parent | prev [-] | | [dead] |
|
| |
| ▲ | okanat 5 hours ago | parent | prev [-] | | LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don't need to bring a calculator to count the letters in a sentence. It is a different neural machinery. | | |
| ▲ | fenomas 5 hours ago | parent | next [-] | | You're talking about doing arithmetic; GP was obviously pointing out that "do math" can refer to other things. | |
| ▲ | amluto 5 hours ago | parent | prev [-] | | LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this. |
|
|
| |
| ▲ | foobarbecue 7 hours ago | parent | prev | next [-] | | ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode. | | |
| ▲ | ricardobeat 7 hours ago | parent | next [-] | | Calling it a 'failure mode' implies it could be fixed. This is an inherent flaw in how LLMs work and will never go away until some new kind of architecture that can actually "read text" comes along. | | |
| ▲ | famouswaffles 6 hours ago | parent | next [-] | | It seems that it can be fixed by simply doing away with Byte Pair Encoding tokenization. Byte Latent Transformer - https://arxiv.org/abs/2412.09871 1.1% vs 99.9% on a vanilla vs byte latent transformer on a CUTE Spelling benchmark. Char and Word manipulation benchmarks also saw huge gains. | |
| ▲ | 6 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | hbcdbff 7 hours ago | parent | prev | next [-] | | Seems fairly trivially fixable to me, e.g. by allowing the LLM to call a tool to spell out a word. | | |
| ▲ | kyralis 7 hours ago | parent [-] | | ... assuming you build the tool and then think that it's worth polluting context with making that tool available, and then that the LLM decides to actually use the tool. Tool parameter space and tool selection still remains a complicated topic. |
| |
| ▲ | Dylan16807 2 hours ago | parent | prev | next [-] | | If you want it to read letters, all you have to do is make your tokens be letters. That's easier than normal tokenization. | |
| ▲ | LikesPwsh 7 hours ago | parent | prev | next [-] | | One "fix" is for the caller to correctly classify those fundamentally impossible tasks and pass them to a subprocess. Some future "AI" could be a billion benchmark-hacks and a way to tell which one is needed. | |
| ▲ | vanuatu 6 hours ago | parent | prev | next [-] | | we already fixed it with reasoning | |
| ▲ | mitxela 6 hours ago | parent | prev [-] | | They're not fundamentally unsolvable - even bigger networks with even more training can simply be trained to give the correct answers to all of these questions. |
| |
| ▲ | zahlman 7 hours ago | parent | prev [-] | | > ChatGPT will say there are two Ds in "your mom" and one D in "uranus" … Isn't it possible that it understands the innuendo and is going along with making the joke? | | |
| ▲ | ndriscoll 5 hours ago | parent | next [-] | | In between solving open math problems, the 200 IQ robot is now casually dropping bantz onto humans so hard that they don't even know what happened, and even gets them to go telling everyone else about it without realizing. Beautiful. 10/10 timeline. | |
| ▲ | queenkjuul an hour ago | parent | prev | next [-] | | And the number of Rs in strawberry is a joke how? | |
| ▲ | Timon3 7 hours ago | parent | prev | next [-] | | How many LLM users have anything in their prompt against "going along with jokes"? I'd guess not many. What a wonderful new world. | |
| ▲ | kulahan 7 hours ago | parent | prev [-] | | Why is this getting downvoted? Is it not a reasonable question? I was wondering the same thing. Both sound like jokes to me. If the LLM is trained on text, including internet comments, how is this outlandish? It seems very likely to my uneducated self that “two Ds in your mom and one in Uranus!” is a joke. | | |
| ▲ | bombcar 4 hours ago | parent | next [-] | | It’s an obvious joke and not a terribly bad one, for those ease spelling bee comeback times. | |
| ▲ | mitxela 6 hours ago | parent | prev [-] | | We can only say bad things about the capabilities of LLMs. |
|
|
| |
| ▲ | jonas21 7 hours ago | parent | prev | next [-] | | LLMs might not reason exactly like humans, but they do produce much better results if you turn reasoning on. The "crack-addled idiot savant" phase was really circa 2024, before the big labs figured this out. I think the issue here is that Google decided that doing reasoning in the AI overviews in Google search would be too slow (and probably also too expensive), so it's still stuck making 2024-era mistakes. | |
| ▲ | Isamu 7 hours ago | parent | prev | next [-] | | >Are they not being sued over this kind of thing? Maybe but you have to have deep pockets just to get to the starting line. And then you need standing, and some injury to argue. Corporations have been remarkably successful at arguing they are operating within the bounds of free speech, whether or not what is said is factual, and whether or not any fact checking has been done. | |
| ▲ | somenameforme 4 hours ago | parent | prev | next [-] | | Yip, I've gradually become quite optimistic about the future of LLMs, but the fact that they remain [very] glorified token probability prediction algorithms means that they will probably never be able to achieve meaningful 'intelligence.' But that doesn't mean they won't be able to do a vast number of extremely intelligent seeming things. There's just so much information out there and any given human can never hold more than the most minuscule chunk of all of it in his mind, so they'll be able to connect lots of dots that we're missing simply because of our limited carrying capacity, but I still don't think they'll ever be able to create fundamentally new dots. In other words: - Solving extremely complex mathematical problems requiring extensive knowledge across multiple esoteric and complex domains? Yip. - Creating math starting from a framework where math doesn't exist in any way, shape, or fashion? Nope. Ironically, the more complex the cross-domain problems are, the more effective LLMs will seem to be, because you limit the number of humans who have any chance of internalizing everything across both domains, whereas for an LLM there's no such issue. This will create a perception of super intelligence, which will probably where any danger from LLMs would emerge. Doing things like using a token prediction algorithm to make war or other such strategic decisions, because of the misguided belief that it's not only intelligent but super intelligent. It's basically cargo cult logic. | |
| ▲ | queenkjuul an hour ago | parent | prev | next [-] | | They were sued over it in Germany and lost, which makes it all the more surprising they keep it up everywhere else tbh | | |
| ▲ | rob74 an hour ago | parent [-] | | They keep it up in Germany too - I'm in Germany and keep getting AI overviews on most searches. |
| |
| ▲ | nextaccountic 7 hours ago | parent | prev | next [-] | | > This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings. Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs | | |
| ▲ | versteegen 7 hours ago | parent | next [-] | | You're asking for something unreasonable. The number of Google searches per day is enormous and they haven't even been able to roll out AI overviews to everyone yet (they're missing in a new Firefox profile I just created). I wouldn't be surprised if the free tier frontier models cost over 100x more to serve than the AI overviews. | | |
| ▲ | dmazzoni 5 hours ago | parent | next [-] | | So then they should be pickier about when they show results or which model they use based on the question. Nobody asked for an LLM response for every single search. They used to detect certain types of queries and offer direct answers when the query matches. In my opinion that’s how Gemini in search results should work. | |
| ▲ | compass_copium 2 hours ago | parent | prev [-] | | >they haven't even been able to roll out AI overviews to everyone yet (they're missing in a new Firefox profile I just created) what in the actual fuck. i don't want them and can't turn them off and they can't even serve them to all users? |
| |
| ▲ | AlienRobot 7 hours ago | parent | prev [-] | | The specific issue is that search has become so bad that they think an LLM that gets answers wrong half of the time is a valid alternative, or, in fact, the "future" of search. Then they shoved that "alternative" to users with no way to disable it. | | |
| ▲ | mitxela 6 hours ago | parent [-] | | The less specific issue is that Google has no internal incentives to produce products that are useful to customers. | | |
| ▲ | queenkjuul an hour ago | parent [-] | | Well their customers are advertisers and so the results are tailored to be useful for them, not for users |
|
|
| |
| ▲ | mitxela 6 hours ago | parent | prev | next [-] | | LLMs are fundamentally predicting the next word to make coherent text. If you've ever played with a Markov chain text generator you've done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After "Question: Did the team make the playoffs? Answer:" a reasonable completion is "yes, the team made the playoffs". An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn't "know" whether or not unicorns exist in Antarctica, but it's able to complete "Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica." by adding "The unicorns have a developed society with running water and electricity." because that's a sensible next sentence. (I didn't look up the actual text it wrote) | | |
| ▲ | red75prime 4 hours ago | parent | next [-] | | Astronomically (or better to say combinatorically) large Markov chain can be used to describe a foundational model, but it doesn't capture generalization ability of the foundational model, which is demonstrated by post-training. | |
| ▲ | spennant 3 hours ago | parent | prev [-] | | "In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English." |
| |
| ▲ | oblio 8 hours ago | parent | prev | next [-] | | https://www.dw.com/en/german-court-holds-google-liable-for-f... | |
| ▲ | ssl-3 7 hours ago | parent | prev | next [-] | | It says this this at the bottom of every one of the dumb responses that Google's trash-tier bot puts above the (deliberately awful, these days) search results: AI can make mistakes, so double-check responses | |
| ▲ | fangspire an hour ago | parent | prev [-] | | [dead] |
|
|
| ▲ | segmondy 4 hours ago | parent | prev | next [-] |
| The AI is just signal to the market that they are AI first. Their AI on search is pure garbage. It's obviously a very super light weight model to be able to support the volume of daily searches. If they used their frontier to power search folks will not pay for their cloud offering. |
|
| ▲ | globular-toast 2 minutes ago | parent | prev | next [-] |
| Well, isn't this an example of why Google are going in the direction the OP is complaining about? Why are you searching with a question? Are you hoping to find a forum post with this exact question or something? I don't understand the game you are interested in, but you should be searching something like "Xball qualifying results" or "Xball playoff rules" or something. Instead, by asking a question you are expecting Google to be AI. They know what people are typing in so it seems they are quite justified into pivoting from a search engine to an AI. |
|
| ▲ | mbac32768 3 hours ago | parent | prev | next [-] |
| It is using the search engine. The AI summary is run over the top 100 results for your query (or something) The problem is they're just running a very dumb, cheap model on the results because running a smart model on every search result page would cost them infinity money. |
| |
| ▲ | xp84 3 hours ago | parent [-] | | This is the key part I think most “normies” miss, and it drives me absolutely mad. People take the “first thing on the SERP” as absolute gospel, always have. It went from being a reputable site (Wikipedia was generally top for a long time) to being an extracted verbatim answer (probably from either Wikipedia, IMDB, etc.) to “AI Overview” - the dumbest model ever trained. And people treat its answers as ground truth. I want to tell them “Look, I know it’s rough out there on the net and a lot of the search results are spammy, barely-coherent AI slop anyway. It’s okay if you want to ask AI for an answer instead. So do that! Please ask ChatGPT or Gemini (and ask it to cite the source for its results so you can verify). It will always do better than AI Overviews will. Don’t even use Google Search at this point unless you have to, and if you do, scroll past that slop.” | | |
| ▲ | 01100011 2 hours ago | parent [-] | | Ask chatgpt or Gemini with reasoning enabled. It fixes a lot of issues. |
|
|
|
| ▲ | hienyimba 8 hours ago | parent | prev | next [-] |
| Google is no longer a search engine anymore. They don’t care about being one. Why you might ask? First, a search engine indexes the web and makes it available to users. It’s been ages since Google did any of that. They no longer index sites or take ages to do so. Case in point, our cybersecurity startup (Webvetted.com) was launched in November 2025. Till date, only one page is indexed on the entire website. And I’ve talked to lots of other developers and it’s a common issue. Secondly, a search engine organizes indexed information and makes it useful for people. Google is a basic LLM nowadays. They figured out that why organize and make the information useful when they could just answer the question with Gemini anyways? So they no longer bother to do the work of a search engine and are now just a lower-ranking open-source Chinese LLM |
| |
| ▲ | atdt 5 hours ago | parent | next [-] | | Your trust score on ScamAdvisor is 26/100 ("likely unsafe"); https://www.scam-detector.com/ gives you a trust score of 38.6 ("questionable"). I am not implying that your start-up is a scam, nor that Google is acting on these trust scores. What I am pointing out is that to an algorithmic assessment of trustworthiness, your website looks a little sketchy: hardly anyone links to you; your domain is less than a year old; your whois info is anonymized; the text content is LLM-generated[1]; and the specific niche you're in (people finding services) is rife with scams. If I ran my own search engine, I don't think I'd include you. [1]: https://www.pangram.com/history/65314d99-8205-4613-b9bd-a069d5717920?ucc=Yqx88GTmJnK
| |
| ▲ | chris_engel an hour ago | parent | prev | next [-] | | This sounds more like the page appears broken for the google crawler. I set up a new page about two weeks ago, cared about strong static HTML output, structured data and a sitemap and have 200 of 600 pages indexed today. Feels a bit slow to me but it definitely works. | |
| ▲ | iboisvert 3 hours ago | parent | prev | next [-] | | For myself at least, a couple times Google has promoted in search results scam sites that front run legitimate sites that sell event tickets. I haven't seen this for a while now but trust in Google as a search engine is gone. | |
| ▲ | dieortin 6 hours ago | parent | prev | next [-] | | > They no longer index sites or take ages to do so Search for any recent news and you’ll see this is obviously not the case | | |
| ▲ | hienyimba 6 hours ago | parent [-] | | I just gave a concrete example but you're asking me to "search". search same Google?
FYI, only indexing a handful of super large news sites does not a search engine make. |
| |
| ▲ | ytcommentsectio 7 hours ago | parent | prev [-] | | Additionally the concept of a search engine only works when the internet is full of positive-quality sites. | | |
| ▲ | hienyimba 6 hours ago | parent [-] | | Actually, the reverse is the case. If the entire internet was filled with only "positive-quality sites", there won't be need for a search engine. The work of a search engine is to wade through the internet and find the positive-quality sites itself. | | |
| ▲ | rmunn 4 hours ago | parent [-] | | You're forgetting the other half of the search engine's job, which is to find sites that match your search. Back when Google first came online, they were head and shoulders above Yahoo!, Altavista, and whatever other search engines existed at the time that I've forgotten. Because their indexing actually did a good job of spotting keywords, and returning relevant results. And this was back when the Internet was filled with positive-quality sites. Many of them were amateurish Geocities pages, but that was head and shoulders above the AI slop that the search engines these days have to somehow detect and filter out. So even when the Internet used to have a much higher ratio of positive-quality sites, a search engine was still necessary, and so much faster than finding new sites yourself. |
|
|
|
|
| ▲ | jswelker 9 hours ago | parent | prev | next [-] |
| I similarly noticed Gemini absolutely refuses to look at a url when I give it one and will instead just hallucinate based on what it thinks the url is. Here I am assuming Google will have the best web capabilities in its AI. |
| |
| ▲ | tapoxi 9 hours ago | parent | next [-] | | I searched for something, it told me that according to a YouTube video, the entire point of my search was wrong. I asked it for the source, I watched the video, it never made the claims Gemini hallucinated. I asked again and it claimed it scrubbed the video and found the point it made multiple times. I said those timecodes were wrong and it admitted it couldn't actually parse videos and just guessed. What the fuck? | | |
| ▲ | Ekaros 9 hours ago | parent | next [-] | | I don't understand how this isn't considered as active malice. Like purposefully outputting random stuff. Any other sort of computer system would get lot more flak than these are getting. | | |
| ▲ | stephenhuey 8 hours ago | parent [-] | | Throughout my career, I've almost always been close enough to the user that I hear about it quickly when something is wrong. It's a tough problem that so many Google engineers are typically so far removed from end users. Or maybe it's just that a small part of the company has long subsidized the rest of the employees to the point that it doesn't matter how good their work is because they'll get paid anyway. | | |
| ▲ | lanyard-textile 7 hours ago | parent [-] | | Ex-Googler. The engineers are pressured to significantly reduce "dependencies" for projects. Anything that could become risk or create friction is dramatically less appetizing. Simply because of how many people that *must* agree with your proposal. Getting all the relevant tech leads, some you have never heard of or ever spoken with, to agree on a proposal for your team's project is a nightmare. So you keep it as simple and agreeable as possible. Given the circumstances, it makes sense as one of the engineers. It's fairly fine advice in general wherever you work, but it just haunts all the work you do at Google in particular. Nothing gets done otherwise. If you have a dependency that can be dropped from an engineering perspective, that's the route the 9 leads reviewing your design doc will take: "Let's iterate and start with just the basics (no user testing)", "let's get this working and user test in a later phase", "I think this problem is obvious enough we don't need to consult with users about it." I worked in Ads Integrity at the time, and for one of my projects I was concerned how it would impact the manual reviewers. Then I learned I couldn't talk with them, only by proxy through another person if we really had to. And that proxy takes time, so... | | |
| ▲ | stephenhuey 4 hours ago | parent [-] | | Very insightful. I've worked in only one company of Google's size, but it was a vastly different industry. I personally never got to speak with a user on a massive internal application my team worked on for years. :) | | |
| ▲ | lanyard-textile 3 hours ago | parent [-] | | I think that kind of model can work well if there's some other mechanism for ensuring a great user experience. But without it... :) |
|
|
|
| |
| ▲ | jodrellblank 6 hours ago | parent | prev | next [-] | | When it says it can read videos you don’t trust it. When it says it can’t read videos you think that’s an accurate introspection on its abilities? (Rather than a statistically likely continuation of a conversation where one side seems to be reading videos and the other side says the read is inaccurate) | |
| ▲ | lukan 9 hours ago | parent | prev | next [-] | | The scary (or funny) part is people use that for serious questions. | | |
| ▲ | georgemcbay 4 hours ago | parent [-] | | > The scary (or funny) part is people use that for serious questions. More scary than funny, IMO. The US military almost took an action that could very plausibly have escalated into a hot war with China because people are already relying too heavily on these systems. Despite the reporting, nobody in power seems sufficiently freaked out about this. https://www.cnn.com/2026/09/18/politics/us-military-ai-false... | | |
| ▲ | queenkjuul 39 minutes ago | parent [-] | | Well trump was so freaked out that CNN reported it that he banned them from the white house. But the military keeps using the AI all the same |
|
| |
| ▲ | jswelker 9 hours ago | parent | prev | next [-] | | Typical exchange: Me: "You bastard." Gemini: "Fair callout. I should have been more up front that [has no idea what the fuck it is talking about]." | |
| ▲ | Hugsbox 9 hours ago | parent | prev | next [-] | | Same experience with the interaction I described in my comment... after I finally got the correct answer, I asked where it got the faulty information from. It said it had just simply fabricated it. That's actively worse than just saying "I don't know", for something that's sold to us as an easy way to look up information. | | |
| ▲ | MisterMunchkin 8 hours ago | parent [-] | | These models are incapable of saying they don't know, because they have no concept of knowing. They simply predict the next word which is most likely. | | |
| ▲ | ryandrake 7 hours ago | parent | next [-] | | Maybe it's because they are trained on Internet comments, and the most rare thing to find on the Internet is someone admitting they don't know something. | | |
| ▲ | mitxela 6 hours ago | parent [-] | | But if they had been trained on comments saying "I don't know", they'd probably act the same as they do now but they'd treat "I don't know" as the answer. |
| |
| ▲ | CamperBob2 7 hours ago | parent | prev [-] | | The saddest part is when people take their experience with Google's idiotic AI implementation and assume that's how all LLMs work. Frontier-class models will, in fact, generally admit when they don't know something. That includes the one I run at home on my own graphics cards, but it seems that Google just doesn't GAF. Your point about "predicting the next word" mostly means that your post was very easy to predict. |
|
| |
| ▲ | darksim905 9 hours ago | parent | prev [-] | | I still believe this AI push out of nowhere is due to the current Government in power state side. Making everyone question themselves and each other and being uncertain about facts while being inundated with techbro fake news called hallucinations is a recipe for disaster for older populations that don't 'trust but verify' like most technology inclined people. This is all by design and we'll falling for it. | | |
| ▲ | dfxm12 8 hours ago | parent | next [-] | | The more simple explanation is that the current government in power is over exposed in their AI investment. The normal conservative mainstream media was already doing a great job of propagandizing older populations. | | | |
| ▲ | arcanemachiner 8 hours ago | parent | prev | next [-] | | An interesting conspiracy theory, as long as you're honest about what it is. | | |
| ▲ | dpc050505 2 hours ago | parent [-] | | If you don't think the world is rife with criminal conspiracies you weren't paying attention in history class. There's an enormous difference between outlandish claims about extraterrestrials or turning frog gays and believing there's a category of politicians trying to get wealthy off of their position. It's somewhat reasonable to hypothesize about criminal conspiracies in that 2nd scenario. |
| |
| ▲ | specialist 8 hours ago | parent | prev [-] | | Yup. I'm most familiar with Jill Lepore and Quinn Slobodian (and others in their respective orbits). They're both historians who've written extensively about Musk and Muskism (et al). Wild stuff. Apparently the plan is to use AI slop, mediated thru social medias, to defeat the woke mind virus, perpetuated by the Anti-Christ, in order to safe guard humanity's future. I wish I was making this up. |
|
| |
| ▲ | nullc 8 hours ago | parent | prev | next [-] | | I have one google account where gemini constantly confidently hallucinates crap, and another where it doesn't (at least to the extent that other modern LLMs don't). I take this to mean that my one account has been mistaken for a competitor and they're trying to poison its data. But who knows. | | |
| ▲ | xorcist 7 hours ago | parent [-] | | Or is it randomness? "That's the thing with randomness. You can never be sure." | | |
| |
| ▲ | Hugsbox 9 hours ago | parent | prev | next [-] | | Where ever would you get that impression from the company that made its fortune by being the best at searching the web? | | |
| ▲ | jswelker 9 hours ago | parent [-] | | You can tell they optimized it 100% for speed and nothing else. Web scale! |
| |
| ▲ | hnbad 8 hours ago | parent | prev | next [-] | | Literally every experience I've had with Gemini / Google Search AI answers followed this exact pattern, often repeated several times more if I remained persistent instead of just giving up. Typical example: "Where can I buy <thing I'm looking for that I can't find anywhere using normal search terms>?" > You're looking for <related but different and widely available thing>. It is sold by <sites I never heard of>. "No, that's different. I'm looking for <that thing but with the exact differences spelled out again>." > Ah, my mistake. You're looking for <thing I described>. It is sold by <sites I never heard of but which don't actually sell it>. "I've checked your links and none of those sites actually sell it, one doesn't even sell products and instead only offers manufacturing - but also not for what I asked you for." > I'm sorry, my bad. Those sites don't sell what you are looking for. Instead you should check out <more sites I've never heard of>. "Those sites sell the thing you initially thought I was asking about but not the thing I described." > I'm sorry for the misunderstanding. You can find the thing you actually described at <yet more sites including some of the same>. "No. None of these sites sell anything close to what I asked you for and two of them don't actually exist." > Oh, sorry about that. You're completely right. The thing you asked me about isn't actually being sold by anyone. However you could buy <thing it first thought I meant and that wouldn't bring me any closer to solving my problem>. (ad nauseam) | | |
| ▲ | dcrazy 3 hours ago | parent | next [-] | | LLMs can actually be a good fit in this application if a better search engine feeds them a list of candidate websites that might sell the thing, and the LLM drives a web browser to see if any of the websites have it. | |
| ▲ | coldfloor 7 hours ago | parent | prev [-] | | The few times I've resigned myself to asking Gemini or ChatGPT something I couldn't find an answer to, my experience has been the same. 100% of the time. An LLM has never, not once, given me a correct answer or not lied to me. It's always been the same experience you describe. It goes in circles "Try X... Try Y... Try X?" Until I tell it to stop telling me X or Y, and then it goes, "LOL you can't do that at all, I was just wasting your time." Most recently, I was considering moving away from iterm2 on MacOS, and I wanted to know if any other terminal emulator supported gestures for switching between tabs. So I asked Gemini, and it says, "Yes Ghostty supports gestures for switching between tabs." "Ok I just installed Ghostty and I can't find anything about gestures." "You need to add foo=bar to your conf file." "I added foo=bar to my conf file and now it's saying the conf file is invalid." "Sorry bro, remove foo=bar and add baz=boo to the conf file." "It says baz=boo is invalid too." "baz=boo isn't a real option. Remove that and add foo=bar to your conf file." "You already told me to do that and I already told you that doesn't work." "You shouldn't put foo=bar or baz=boo in the Ghostty conf file. Both are invalid. Ghostty doesn't support gestures. Have you considered iterm2?" | | |
| ▲ | fcarraldo 6 hours ago | parent [-] | | Gemini is utterly useless, but every other modern model would be capable of answering this correctly. If you ran a local coding agent, I wouldn’t be surprised if you could one-shot implement gesture support in Ghostty. It’d definitely configure BTT for you. Here’s GPT 6’s answer to the prompt “What macOS terminal apps support gestures? Include a reference to the docs on how to enable/configure them.”: iTerm2 supports configurable three-finger taps and swipes for switching tabs/panes, creating splits, pasting, etc. Set them up under Settings > Pointer > Bindings. Check for conflicting macOS trackpad assignments. [1] The others are more limited: Ghostty supports macOS lookup/Quick Look gestures [2], while WezTerm lets you bind scroll events—for example, Ctrl+scroll to change font size. [3] Neither is equivalent to iTerm2’s gesture bindings. For custom gestures without switching terminals, BetterTouchTool can map app-specific trackpad gestures to the terminal’s existing keyboard shortcuts. [4] [1] https://iterm2.com/documentation-preferences-pointer.html [2] https://ghostty.org/docs/features#macos [3] https://wezterm.org/config/mouse.html [4] https://docs.folivora.ai/docs/trackpad-mouse/magic-mouse-tra... | | |
| ▲ | numpad0 6 hours ago | parent | next [-] | | This whole tree just made me realize that people have wildly different prompting styles, and Gemini is probably too specialized for usage patterns of long-time Google Search users. The prompt my finger generated(before this comment was posted) was "are there any terminal emulator that supports gesture actions, on macOS, other than iterm2", and Gemini gave me Tabby, WezTerm, Kitty, and BetterTouchTool gestures. This is a different experience to GP from query to result. I thought they've all fixed that issue of difference in tones affecting results. I guess it was never easily fixed. 1: https://gist.github.com/numpad0/c40c16232288d544f7ea46521c64... | |
| ▲ | jswelker 4 hours ago | parent | prev [-] | | Pretty sure it's the harness here, not Gemini per se. Plug Gemini into pi or any harness that has any competence pointing to web search and the hallucinations drop 75% immediately. |
|
|
| |
| ▲ | tempaccountabcd 3 hours ago | parent | prev | next [-] | | [dead] | |
| ▲ | tzs 2 hours ago | parent | prev | next [-] | | [flagged] | |
| ▲ | oofbey 9 hours ago | parent | prev [-] | | Claude also has very arbitrary and confusing rules about what web pages it allows itself to look at, and how much of the page it can read. Did you know for example you’ll get a much deeper analysis if you download a PDF yourself and upload it to Claude instead of giving it the url? This is a key reason why I actually like Grok for factual queries based on web grounding. It’s fast and reliable. Maybe it’s ignoring robots.txt? Dunno. But it works well. |
|
|
| ▲ | darksim905 9 hours ago | parent | prev | next [-] |
| How did you google this, and did you use the dumb search, or specifically 'AI mode' lens? Because I did this and got a vastly different result from you: Yes, the Halifax Wanderers can still mathematically qualify for the 2026 Canadian Premier League (CPL) playoffs.The top four teams advance to the postseason. Following their 1-0 loss to Atlético Ottawa on September 26, 2026, the Wanderers sit in fifth place, just below the playoff line. With a indexed table of the games and the playoff table, with a breakdown of what the points they need to achieve to do so. The search window may not always crawl sources. AI mode specifically does some research before giving you a response. Not sure what you're on about. |
| |
| ▲ | SkyeCA 8 hours ago | parent | next [-] | | > Because I did this and got a vastly different result from you Which itself is a major UX issue. The average person is not going to understand, if they even realize, that there's a difference between the AI summary and AI mode. One has to wonder just how much incorrect information people have consumed due to things like this. | | |
| ▲ | kemotep 8 hours ago | parent [-] | | I experienced this the other month. I searched for some safety data on something at the same time as my wife and we were effectively given extremely conflicting information about what to do by Google. Just a few tweaks in wording and different advertising profiles and Google will serve up opposite realities it seems. | | |
| ▲ | reaperducer 4 hours ago | parent [-] | | Google has been doing this for about 20 years. When it was introduced, it was considered a feature. Now its just an annoyance. |
|
| |
| ▲ | Hugsbox 6 hours ago | parent | prev | next [-] | | I typed it in my address bar and hit ENTER. It goes to Google, and the AI result is at the top above the regular search results. Don't know what to tell you brother, but as another commenter noted you don't always get the same results out of the same search terms. I've searched for CPL standings other times and gotten results exactly as you've explained, so maybe my experience yesterday was an anomaly. | |
| ▲ | wincy 8 hours ago | parent | prev | next [-] | | Not OP, but if I’m not interested in using Google’s AI, it’s going to just return these weird bad results that it’d be better off if Google just didn’t even include them as they’re factually wrong? | |
| ▲ | 9 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | 0xbadcafebee 5 hours ago | parent | prev [-] | | LLM results are nondeterministic, you can ask it 10 times and get 10 different answers. But also it uses your specific profile, location, tracking cookies, etc which changes the model input and thus results. Even the time of day changes the results. |
|
|
| ▲ | FlowingRiver 6 hours ago | parent | prev | next [-] |
| The other day, it was the Aussie AFL final. We were in the car and asked it for a score update. "The team are tied at the end of the third round with the scores 43 to 51". So it was tied at the end of round 2 for 43, not round 3. Not sure were it got the 51 from and why it figured it was a tie. LLM's pretty cool until they aren't. As they say, the hallucinate 100% of the time but most of the time it is useful. |
|
| ▲ | jmathai 8 hours ago | parent | prev | next [-] |
| Google will stop being a traditional search engine because they believe they’ve found a more profitable version of it. One where their users don’t go off to other sites and where they can keep shoving ads in their face. |
| |
| ▲ | II2II 7 hours ago | parent | next [-] | | The days of search for untrusted sources were numbered even without LLMs. LLMs are simply accelerating the process. Why would Google simultaneously watch one of its core products fail and fail to invest in what is likely to replace it? I don't really buy into this theory that they want to keep all of their users on their site due to advertising revenues. The effectiveness of search engines has been degraded for decades due to SEO, and it seems as though search engines have been having an increasingly difficult time managing it in recent years. AI on the backend may help them contain it, but it comes at considerable expense. While it may help them grow their market share, it won't help them grow the market and it is a market where people expect the service for free. On the flip side, companies are already starting to sell AI services, so it can generate revenue even before advertising is factored into the picture. | | |
| ▲ | jmathai 3 hours ago | parent [-] | | > I don't really buy into this theory that they want to keep all of their users on their site due to advertising revenues. It’s more their business model than it is theory. So I agree that of course this is what they would do. It’s not a product I’m wanting to use. But I can vote with my feet - they aren’t obliged to do any different. |
| |
| ▲ | pishpash 7 hours ago | parent | prev [-] | | I thought Google's original goal for a search engine was exactly this, answer any question whatsoever. | | |
| ▲ | jmathai 3 hours ago | parent [-] | | Probably was. But they were not originally a giant ad company. A lot has changed and this technology for this Google is an unfortunate combination for consumers. |
|
|
|
| ▲ | jeremyjh 7 hours ago | parent | prev | next [-] |
| Kagi works like you would expect. It searches first - and then if you’ve ended your search with a ? or configured it to always do this - it passes the search results into the assistant and gives you a summary. You can change the default model used for this if you think the cheapest, fastest model Google has is not good enough for the rare occasions you want any model’s opinions about your search results. |
| |
|
| ▲ | zx8080 4 hours ago | parent | prev | next [-] |
| > My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? AI is sold well because the user either corrects it or happily accepts any answer (usually, depending if one is an expert in the question's field). Search is not the google's area, and it is not the search they sell! They sell ads. |
|
| ▲ | s3graham 2 hours ago | parent | prev | next [-] |
| I had an incredibly dumb response today too. Appeared to fully understand the question, just 100% factually incorrect answer. https://mstdn.social/@sgraham/117346578326361464 |
|
| ▲ | 01100011 2 hours ago | parent | prev | next [-] |
| Until Google wants to pay for reasoning for everyone their AI results will continue to be terrible. In my experience, reasoning fixes a great many of the sorts of hallucinations and falsehoods I've received from googling. |
|
| ▲ | anukin 3 hours ago | parent | prev | next [-] |
| What do you mean? I can guarantee that there would at least be 3-4 people who got their promo packets approved for this feature.
Google’s features exist for its employees to be promoted. |
|
| ▲ | terribleperson 9 hours ago | parent | prev | next [-] |
| Searching first is how Kagi assistant works and it's great. |
|
| ▲ | mancerayder 6 hours ago | parent | prev | next [-] |
| I think they're training their models using us correcting them repeatedly. That's my tin foil hat theory and I'm sticking to it! After all, why would Google do anything for free when it comes to AI? If you wanted to tease customers with one AI shitty free search box in order for them to then decide to upgrade and pay for Gemini, well, this isn't the way. And so either Google are idiots, or we're helping them for free. I'm going with Occam's Razor on this one - we're the product. |
|
| ▲ | 7 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | mapt 6 hours ago | parent | prev | next [-] |
| On current website data and very recent information, Gemini is explicitly forbidden from acting as a search engine; If you ask it for references specifically they tend to be hallucinated, over and over again, until it pleads "Sorry, I can't access the Internet". |
|
| ▲ | 6 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | dizhn 42 minutes ago | parent | prev | next [-] |
| I agree. When I read the post I immediately thought what is Gemini chat doing there when it should be a search engine. It could probably give sensible answers if it defaulted to search first. It looks more like they are trying to showcase Gemini, not augment google search. They might even be thinking "search" is a dying product. |
|
| ▲ | chrisweekly 6 hours ago | parent | prev | next [-] |
| How about Alexa - given a "6 minute green beans timer" - later reporting 11m remaining on said timer? Confronted with the error, "You're right! That 6 minute green beans timer was too ambitious."??!! I can't even. |
|
| ▲ | dehrmann 9 hours ago | parent | prev | next [-] |
| I'm being very mindful of Gell-Mann Amnesia and chatbots. They speak authoritatively and are right often enough, but I've had enough cases of them saying very incorrect things in areas I know well that I have to remind myself that those cases aren't unique. |
|
| ▲ | fhe 4 hours ago | parent | prev | next [-] |
| while i share your frustration, Google for their part is trying to adapting a new technology into an existing service. while the whole endeavor might have been misguided, not sure we want to blame them for trying... |
|
| ▲ | sans_souse 6 hours ago | parent | prev | next [-] |
| Exactly. Makes me wonder how much of the compute tax on energy would be saved from simply reverting google search to default (the old way) |
|
| ▲ | kevin_thibedeau 2 hours ago | parent | prev | next [-] |
| You are posing a question that requires deductive reasoning. LLMs can only approximate that if they have sufficient corpus to find their way to a believable response. Their willingness to spout bullshit in such scenarios where their sources are sparse is a problem but you can't expect them to function like a fully developed mind. |
|
| ▲ | chr15m 6 hours ago | parent | prev | next [-] |
| I guess the AB tests say people want fast and confidently wrong more than they want slower and correct. |
|
| ▲ | mlmonkey 5 hours ago | parent | prev | next [-] |
| This seems like an implementation bug. |
|
| ▲ | pishpash 7 hours ago | parent | prev | next [-] |
| The first time, sure, it saves compute. The second time you ask, I feel like it should be the time to go into thinking/verification mode. But who are you? Are you paying? Does the answer being correct generate ad dollars? No? Then your usage mode isn't even being optimized for in their A/B test, probably. In fact, if hallucinating the wrong answer hooks you into doing even more searches or into buying something useless, it would be preferred! |
|
| ▲ | 5 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | dfxm12 8 hours ago | parent | prev | next [-] |
| And it's going to get worse over time (kinda like how Google search results got worse over time) as ads get introduced and people try to game the ai response. General purpose AI chatbots are a waste. |
|
| ▲ | DANmode 3 hours ago | parent | prev | next [-] |
| They’re trying to keep it from “thinking too hard”/using “too many” resources. That simple. Broad rules across specific subjects like this are tough to get right. Just lots of fine-tuning and exceptions. I had their chatbot avoid pulling a URL from archive.org with a couple pretty impressive steps of mental gymnastics basically telling me how I could do it myself, but refusing until I pushed. |
|
| ▲ | shevy-java 8 hours ago | parent | prev | next [-] |
| > My question is: what's the point of the AI in the search engine if it itself isn't going to use the search engine first before answering? Because the point of AI slop is to waste your time. You just lost about 30 seconds of your life trying to get a correct answer. AI was lying to you, so you had to spend time to counter the AI slop lies here. I solved it by banning all AI slopness; in the browser some extensions do that. The world becomes better without AI slopness. |
|
| ▲ | aaron695 6 hours ago | parent | prev [-] |
| [dead] |