| ▲ | Gander5739 4 hours ago |
| If the output can't be trusted, and you use another llm whose output can't be trusted to check the untrusted output of the first llm, then you're back where you started. |
|
| ▲ | duncangh 2 hours ago | parent | next [-] |
| Yeah this seems to me similar to how the mortgage backed security risk concentration occurred leading up to the global financial crisis. Whereby the risk from exposure to low grade / risky single mortgages was eliminated via diversification but the diversification was simply packaging all of the risky MBS’s together and in no way diversified or de-risked the entire portfolio |
| |
| ▲ | cromka an hour ago | parent | next [-] | | I don't see it. To me it's like having e.g. 3 drunk PhDs arguing between each other to settle on truthful answers to questions. | | |
| ▲ | DrewADesign 4 minutes ago | parent [-] | | But PhDs are PhDs because they’d look it up in an authoritative source, or actually find out through research and experimentation. The whole point is that facts aren’t a matter of opinion. The only people that argue over documented, findable facts are idiots that nobody should listen to. |
| |
| ▲ | Terr_ 2 hours ago | parent | prev [-] | | I'm hoping that the Big Horrible Realization comes sooner rather than later, when we have less collective damage and pain riding on it. (Plus I'd feel personally vindicated.) |
|
|
| ▲ | daishi55 2 hours ago | parent | prev [-] |
| Not really. Take hallucinations for example. If they are 1 in 100 (actually they are much rarer, but for the sake of argument), then the chances that 2 LLMs or even just 2 runs of the same LLM have the same hallucination is, well, a lot less than 1 in 100. |
| |
| ▲ | Terr_ 2 hours ago | parent | next [-] | | That rests on a false-assumption that the errors are statistically independent events, and have nothing to do with the shared nature of the judges. | | |
| ▲ | daishi55 an hour ago | parent [-] | | Are there any reproducible hallucinations on any of the currently available OAI/Anthropic models? I’m not aware of any. And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output. | | |
| ▲ | Terr_ an hour ago | parent [-] | | If simply running things thrice-over was enough to stop "hallucinations" (and not incur other problems) we wouldn't be here talking about it today, it'd have been "solved" months or years ago. |
|
| |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|