| ▲ | bluecalm 2 hours ago | |
So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers. I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail. I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly. I am sticking with my first intuition: people are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well). They also have vested interest in the conclusion (they are financial advisory firm) so there is that to consider. People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess. | ||
| ▲ | Dwedit 36 minutes ago | parent [-] | |
LLMs use random numbers, so a single test won't necessarily match someone else's experience. | ||