| ▲ | simianwords 3 hours ago | |||||||
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much. There's no reproducible set either. I'm not gonna trust this report. | ||||||||
| ▲ | stymaar 2 hours ago | parent [-] | |||||||
Most people[1] interacting with chatbots don't have a paid subscription and they do interact with the free-tier LLMs that are Luna and Haiku, so I still think it's relevant. [1]: not on HN obviously, but IRL, and probably among FT's readership as well. | ||||||||
| ||||||||