| ▲ | rao-v 5 hours ago | |
Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed) | ||
| ▲ | bunderbunder 5 hours ago | parent [-] | |
We should be careful about extrapolating from benchmarks to real life though. In medical applications, various forms of AI have been beating health care practitioners at specific benchmark tasks since the 1990s. They still have a pretty poor track record of real world success. The real world is not all that similar to a benchmark, as it turns out. | ||