A nice talk about a researcher's experience/benchmarks with raw GPT-4, before and after RLHF:
https://www.youtube.com/watch?v=qbIk7-JPB2c
Yup, I remember that! Microsoft removed that part of the paper.