| ▲ | mehmetoguzderin 11 hours ago | |
Good one; reading this briefly took me back to the days of “Needle in a Haystack” being super challenging for LLMs. Maybe there needs to be a benchmark of “Rule in a Haystack” (similar to information one, testing not only independent rules also the ones that need hops) to clarify model performance regarding this. Thank you for the resource. | ||