| ▲ | Orca-Bench: How Ready Are Language Model Agents for Oncall?(arxiv.org) | |||||||||||||
| 13 points by yruzin 2 hours ago | 4 comments | ||||||||||||||
| ▲ | 4di an hour ago | parent | next [-] | |||||||||||||
looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben... This doesn't work anymore. Is there a newer link? | ||||||||||||||
| ▲ | dash2 an hour ago | parent | prev [-] | |||||||||||||
Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them. | ||||||||||||||
| ||||||||||||||