| ▲ | xpct 2 hours ago | |||||||
You can definitely try to regularize against ruleset changes by generating a bunch of cards and making the agent play in randomized subsets of those cards. I didn't look for prior work on this, but my estimate is that it's probably within 2-3 orders of magnitude of additional training compared to a static game. (Still a lot!) | ||||||||
| ▲ | hnedeotes 2 hours ago | parent [-] | |||||||
But wouldn't (couldn't) the model then hallucinate play patterns and get itself into problems when playing against a real opponent? | ||||||||
| ||||||||