Remix.run Logo
▲ hnedeotes 3 hours ago

No, well, in MtG you could interpret it as meaning such but what I mean is that if in the training set sequence A-B-B-A when state is C-A-X-Y is the play 80% of the time, then you have a new card (that doesn't need to be combo) that by sheer mechanics thwarts that then that strategy won't stick by the addition of that single card to the opposing deck (that you can't know if your opponent is playing or not) and having one or 2 or 3 or 10 different cards renders every calculation very problematic as a play can be the best or the worst depending on such simple things diluting further the best play as the pool grows. Then you need to take into account in MtG shuffling and drawing. I think it's fair to say it's much more difficult to model... And while an agent can learn new combos, you just need to read the card once, the agent needs to be retrained.

▲ironSkillet 2 hours ago | parent [-]

Doesn't this entirely depend on the latent embeddings of strategies and game space in the AI model, which may not be so concrete and explicit as you've described? That's kind of the magic of LLMs with coding, they can generalize because the abstract patterns are encoded in latent space, not the specifics.

▲hnedeotes an hour ago | parent [-]

I might be wrong but what I was thinking was that in chess (or even imperfect information games with a much smaller "range" such as Stratego,) a model can calculate all possibilities for all moves and following moves, by itself and opponent up to a depth that the human cannot. So it can see everything that can happen if it does move X-Y, then Y-Z, then A-C and figure out one that is unbeatable no matter what (or at worse leads to a draw).

But on MtG in particular that never really applies in full due to drawing new cards. You can play perfectly and still lose due to sheer randomness of draws.

The latent space I'm not sure how it translates to a game playing bot, but I would imagine that it would open it up to fail in the same ways a human fails.

On the game I'm designing it could do that (calculate all possibilities up to X depth, for all possible scrolls and table states) but it would be extremely expensive to do so (not a very good argument if compute power keeps increasing), but more than that, in contrast to something like chess, there can be many more paths and decision points where a bad decision turns into a loss, so if it assumes that the best play is X at some point, a sequence that it discarded due to not being the most probable can exist and the bot can never be sure, so if it makes a decision that plays into a "trap" he can't undo to a favourable position. While in Chess it's much clearer what is possible from a given state, it's unambiguous and the rules are fairly limited.

In stratego you have a 10x10 board game, a very clear objective and at most 40 pieces (with repeated pieces and simple mechanics amongst them), while in MtG and similar games a single piece (card) can have probably hundreds of different interactions depending on everything else going (and everything else hidden), at many points of decision. In stratego it also seems that for humans at least, most moves are "inconsequential", as it probably plays more at the psychological/bluff level. Maybe a human player that was given the same budget for training could spend a month training against bots might fare better as the strategies might be then better understood (by the article it's mentioned that the agent recovered from bad positions, so it seems that it was mostly human error, as the human was playing better up to that point).

While on MtG or Asummon, although there can be inconsequential moves (they don't matter given the context/stage of the game), every move carries with it a possibility of being consequential in unpredictable ways. Anyway, there should be ways of training models with just a rule abiding client for these games, without codifying all rules, that they can just keep playing to figure out the interactions, so if that theory is true then it should be possible to create an unbeatable bot - I'm just not sure it is without infinite time/compute and less so if the "meta" keeps changing rendering possible training inconsequential regularly.