Remix.run Logo
▲ j2kun 4 hours ago

Functional equivalence here, of course, depends on the completeness of the test suite, where byte-identical compiled artifacts does not.

(For example, your approach would not necessarily catch all the same overflow behaviors; the OP expressly claimed that "replicating all bugs" was also important, and many bugs are caused by certain overflow behaviors)

▲nine_k 2 hours ago | parent | next [-]

> catch all the same overflow behaviors

So you're looking not just for functional but also dysfunctional equivalence %)

▲8note 2 hours ago | parent [-]

no, thats still functional here. the bugs have to be the same, such that speed runs could still run correctly

▲SubiculumCode 4 hours ago | parent | prev [-]

Byte exact seems only of interest to preserve known bugs etc for cheats/shortcuts/etc.

▲j2kun 4 hours ago | parent | next [-]

That may be true, but I hate it when people repeat the false idea that functional equivalence requires only a test suite that has full branch/line coverage. Call me triggered :)

That said, I would probably follow this same approach if I were to do this, but with extensive randomized testing as well.

▲hedgehog 4 hours ago | parent [-]

You can do the process in stages. Do the first decompilation mechanically (no LLM), use a SMT solver to show it builds to an equivalent binary to the original, and then use LLM to clean up the code into something idiomatic with the benefit of a correct binary built with the new toolchain. This helps when you want to port across languages or toolchains, and helps protect against toolchain bugs.

▲kg 2 hours ago | parent | prev [-]

For anything with recorded replays or multiplayer you need to preserve known and unknown bugs for compatibility reasons, not just for cheating.