| ▲ | Retr0id 5 hours ago |
| > The demo is run on an x86 emulator that records every block of code the CPU actually executes, over the whole demo. > That recording is translated into C - the original instructions, one for one, with the exact cycle timing of the emulated machine. What's the advantage of this approach vs cycle-accurate emulation? (I'd guess it can go faster, due to compiler optimization?) > The result is checked against the emulator event for event: every interrupt, port access and frame at the same moment of emulated time. It's funny how LLMs leak their test procedure docs into user-facing output. |
|
| ▲ | vardump 4 hours ago | parent | next [-] |
| Cycle accurate emulation helps little with this era PC demos, because back in 1993 PC hardware performance variation was already massive. From 386SX-16 to Pentium 66 MHz, that's over 20x difference in performance. |
| |
| ▲ | dspillett 3 hours ago | parent [-] | | > Cycle accurate emulation helps little with this era PC demos It works, and has done for some time, surely. What more is there to add here above accurate playback? > From 386SX-16 to Pentium 66 MHz, that's 10-20x difference in performance. Yep, and some demos had specific requirements for not only minimal CPU but maximal because the assumptions inherent in small-code timing tricks would break beyond a certain speed or because of significant differences in relative instruction execution times (and sometimes the unpredictability of those execution times as the P5 architecture and some of its frankenstein-486-like competitors added branch prediction). Larger demos were more flexible as you didn't need the small-code tricks to squeeze into 4Kb or sometimes less, so 64K demos and larger could be more flexible wrt target CPU. But cycle-accurate emulation has this sorted to: you just need the instruction cycle time accuracy to be matching a particular CPU running to the pace of a particular clock. And given the description of how this is being done (“The demo is run on an x86 emulator that records…”) I assume this is actually using cycle-accurate emulation! I'm guessing the benefit here is that the overhead of playing back the demo on each client machine is lower when playing this recorded version, compared to each viewer's browser running the initial emulation live in the browser. That and doing something a new way was fun or otherwise intellectually stimulating for the dev(s) involved. | | |
| ▲ | vardump 2 hours ago | parent [-] | | Cycle accurate compared to what? Even with a cycle accurate CPU you'd still have a lot of variation caused by the motherboard chipsets, DRAM speeds, cache chips and graphics cards. Tseng ET4000 was on a completely different level than something like Cirrus Logic CL-GD510 or heaven forbid, Oak Technologies card. If the era appropriate PCs had huge variations, what extra does cycle accuracy really bring at this point? The same PC demo could perform very differently on two 486DX2 66 MHz PCs. |
|
|
|
| ▲ | rmnclmnt 4 hours ago | parent | prev | next [-] |
| > It's funny how LLMs leak their test procedure docs into user-facing output. You have to constantly fight so hard to not get this. And for the past few months it seems most people do not even care to remove it and have proper user facing docs |
| |
| ▲ | dspillett 4 hours ago | parent | next [-] | | > You have to constantly fight so hard to not get this. By “fight so hard” you mean “do a bit of basic editing before publishing the text”? You wouldn't have thrown a junior's text straight at end users without any review in the past (at least I hope not, though obviously some teams actually were and still are that lax), why do you expect to get away with skipping the review/edit step with your clockwork colleague? | | |
| ▲ | Sharlin 4 hours ago | parent | next [-] | | Because we expect a junior colleague to learn from feedback and not have to fix the same problems in their output year after year. | | |
| ▲ | dspillett 3 hours ago | parent [-] | | But having learned from the feedback they or you then go off to work on something else and, unless you move onwards/upwards/both in perfect sync, and you end up with new juniors who need the feedback afresh. Also: I've always avoided being anything like a lead or manager, but I do effectively have a couple of relative juniors ATM who I wish would have a better long-term learn/forget ratio when it comes to feedback from myself and elsewhere! |
| |
| ▲ | rmnclmnt 2 hours ago | parent | prev | next [-] | | Okay yes "fight so hard" might have been a bit stronger sorry for that. But it was more in line with "this has become so annoying recently" rightly so because yes you do need to remove the cruft that has nothing to be here and I don't remember being such as issue 6 months ago. The other part of it is to regularly remind colleagues of this rising issue because as I said, even experienced people tend not to care anymore, and this is saddening me. And while in 2025 I considered AI coding agent in the realm of junior devs, we have been way way past that since early 2026 honestly. | |
| ▲ | Retr0id 3 hours ago | parent | prev [-] | | The people who don't care can output slop orders of magnitude faster than those that do. | | |
| |
| ▲ | christophilus 4 hours ago | parent | prev [-] | | I didn’t notice it when using Codex. Recently switched to Claude to test Opus 5.5, and it is constantly adding useless commentary. | | |
| ▲ | rmnclmnt 4 hours ago | parent [-] | | I don’t think that’s specific to Claude though. I’ve seen people use GH Copilot harneds (don’t with which models) and it was even worst |
|
|
|
| ▲ | bananaboy 3 hours ago | parent | prev | next [-] |
| > What's the advantage of this approach vs cycle-accurate emulation? (I'd guess it can go faster, due to compiler optimization?) I don't think anyone is doing it this way because they chose to do it this way. All the AI decomp/recomp projects I've seen recently have done it this way. It just seems an easy way to have an AI brute force "port" something. |
| |
| ▲ | Retr0id 3 hours ago | parent [-] | | This only works for something noninteractive that always executes the same code (unless I'm misunderstanding the LLM's description of what it did) | | |
| ▲ | bananaboy 3 hours ago | parent [-] | | Ah yeah true, no I don't think you're misunderstanding. Actually sorry I misspoke - what I've seen in the game re/decomp projects is that the binary gets turned into a direct C implementation of the instruction stream, and it's basically a sort of emulation with CPU/memory state and emulation of whatever hardware it might need to talk to. | | |
| ▲ | _the_inflator 2 hours ago | parent [-] | | I agree on this description. It looks like a wrapper for an app alias demo, not a generalization for all app following the hardware possibilities at the time. On the other hand, at least these folks build and create and I prefer wrapper demos instead of a perfect emulator that will never finish. Nevertheless, on the feature side of these wrappers I really miss a fast-forward option. If you go pseudo-emulation, then at least come up with some convenience features. Maybe this is the irony: the ff button is perfectly legit, I used mine many times with the 486DX3-100 and Pentium 133. Back then it was called turbo button. |
|
|
|
|
| ▲ | pjc50 3 hours ago | parent | prev [-] |
| Yes, cycle-accurate is very expensive because it prevents optimization. But as described they say they're doing an AOT version of cycle-accurate anyway? In some ways this is similar to the concurrent-systems problems that memory barriers are a tool to address. Your emulator will come in two halves, a CPU side and a gfx/sfx side, as well as control inputs, and the "intermediate" approach is when you're just trying to preserve the sequence of load/stores between the two with an accuracy relative to frame timing. Maintaining ultra precise timing within a frame imposes more detailed coordination requirements on both sides of the emulation. |
| |
| ▲ | whizzter an hour ago | parent [-] | | Being cycle-accurate is annoying when you want full realtime if the hosting platform doesn't have enough CPU power to emulate all parts at full speed. I guess the point is that a pre-compiled version could have HW checkpoints from a slow emulated run, so that it can then run at full speed and just advance the HW to the required timestate when a write is to occur (ie, instead of costly cycle-accurate emulationg, most HW operations can be batched inbetween time-sensitive checkpoints). Then again, why not just record a video at that point? :P |
|