Remix.run Logo
mghackerlady a day ago

I miss itanium. I feel like a madman for saying it, but I do. Something about it is just alluring to me. Alas, the problems were too hard to solve or weren't worth solving anymore

Ironically, it died around the time LLMs/AI started becoming good. I feel like the compiler problems with vliw could be solved to a degree with a purpose built ai

monocasa a day ago | parent | next [-]

So I feel like I need to write a blog post about this, but succinctly I think there's a good argument to be made that the issue wasn't the compiler despite popular wisdom. I don't even think it was the nature of unpredictable memory access times either as the itanium has a ton of special architectural hardware to handle unpredictable memory accesses (a lot of which are essentially some of the primitives that an OoO core uses for internal bookkeeping, just exposed architecturally).

I just think the arch has a similarity to archs like cell where it was planned for a world without the end of dennard scaling and just stopped making sense when we weren't targeting scaling to 10Ghz consumer CPUs and beyond.

The relatively fixed clock period that makes sense post ~2006 also means that the CPU architecture of that made the most sense ~2006 (Tomasulo OoO cores) continues to make sense, with most of the process gains going to just making bigger, wider cores.

twoodfin a day ago | parent | next [-]

It was also planned for a world where high-end CPUs were differentiated by their ability to run floating-point-heavy workloads with relatively predictable memory access patterns.

The rise of the web—and databases behind it—as the dominant high-end, high-margin workload obsoleted that assumption.

There’s a great presentation floating around where a Compaq-acquired-DEC engineer is trying to justify how great the OpenVMS port from Alpha to Itanium is going, despite benchmarks showing Alpha smoking Itanium running Apache.

wbl a day ago | parent | next [-]

If only the Alpha had hung on.

icedchai a day ago | parent | next [-]

I have an AlphaServer in my collection (DS10.) For a 25+ year old system, it still feels pretty snappy!

hypercube33 a day ago | parent [-]

I think I read at one point alpha had a 64bit cpu in the 90s that could have hit 1ghz but some bug held it back. I am having so much trouble finding these articles and stories about processors these days with AI and shitified Google.

The other one I can't find I swear was a former AMD engineer talking about how Intel bet the farm on one ibm mainframe design but AMD went with an older architecture (under the hood of x86 for both) that was harder to deal with but ultimately runs faster and that's how we got ryzen. Maybe I'm just losing my mind.

touisteur 20 hours ago | parent [-]

I think the "Jim Keller" story around Zen is a bet on modularity, core-complexes, then chiplets. Smaller, less monolithic designs and a clean re-design of the x86 cores for compacity, ease of validation and scalability (in core count).

shawn_w a day ago | parent | prev [-]

My college was a DEC shop; I learned programming and how to use unix in labs full of dumb X terminals connected to Alpha servers.

Still have a lot of nostalgia for the architecture.

eej71 a day ago | parent | prev [-]

I worked for a company that had a fairly large OpenVMS installation and had to make the transition from Alpha to Itanium. It required a considerable amount of work and the early iterations of Itanium did not provide a clear performance improvement over the final Alpha EV7z that we had been using in some GS1280s.

Going by my faulty memory, I'd say it wasn't until Tukwila that it was a clear win over Alpha EV7z. By the time Tukwila arrived, it was pretty clear that Itanium's goose was already cooked.

p_l 19 hours ago | parent [-]

EV7 were introduced after customers refused to upgrade to Integrity over crap performance.

jcranmer a day ago | parent | prev | next [-]

I recall opening up the Itanium manual, and by the end of the architecture description, just despairing of the thought of trying to write a compiler for it. Itanium, I think, was ultimately a victim of its weirdness: it's too weird to really comfortably write assembly by hand; the compilers weren't really capable with its weirdness, so "regular" code was worse off than you'd normally expect. Raymond Chen has pointed out in several articles how the hardware took advantage of C's UB to do some really weird things--and this is an era where most developers expected UB to really be just implementation-defined behavior.

Combine that with the fact that the hardware development process seems to have been compromised from the start (if you told me the hardware architects never looked at anything other than 30-instruction traces of BLAS kernels, I'd believe you), and the insane hype that was built up for it... it's not surprising that it had an extremely underwhelming launch.

bluedino a day ago | parent | prev | next [-]

I wonder if most people kept them around just so they don't have to port their apps or migrate to something new.

I worked at an HP shop, and Itanium ran HP/UX so they kept running their business on their PickBASIC (and whatever database that I've forgotten the name of) system

acdha a day ago | parent | prev | next [-]

There was also a long piece by a former Intel chip designer who was incredulous about how much the Itanium team was promising numbers based on a very few hand-scheduled routines for FPU-limited code. I think there’s a solid argument that the design just wasn’t based on a correct understanding of what most CPUs did and over-indexed on the most performance-sensitive HPC code. I once helped run some HPC code on a test Itanium system and even there it was just so easy to fall out of the only patterns which performed well and end up slower than older Pentiums even before factoring price into the evaluation.

> I said, wait I am sorry to derail this meeting. But how would you use a simulator if you don't have a compiler? He said, well that's true we don't have a compiler yet, so I hand assembled my simulations. I asked "How did you do thousands of line of code that way?" He said “No, I did 30 lines of code”. Flabbergasted, I said, "You're predicting the entire future of this architecture on 30 lines of hand generated code?" [chuckle], I said it just like that, I did not mean to be insulting but I was just thunderstruck. Andy Grove piped up and said "we are not here right now to reconsider the future of this effort, so let’s move on".

https://www.sigmicro.org/media/oralhistories/colwell.pdf

> Davidson also pointed out two areas where academic research could create a blind spot for architecture developers. First, most contemporary academic research ignored CISC architectures, in part due to the appeal of RISC as an architecture that could be taught in a semester-long course. Since graduate students feed the research pipeline, their initial areas of learning frequently define the future research agenda, which remained focused on RISC. Second, VLIW research tended to be driven by instruction traces generated from scientific or numerical applications. These traces are different in two key ways from the average system-wide non-scientific trace: the numerical traces often have more consistent sequential memory access patterns, and the numerical traces often reflect a greater degree of instruction-level parallelism (ILP). Assuming these traces were typical could lead architecture designers to optimize for cases found more rarely in commercial computing workloads. Fred Weber echoed this latter point in a phone interview. Bhandarkar also speculated that the decision to pursue VLIW was driven by the prejudices of a few researchers, rather than by sound technical analysis.

http://courses.cs.washington.edu/courses/csep590/06au/projec...

aleph_minus_one a day ago | parent | next [-]

> http://courses.cs.washington.edu/courses/csep590/06au/projec...

This link requires a log-in.

monocasa a day ago | parent | next [-]

It apparently works with https. Weird. Quite a journey to figure that out.

https://courses.cs.washington.edu/courses/csep590a/06au/proj...

acdha 13 hours ago | parent | prev [-]

Sorry about that and thanks @monocasa for correcting the link. I’ve had that bookmarked for a decade.

kmeisthax a day ago | parent | prev [-]

It's doubly funny knowing that basically nobody bothers doing these kinds of workloads on CPU if they can help it. And everything you have to do to get GPU floating point performance also makes GPUs really, really bad for normal CPU code. Hell, at one point AMD actually was shipping VLIW for shader code...

aleph_minus_one a day ago | parent | next [-]

> It's doubly funny knowing that basically nobody bothers doing these kinds of workloads on CPU if they can help it.

When the Itanium was developed and introduced (2001), nobody was thinking about general-purpose computations. DirectX 8.0, which introduced Shader Model 1.1 (which was far away from being suitable for GPGPU; Shader Model 1.1 was rather about strongly (also size-)limited programs for the vertex and pixel processing stage), was only introduced in 2000, the first release of CUDA was in 2007, and the first release of OpenCL was in 2009.

monocasa 9 hours ago | parent [-]

Yes, but even farther back. Apparently the Pentium Pro team was forked off to work on Itanium in the mid 90s. At that point a GPU was at best just the rasterizer and ROP phases from an acceleration perspective.

aleph_minus_one 8 hours ago | parent [-]

> At that point a GPU was at best just the rasterizer and ROP phases from an acceleration perspective.

For those who are not so deep into GPU architecture:

ROP: Raster Operation Pipeline

> https://en.wikipedia.org/w/index.php?title=Render_output_uni...

acdha 14 hours ago | parent | prev [-]

I think that’s really illustrating how the problem isn’t what they expected: nobody was doing GPU computing in the 90s and when it became huge that was specialized for certain classes of work and involved custom toolchains. The fact that even AMD’s VLIW ended up diverging on later models makes me think Itanium was doomed even if it had shipped on time and budget.

alphabeta3r56 18 hours ago | parent | prev | next [-]

It's more simple. Online bin packing problem is simpler to solve and has more optimal solutions with smaller chunks

a day ago | parent | prev [-]
[deleted]
20k a day ago | parent | prev | next [-]

People have tried all kinds of techniques for VLIW, including techniques that are much better than a purpose built AI, AI isn't a magic silver bullet. Fundamentally there's no reason you can't analyse a piece of code to death, and maximally extract parallelism out of it

The fundamental issue is that there simply doesn't exist enough information to be able to extract the necessary parallelism without a rewrite, its the same issue as trying to autovectorise. You can do it to some degree, but it doesn't work in practice to be able to fill out a very wide architecture with reasonable efficacy

The SIMT programming model has proven to be much more successful vs trying to autovectorise or mash things into a VLIW architecture

ggm a day ago | parent | next [-]

From memory, Glasgow University CS had serious buy in to VLIW models of computation for a while, predating Itanium. There was good reason for believing it might have some interesting behaviours. I think they worked on languages targetting it, data models, things like reversible computation, long lived processes.

adrian_b 19 hours ago | parent | prev [-]

The so-called "SIMT programming model" is a term from the obfuscated jargon that NVIDIA has introduced for CUDA, where they have renamed almost all traditional terms used for decades in the computing literature.

The "SIMT programming model" is the same thing that in 1963 was called "Parallel DO" (from the name of the loop statement in FORTRAN), and later it was more frequently called "Parallel FOR", like in the C/C++ version of OpenMP. In the influential research paper on which later the programming language Occam was based, C.A.R. Hoare referred to the same thing as "an array of processes".

The "SIMT programming model" just means that you write iterative structures where the programmer guarantees that the iterations are independent (unless specified otherwise), so that their order of execution does not matter.

In CUDA it is slightly less obvious than in OpenMP that the so-called "CUDA kernel" is the body of a loop, because the header of the loop is not adjacent, but it is placed elsewhere in the file.

The essential difference between the "SIMT programming model" and writing a "for" loop in C is that the compiler is certain that the iterations are independent. For auto-vectorization, the compiler must prove that they are independent, which can be difficult.

Compiling for a VLIW CPU is in general a much more difficult problem than auto-vectorization (which implements data parallel processing with multiple cores and/or SIMD cores).

A VLIW CPU is able to execute in parallel distinct instructions, not only the same kind of operation like a SIMD CPU, but in most cases there are complex restrictions about what kind of instructions may be combined.

It can be too difficult for a compiler find a schedule for the instructions in such a way that this would allow a maximum number of instructions executed per clock cycle.

Probably the worst part is that it is unlikely to be able to keep busy all execution units without speculative execution of the instructions that are beyond conditional jumps. For these, it is pretty much impossible to guess an optimal schedule at compile-time, because it would change during execution when the same code (in a function or in a loop body) is executed again.

A CPU with dynamic instruction scheduling a.k.a. with out-of-order execution, will change the instruction schedule depending on the predicted branches, with much better chances of keeping busy the execution units.

In HPC applications, where Itanium worked best, it is possible to predict the branches at compile-time quite well, so a compiler has a chance to find an acceptable instruction schedule, unlike in more general-purpose applications, which are hard to predict.

A VLIW CPU also has speculative execution and a branch predictor, because these are needed in any pipelined CPU.

But unlike in an OoOE CPU, the speculative execution handles only complete bundles of instructions that are executed in a given clock cycle. A bundle of instructions may be executed speculatively or not, but the CPU cannot extract speculatively individual instructions from a set of bundles and combine them into a bundle that would occupy all execution units, like in an OoOE CPU.

spott a day ago | parent | prev | next [-]

I think it died before that, it was just that buried it then.

MBCook a day ago | parent | prev | next [-]

It was neat to live through the era where CPUs constantly got faster and they were willing to try such oddball stuff.

For such a long time it became “faster and more cores, don’t be different” and just didn’t seem as interesting.

Apple Silicon had been very interesting to me. I’m really hoping to see a stronger ARM push on Windows, both because I know it can be great and because it’s just interesting. Windows has never had to switch architectures (for consumers) or support two at once for any reasonable population.

Also, whatever happened to mill? We used to get posts about them all the time.

And I wonder what would have happened to Power if they had the 3rd party fabs that exist today instead of being stuck with what IBM could make in-house.

icedchai a day ago | parent | prev | next [-]

Itanium was effectively dead once Intel adopted AMD's x86-64. It just remained a zombie for another 15+ years.

ndiddy a day ago | parent | prev | next [-]

Even if the magical Itanium compiler did exist, Itanium would have still lost to AMD64. As soon as you introduce anything that doesn't behave in a statically predictable manner (multitasking, or virtualization, or even an application that processes unpredictable input like a web backend or database), your performance drops down to a fraction of what a similarly priced AMD64 chip could do. VLIW is great for some very specific workloads like HPC, but Intel should have never tried to replace x86 with it.

MBCook a day ago | parent [-]

The lack of licensing to other companies was clearly a POWERFUL incentive.

Imagine how much money they could make if that pesky AMD went away.

p_l 19 hours ago | parent [-]

IIRC it was major part of intel's roadmap to establish it so that "future" of PC cpus would be locked down to Intel/HP partnership.

And back then VIA was still noticeable competitor!

nl a day ago | parent | prev | next [-]

> Ironically, it died around the time LLMs/AI started becoming good.

What?

I think maybe you are confusing Itanium with something else?

Development on itanium stopped in 2013:

> On 31 January 2013 Intel issued an update to their plans for Kittson: it would have the same LGA1248 socket and 32 nm process as Poulson, thus effectively halting any further development of Itanium processors.[1]

It's true that it shipped until 2021, but I think you had to already have previous orders to get that.

[1]https://en.wikipedia.org/wiki/Itanium

486sx33 a day ago | parent | prev [-]

[dead]