Remix.run Logo
▲ nananana9 5 hours ago

It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.

Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.

The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

▲tczMUFlmoNk 5 hours ago | parent | next [-]

I don't disagree with your general thesis, but I don't think it's right to say that defining behaviors like dereferencing a null would require a null check before every dereference. For example, Java defines the behavior of dereferencing null—it throws a `NullPointerException`. My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled, so that accesses are guaranteed to segfault, and (c) trapping `SIGSEGV` to translate that back into a Java exception.

(Here's a link with a bit more detail: https://courses.cs.vt.edu/cs3214/spring2026/questions/catchi...)

So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!

▲nananana9 4 hours ago | parent | next [-]

> My understanding is that JVMs implement this by (a) representing Java `null` references as the zero pointer, (b) mapping the zero page with write permission disabled

If you want to run everywhere where C does, you can't rely on that. I gave WASM as an example - that's a widely used target that just exposes a flat memory model where 0 is literally just a normal address and there's no way to trap it (unless they've released extensions I'm unaware of). Same deal with most microcontrollers as far as I'm aware of, although I don't do embedded.

I can't think of a way you'd implement null trapping efficiently on those platforms.

▲charleslmunger 4 hours ago | parent | prev [-]

Sure. But consider what would happen if you had an array, and you looked up the nth element. If that base address is a null pointer and the index is greater than the size of your zero page reservation, it'll get some other address which is holding stuff. There's ways to deal with that too, of course, but not for free and it carries implications for other things. In Java this is avoided because arrays carry their length at the beginning, and you check that first for bounds, so if the array pointer was null you'd fault a small number of bytes past 0 and it still works.

Fil-C is amazing and a prime example that undefined behavior means implementor freedom, and the implementor can choose to always trap on null pointer use. Sometimes the implementor freedom doesn't buy you much; for example why should it be UB to do

    (const char*)NULL + 1
Dereferencing null is and should be UB but why is just calculating a pointer problematic? I just did some research and some old architectures would actually trap on creating an invalid address. So if we want C to support those machines, the standard can't define the behavior to do something other than what the hardware does.
▲eru 4 hours ago | parent [-]

That's another argument in favour of implementation defined behaviour, not undefined behaviour.

Btw, a pointer in C doesn't necessarily need to mean an address (invalid or not) in your underlying machine. C is a formally defined abstract language, not portable assembly.

▲eru 5 hours ago | parent | prev | next [-]

> It doesn't make sense to define what happens when you e.g. read from NULL because it's hardware specific - if you have virtual memory of some sort, you'd probably get a page mapping error. If you don't (embedded, WASM), you'd read back whatever value is at that address.

That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.

▲afdbcreid 4 hours ago | parent | prev | next [-]

> The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

IMO UB as "undefined but don't be crazy please" was the original meaning of the standard but people argue on that. It is a fact that compilers didn't exploit UB as strongly back then. However there is a good reason for this change: if you want formal semantics (which you do want, at least possibly) it is pretty much impossible to distinguish the two. If "undefined behavior" is undefined in the math sense, or in formal semantics of languages - the operation can reach any Abstract Machine state, then the fact that you cannot reason about anything follows immediately. The only dubious thing is time-travel, and this was indeed removed in the last version of the standard (and also for Rust now).

▲p1necone an hour ago | parent | prev [-]

> The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".

This is the bit that seems insane to me. Some of the behaviour of C/C++ compilers when they encounter undefined behaviour seems less like 'this is undefined, we'll just do something vaguely reasonable given the context/produce an error' and more like 'ahaha, the user has fallen into our trap, lets fuck them up'.