Remix.run Logo
▲ mort96 8 hours ago

Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds: /linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.

I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.

▲amluto 7 hours ago | parent | next [-]

The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.

▲mort96 7 hours ago | parent [-]

The design of Git, as a Merkle tree, is meant to allow for use cases like this:

* I host a mirror of the Linux git repo.

* You download Linux from my mirror.

* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).

* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.

This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.

If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.

▲amluto 6 hours ago | parent [-]

> You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism

That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.

Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:

a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.

This breaks if there is a malicious collision already in the tree.

b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.

I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.

▲mort96 5 hours ago | parent [-]

For A), I don't understand what the point is? I never mentioned what Linus is confident about, I talked about what you can verify when you pull from my mirror. I could replace a commit from 2010 with a malicious one

For B), I would think this could work, but it's a completely different solution from what you proposed and what I responded to.

▲amluto 4 hours ago | parent [-]

> I could replace a commit from 2010 with a malicious one

How? Remember, there are (currently, anyway) no known SHA-1 preimage attacks.

▲mort96 3 hours ago | parent [-]

We're discussing a hypothetical situation where SHA-1 gets even more broken. From my original comment in this thread (https://news.ycombinator.com/item?id=49924179#49925367):

> If I can forge commits with any SHA1 hash at will

We probably don't want to wait until there are practical pre-image attacks discovered to change away from SHA-1.

▲amluto 3 hours ago | parent [-]

This is a fair point.

I think I stand by my second proposal. I also think it's absurd that, after all these years, upstream git still can't figure out a credible migration plan.

▲PunchyHamster 3 hours ago | parent | prev [-]

If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?

The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered