Why do cryptocurrencies use Merkle trees instead of hashing all the data in the block in one go?
Reddit r/cryptography discussion explaining why blockchains use Merkle trees rather than a single block-wide hash: inclusion proofs, SPV light-client verification, partial verification without full data, and second preimage concerns.
Overview
This is an archived question-and-answer thread on Reddit’s r/cryptography asking why cryptocurrencies structure block data as a Merkle tree rather than concatenating every transaction and computing a single hash. The core answer from participants is that a plain block-wide hash proves only that the whole block is intact, whereas a Merkle tree lets anyone prove that one specific transaction is included in the block using a short logarithmic-size path, without possessing or downloading the rest of the block’s data. This capability is what makes lightweight (SPV) clients and selective verification practical.
Key points
- A single hash over all block data proves only integrity of the entire block; it cannot demonstrate membership of one transaction without revealing and re-hashing everything.
- A Merkle tree yields an inclusion (audit) proof: a logarithmic number of sibling hashes that let a verifier recompute the root and confirm a leaf’s presence, without the full transaction set.
- This underpins Simplified Payment Verification (SPV) and light clients, which hold only block headers (each containing a Merkle root) and request compact proofs on demand.
- The thread also cites an efficiency angle: recomputing the root after a single leaf changes touches only the sibling hashes along one path rather than re-hashing the whole data set. The property most emphasized for cryptocurrencies, however, is compact partial verification via inclusion proofs.
- Discussion touches on second preimage weaknesses in naive Merkle constructions (for example ambiguity between leaf and interior nodes, and Bitcoin’s duplicated-last-hash quirk), which formalized designs mitigate with domain separation.
Relevance to Truestamp
Truestamp relies on the same property this thread explains: it publishes only a Merkle root per block and issues compact inclusion proofs so a single item can be verified against that root without exposing other items. Its tree uses the RFC 6962 leaf and node domain separation (Truestamp’s Merkle tree), directly addressing the second preimage concerns raised in the discussion.
Citations
- Why do cryptocurrencies use Merkle trees instead of hashing all the data in the block in one go?. Reddit r/cryptography discussion thread (2022).