Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions src/ipips/ipip-0550.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
---
title: "IPIP-0550: PBNode field ordering"
date: 2026-08-24
ipip: proposal
editors:
- name: Alex Potsides
github: achingbrain
url: https://achingbrain.net/
affiliation:
name: Shipyard
url: https://ipshipyard.com
relatedIssues:
- https://github.com/ipfs/specs/issues/533
order: 550
tags: ['ipips']
---

## Summary

Encode `Data` field first in `PBNode` protobuf messages

## Motivation

Regular UnixFS and HAMT-sharded directories are encoded as `PBNode` protobuf
messages.

HAMT-sharded directory entries have the characteristic of prefixing the name of
each entry with a number of characters drawn from the hash of the directory
entry name.

Where hashes collide, a new sub-shard is created with it's own CID/block that
contains a sub-portion of the shard.

The settings used to derive the prefix characters is stored in the `Data` field
of the `PBNode` protobuf message.

This means that all `Link` messages must be read from the `PBNode` message
before we can read the hash algorithm name and fanout values that let us
calculate the prefix length for a given directory entry.

When the reader is attempting to traverse to a single entry deep in the shard,
they are forced to read all entries for the current sub-shard before they can
move deeper within the shard, which leads to inefficient traversals.

## Detailed design

If we allow content authors to write the `Data` field first, readers can apply
a more efficient streaming parser for protobuf messages, since they will no
longer need to read all of the `Link` messages before they can process any of
them.

## Design rationale

Traversing HAMT shards is more expensive than it needs to be, which
disproportionately affects resource-constrained environments and inefficient
runtimes.

### User benefit

Traversing HAMT shards will become faster in resource-constrained environments
and inefficient runtimes.

### Compatibility

Protobuf has no requirement to write fields in any particular order, including
not in the numerical order of the field IDs, so this is a backwards compatible
change, unless custom protobuf parsers are used that expect fields defined in a
certain order. The UnixFS spec does not disallow this so any parsers enforcing
ordering on read may be considered buggy.
Comment thread
lidel marked this conversation as resolved.
Outdated

### Security

N/a
Comment thread
achingbrain marked this conversation as resolved.
Outdated

### Alternatives

N/a

## Test fixtures

N/a

Comment thread
lidel marked this conversation as resolved.
### Copyright

Copyright and related rights waived via [CC0](https://creativecommons.org/publicdomain/zero/1.0/).
25 changes: 19 additions & 6 deletions src/unixfs.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,19 @@ More complex nodes use the `dag-pb` (`0x70`) encoding. These nodes require two s
decoding. The first step is to decode the outer container of the block. This is encoded using the [`dag-pb`][ipld-dag-pb] specification, which uses [Protocol Buffers][protobuf] and can be
summarized as follows:

:::warning
In a earlier version of this spec, the `Data` field of the `PBNode` was ordered
after the repeated `Links` field. This lead to inefficiencies processing HAMT
shards as all links for a node must be read before the hash type and fanout
values could be read, which in turns means that if the ready is looking for a
specific path within the shard, they cannot abort reading links early.

To support legacy data, implementations MUST be able to read and write `PBNode`
messages in the legacy format as well as the current format described below.

Field IDs were the same in the legacy format.
:::
Comment thread
achingbrain marked this conversation as resolved.
Outdated

```protobuf
message PBLink {
// Binary representation of CID (https://github.com/multiformats/cid) of the target object.
Expand All @@ -127,11 +140,11 @@ message PBLink {
}

message PBNode {
// refs to other objects
repeated PBLink Links = 2;

// opaque user data
bytes Data = 1;

// refs to other objects
repeated PBLink Links = 2;
}
```

Expand Down Expand Up @@ -368,12 +381,12 @@ The HAMT directory is configured through the UnixFS metadata in `PBNode.Data`:
- `decode(PBNode.Data).fanout` is REQUIRED for HAMTShard nodes (though marked optional in the
protobuf schema). The value MUST be a power of two, a multiple of 8 (for byte-aligned
bitfields), and at most 1024.

This determines the number of possible bucket indices (permutations) at each level of the trie.
For example, fanout=256 provides 256 possible buckets (0x00 to 0xFF), requiring 8 bits from the hash.
The hex prefix length is `log2(fanout)/4` characters (since each hex character represents 4 bits).
The same fanout value is used throughout all levels of a single HAMT structure

:::note
Implementations that onboard user data to create new HAMTDirectory structures are free to choose a `fanout` value or allow users to configure it based on their use case:
- **256**: Balanced tree depth and node size, suitable for most use cases
Expand All @@ -382,7 +395,7 @@ The HAMT directory is configured through the UnixFS metadata in `PBNode.Data`:
- Trade-offs: Larger blocks mean higher latency on cold cache reads and more data
rewritten when modifying directories (each change affects a larger block)
:::

:::warning
Implementations MUST limit the `fanout` parameter to a maximum of 1024 to prevent
denial-of-service attacks. Excessively large fanout values can cause memory exhaustion
Expand Down
Loading