Repository navigation
Conversation
There was a problem hiding this comment.
Pull Request Overview
This PR introduces SWIP-39, a smart neighbourhood management system for decentralized service networks. The proposal aims to solve the "one operator, one node in a neighbourhood" problem through a balanced assignment mechanism that ensures fair load distribution and prevents sybil attacks.
Key changes include:
- A comprehensive specification for balanced neighbourhood registry with random assignment
- Smart contract implementation for managing node registration and neighbourhood assignments
- Mathematical formulations for neighbourhood depth calculation and overlay address validation
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
Thanks for the well-thought-out SWIP — the design is elegant and clearly addresses Sybil resistance and balanced assignment. I had a few questions and points I’d like to discuss for further clarity and robustness:
|
|
Could _upgradeDepth() become too expensive to execute as the number of assigned nodes grows? The _upgradeDepth() function doubles the assignment and remaining lists, copies all existing nodes to their new positions using bitwise logic, and clears/rebuilds state — all in a single call. If the number of nodes reaches high volumes (e.g. 1,000+ or 10,000+), this could approach or exceed the block gas limit, making the function fail or stall the system. Proposed solutions: |
|
Can the committers array become inefficient with a large number of registrants (e.g. 20k nodes)? The committers[] array is iterated over in _expire(), _findEntryFor(), and _removeCommitter() using for loops. If a large number of nodes register, or if expired entries are not promptly cleared, the gas cost of these operations can grow linearly and become prohibitively expensive. Proposed solutions: |
I did not consider it realistic, since each registrant entry expires in max 256 blocks, that is in a matter of <4 game rounds and they lose their deposit if they refuse to pay, so likely all the potential players may organically wait out.
but they need to be removed at some point.... and I am not sure how a mapping that needs to be reindexed after every entry removed, will solve this. |
well, maybe. To be honest, there is also another way. We do not need to allow, just any length of the committer list. The length represents the queue, and the length of valid entries are the ones in the queue you can skip. This effectively quantifies the tries that you got (effectively mining) but also the realistic probability that that someone will come in and change the neighbourhood you (thought you were) assigned to. If this probability is high (there is a lot of nodes that can submit mined overlays), then it can easily happen, that whenever an assigned neighbourhood is read off, nodes will frontrun. So it would just make sense to limit this skip queue to a fix constant number. But this means that the committers list should effectively have a limited length. Now if we siply reject registrations beyond this limit, then the shorter this length, the harder it is for the same amount of currently aspiring nodes to commit. Now in order to avoid that the registration tx needs to be continuously retried (due to it most likely be frontrun by competing resistrants), we should introduce another proper FIFO queue (that is unlimited but does not need iteration). In this case the validity period starts when you enter the limited queue.
Not sure I get how these structures would be useful: index needs reindexing or keeps inactive entries, head pointer just delays the problem and so does the inactive flag. |
i. the mining step is just offloading computation rather than POW, strategic placement is prevented by random allocation, economic disincentives to be quantified forwith |
agree with this, some discussion around implementing binary trie or similar datastructure which will ensure uniform gas usage while providing for the necessary functionality |
|
very good swip, a few thoughts for discussion and expansion in the doc:
|
|
First: the PR was heavily restructured on 2026-07-27 (two-transaction join, no 1. Why do we need on-chain topology? Because assignment must be litigable. The redistribution game checks, at claim 2. Sorted ring + sparse buckets — formalize and compare. Happy to add a comparison subsection to the SWIP; here is the summary. The layout
Where the ring genuinely wins: iterating a neighbourhood's members in overlay order 3. High-level operations. Now in the SWIP: join = 4. Random/balanced assignment without A and R. Reading "A and R" as the ordinal mapping and the reservation from the old draft: 5. Comparison with compacted binary trie. A compacted (path-compressed) trie saves storage when keys are sparse and clustered 6. Depth transitions. Walked through in the SWIP (§Balance invariant: preservation under 7. Complexity / gas. §Gas and performance analysis: selection O(d) reads, activation/departure O(d) 8. The earlier notes (shrink, queue, gas of updates).
|
Promised in the PR #74 review reply: a comparison subsection in the implementation notes. Sorted ring + sparse buckets rebuilds the ICBT once selection counts are added and hits an O(N) re-keying cliff at depth transitions; path compression buys nothing on a tree that the invariant keeps dense by construction. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
@significance Dan — your Aug/Sep comments all predate the 2026-07-27 restructure. Several of them are now addressed in the text: the binary trie you endorsed is in as the ICBT (implicit complete binary trie, with a layout-comparison subsection), de/registration runs through commit queues with bounded expiry, and there is now a proper Terminology section separating these tree depths from storage depth. Still open from your list: activity/liveness tracking (squat-attack eviction), quantified economic disincentives (explicitly deferred to deployment/staking spec), key decoupling, and onboarding adherence proofs. Please re-read the current version and leave a proper GitHub review (approve / request changes with the open points) rather than comments — it would help move this toward Accepted. |
…ts as its two readings Replace the stored splitCount/donorCount pair with a single leafCount n(i). Split and donor counts are complementary within a level (they sum to the depth-d slot count 2^(d-l)) and are read off n(i) and the current d. Spell out counter maintenance as pseudo code: every join, direct departure and donor draw is one ±1 root walk; a completed relocation writes no counter. Show that depth transitions cost no writes, and why the leaf count rather than a split count reduced modulo 2^(d-l) is stored (the residue cannot tell an all-leaves subtree from an all-pairs one). Pending-departure exclusion is applied by rejection at selection, marked (?) for review. Also fix the donor's target: the departing prefix itself, not its sibling, matching the worked example. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
There was a problem hiding this comment.
🟡 Changes recommended
The SWIP text still contains an unresolved (?) ambiguity and it contradicts the PR description about appended/generated Solidity material.
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review details
- Files reviewed: 1/1 changed files
- Comments generated: 2
- Review effort level: Lite
| The implementation SHOULD also generate random valid sequences of joins, direct departures, donor relocations, expiries, and redraws, comparing contract state against a simple off-chain model after each completed transition. | ||
|
|
||
| The exact Solidity ABI, client API paths, economic parameters, and deployment addresses remain to be supplied before this SWIP can advance beyond Draft. |
|
|
||
| The four removal cases map onto these as follows: case 1 clears the root record; cases 2 and 3 are `removeLeaf` of the departing leaf (at depth $d$ and $d+1$ respectively — in both the sibling is an active leaf); case 4 is `removeLeaf` of the drawn donor at draw time, followed by `replaceLeaf` at the departing prefix when the donor activates. So a join, a direct departure, and a donor draw each write one root path of at most $d+2$ counters, and a completed relocation writes none. There is no other write to the tree. | ||
|
|
||
| The eligibility rule that a leaf with a pending departure is not a split candidate is applied at selection, not in the counter: `targetPrefix` treats a descent that lands on a pending leaf as a rejection and re-samples with the next rank drawn from $\rho$. This is exact rejection sampling over the non-pending candidates, and it degrades no worse than an explicit exclusion would — if every candidate is pending, neither yields a target until a relocation completes. **(?)** |
|
it was discussed that the scheme is at direct risk of a secondary market arising if there is some value to acquiring contiguous operating addresses which in the current swarm architecture there indeed is it was discussed that the possibility of a secondary market becoming established could be to some extent mitigated by forcing an EOA account which cannot be changed and is allowed to withdraw stake in a way that cannot be prevented it is further emphasised that neutering the ability should be explicitly avoided (and a comment retained in the relevant parts of the code) despite this approach necessitating undesirable ux as it requires forcing unstaking/restaking and reallocation of overlay address location if this withdrawal address needed to be changed eg. in the event of compromise further note: the withdrawal address must be proven owned to satisfy this requirement - i.e. cannot be an obviously fake address eg. the zero address |
|
relevant.... Forcing EOA-only calls in SolidityThe checkrequire(msg.sender == tx.origin && msg.sender.code.length == 0, "EOA only");
Either clause alone is insufficient: Risks
|
|
@zelig leaving this here so it can be considered while i continue the review - the extreme manifestation of this is when viewed from the perspective of the last unmarried sister in the depth - for them it is a no brainer to stake the contiguous slot presuming there is economic benefit from doing so as there is in eg. the redistribution game. the obvious solution leads to a trade off of eroded efficiency in depth saturation, which can be mitigated by redundancy/shared responsibility, but this now becomes an essential constituent and hence should form part of this specification. perhaps there is a better way? The problem (at it's most extreme) is that when a depth |
|
I'm writing my first batch of feedback here rather than commenting on specific paragraphs in code review. These comments are all about high-level motivation rather than implementation. The only concrete motivation given is to make it more costly to execute what I will call a deduplication attack, i.e. running two node identities on the same underlying storage resources, masquerading it as replicated storage. To accept this as motivation, we need an assessment of the likelihood and harm of this type of attack. (The motivation in terms of DSNs is not concrete because it doesn't name a specific issue that Swarm faces today; see below.) A second possible motivation that we discussed on our 1-1 call is to make it harder to Sybil the consensus over the reserve. However, since it didn't seem like that was your primary focus, and you haven't mentioned it in the SWIP document, I won't address that point in this review. Harm. How harmful is the attack? Basically, a node that represents itself as multiple nodes does two things:
Status quo. The current system defends against the deduplication attack via the elective (self-selected) stake system: if you want to scale your revenue in a single neighbourhood, it is cheaper and simpler to do so by scaling stake alone instead of deploying a full Sybil identity. This shrinks the population of entities who might attempt this attack, i.e. it makes it less likely. Given the elective stake system in place, it is not clear to me that there is any urgent need for further defense against the deduplication attack; especially when the measure involves such a dramatic overhaul of the current system. Under centrally quoted stake. If we take it as given that we want to do away with the elective stake system, replacing it with a mechanism where the system tells you how much you must stake — and I suspect this is among your assumptions, even though it is not stated explicitly in the SWIP — then we may need a new mitigation. Long term roadmap. I think we both believe that in the long term, Swarm ought to move towards a ~1 average replication rate and rely on erasure coding for fault tolerance. In that régime, it's not clear to me that a deduplication "attack" is actually harmful:
In summary, I find the motivation plausible in the medium term (before Swarm does away with replication as the primary means to underwrite durability), conditional on getting rid of elective stake. This should probably be clarified in the Scope section. Recommendations
|
|
|
||
| The mechanism is designed to achieve: | ||
|
|
||
| - complete coverage of the address space by disjoint areas of responsibility; |
There was a problem hiding this comment.
The use of "areas of responsibility" is a bit confusing since the proposal doesn't specify anything about the responsibility of the nodes. Is the intention that nodes continue determine their AoR as they do today, i.e. using the depth that gives them between 2^21 and 2^22 chunks? If so, what does this proposal have to do with how those AoRs cover the address space?
OTOH if this is a statement about the set of balls defined by leaf prefixes (which are not generally the AoRs of the assigned nodes unless the method of determining AoR has changed), you are asking that these balls partition the address space. Why is that a goal?
| The mechanism is designed to achieve: | ||
|
|
||
| - complete coverage of the address space by disjoint areas of responsibility; | ||
| - a maximum factor of two between the largest and smallest areas; |
|
|
||
| - complete coverage of the address space by disjoint areas of responsibility; | ||
| - a maximum factor of two between the largest and smallest areas; | ||
| - uniformly random selection among currently eligible split candidates or donor pairs; |
There was a problem hiding this comment.
This was hard to understand on a first reading because you are using jargon from the construction. It reads as though the author defined the construction first and repackaged part of the construction as a goal, which it tautologically achieves.
The condition can be expressed in higher-level language, e.g. "Newly joining nodes are assigned prefixes uniformly at random from a set of available prefixes," which at least aids comprehensibility (though it is still unclear why this is a goal).
| - complete coverage of the address space by disjoint areas of responsibility; | ||
| - a maximum factor of two between the largest and smallest areas; | ||
| - uniformly random selection among currently eligible split candidates or donor pairs; | ||
| - assignment that is unknown to an applicant when it commits stake and fees; |
There was a problem hiding this comment.
If we accept that the main objective of this proposal is to centrally assign prefixes, it's easy to see why you want this.
This proposal does not actually achieve this requirement in all cases: if there are exactly
| - a maximum factor of two between the largest and smallest areas; | ||
| - uniformly random selection among currently eligible split candidates or donor pairs; | ||
| - assignment that is unknown to an applicant when it commits stake and fees; | ||
| - a bounded, explicit procedure for failed joins and failed donor relocations; |
There was a problem hiding this comment.
Again, jargon. Suggest "a bounded procedure for handling nodes that fail to take up their assigned prefix."
| - assignment that is unknown to an applicant when it commits stake and fees; | ||
| - a bounded, explicit procedure for failed joins and failed donor relocations; | ||
| - logarithmic contract work in the number of active nodes for selection and tree updates; and | ||
| - deterministic validation of assignments by the contract and swarm node clients. |
There was a problem hiding this comment.
What exactly is being validated here and by whom? That a given prefix is assigned to a node? This sounds like a very weak condition, why is this being mentioned?
|
|
||
| The registry associates each active node with one leaf prefix, and the node's overlay MUST fall within the neighbourhood designated by that prefix. | ||
|
|
||
| ### Balance invariant |
There was a problem hiding this comment.
In this section, you appear to be defining a property of a function
|
|
||
| Because a depth-$d$ neighbourhood has twice the address-space volume of a depth-$(d+1)$ neighbourhood, the largest area of responsibility is at most twice the smallest. | ||
|
|
||
| #### Preservation under insertion |
There was a problem hiding this comment.
Here you are defining a procedure to construct a balanced
| If $N\geq1$, a join selects one split candidate $p$ at depth $d$. Let the incumbent overlay's next bit be $b=o_{\mathrm{inc}}[d]$. The incumbent is reassigned to $p\mathbin\Vert b$, and the joining node is assigned: | ||
|
|
||
| $$ | ||
| p_{\mathrm{new}}=p\mathbin\Vert(1-b). |
There was a problem hiding this comment.
Suggest using a more compact notatation for the complement of a byte, e.g.
|
|
||
| Replacing $p$ by its two children preserves disjointness and full coverage. When the insertion increases $N$ from $2^{d+1}-1$ to $2^{d+1}$, all leaves are at depth $d+1$, which becomes the new minimum depth. | ||
|
|
||
| #### Preservation under removal |
There was a problem hiding this comment.
Defining balanced
|
|
||
| **Active node count** — $N$, the number of activated nodes in the balanced partition. A pending registration is not active. | ||
|
|
||
| ## Motivation |
There was a problem hiding this comment.
Suggest moving this section to the top, after Abstract.
awmacpherson
left a comment
There was a problem hiding this comment.
I write this review under the assumption that the high level goal is to ensure that the number of nodes serving each point in the address space is "as close to constant as possible." The proposal doesn't specify whether or how nodes' responsibilities would change, so I'll assume they continue to reckon their AoR as they do today based on the overlay address they mine and the number of chunks they can find in proximity to their address. I will refer to the view of depth stored by each node based on this count as the reserve depth, and the
The healthy situation is that reserve depth is less than registry depth. In this case, the proposal guarantees that the number of prefixes assigned in two distinct neighbourhoods (at the reserve depth) differs by at most a factor of two. (Of course, it doesn't guarantee that assigned prefixes correspond to currently active nodes, but we can likely get fairly high confidence of that by tuning incentives.) I don't know whether the distribution of node count in each neighbourhood at a given
|
|
||
| Any number of departures MAY pend concurrently: each records its own donor and relocation deadline, and a stalled donor delays only its own departure. Concurrent draws never collide, because the donor's sibling takeover is applied to the tree at draw time, so every draw selects from the currently remaining pairs. A leaf whose departure is pending is excluded from split-candidate eligibility, so a join cannot split a neighbourhood that is about to be vacated. | ||
|
|
||
| #### Data handover |
There was a problem hiding this comment.
I'm sure this has already been raised, but a major concern is churn triggered by node exits that require a node to be transplanted from another prefix. The cost of exiting would have to be tuned to reflect the costs to the donor node (who would likely stop earning while they sync the new neighbourhood) and the externality on the network. This would likely make it very expensive to exit large numbers of nodes, which in turn makes it less attractive to operator fleets at scale.
Exiting large numbers of nodes would cause more harm to the network under SWIP-39 than it does today, because as well as the loss of replicas you are also likely triggering large numbers of overlay changes at the same time.
UPDATED AUGUST 2026
First: the PR was heavily restructured on 2026-07-27 (two-transaction join, no
reservation held, identity-keyed registrations, concurrent departures, ICBT as a
separate container contract) — several of the points below are answered in the text
now, so please re-read the current version before diving into the old one.
Point-by-point:
1. Why do we need on-chain topology?
Because assignment must be litigable. The redistribution game checks, at claim
time, that a node plays in the neighbourhood it was assigned — that check has to run
in the contract, against state the contract trusts. Off-chain assignment can be
neither enforced (nothing stops a node self-selecting) nor verified (the contract
has no source of truth to check an overlay against). The same goes for the two
enforcement events: forfeiting an expired registration and slashing a defaulted
donor — both need the assignment state on-chain to be adjudicable. And grind-proof
randomness (stake locked before entropy known, seed from a committed block height)
only means anything if the commit itself is on-chain. What stays off-chain is
everything that can: target-prefix computation is a read-only call, mining is local,
and the join costs exactly two transactions.
2. Sorted ring + sparse buckets — formalize and compare.
Happy to add a comparison subsection to the SWIP; here is the summary. The layout
(doubly-linked list in overlay order +
mapping(prefix at staking depth => [first node, count])) is good at what rings are good at: O(1)-writeinsertion/removal once the position is known, O(1) neighbourhood membership query,
cheap ordered iteration. It is weak exactly where this SWIP lives:
layout either scans buckets — O(2^d) — or maintains hierarchical per-subtree
counts to support O(log N) rank selection. The moment you add the counts
hierarchy (your "buckets hierarchy?" bullet concedes this), you have rebuilt the
ICBT: the trie is the counting structure, with the ring's information implicit
in it.
when
lazy-migration scheme with its own bookkeeping. The ICBT never re-keys:
derived, indexes are stable, a depth transition is zero writes.
(path to root, ~30 slots at a million nodes). Ring: O(1) link writes + O(log N)
anyway for whatever counting structure supports selection. We are comparing
log-vs-log; the constant matters less than the re-keying cliff above.
Where the ring genuinely wins: iterating a neighbourhood's members in overlay order
(we never need this — one node per leaf by construction) and finding the successor
of an arbitrary address (the ICBT does it in O(log N) via
nodeFor, good enough fora view function). So: formal comparison in the SWIP yes, layout change no.
3. High-level operations.
Now in the SWIP: join =
register+activate(§Join protocol), departure =deregister(+ donor redraw path) (§Departure and rebalancing), neighbourhoodqueries =
getPrefix/nodeFor(read-only). Redistribution eligibility is oneprefix check against the assignment record — the staking contract calls
getPrefix(identity)and compares against the claimed neighbourhood. If a specificoperation list is wanted verbatim in the issue's terms, point me at it.
4. Random/balanced assignment without A and R.
Reading "A and R" as the ordinal mapping and the reservation from the old draft:
the restructured version already dropped both. There is no reservation
(target prefix is a read-only computation, nothing locked, activation revalidates
against current state) and no request IDs / ordinal indirection (registrations are
keyed by staking identity; the seed is H(domain ‖ identity ‖ blockhash)). If A and R
meant something else, tell me what and I'll answer that instead.
5. Comparison with compacted binary trie.
A compacted (path-compressed) trie saves storage when keys are sparse and clustered$d$ or $d+1$ to be occupied, i.e. the trie is always complete to
— but our key population is dense by construction: the invariant forces every
prefix at depth
within one level. There is nothing to compact — path compression on a complete tree
adds skip-pointers that must be maintained on every split/collapse and saves zero
levels. The implicit heap layout additionally removes all pointer storage: parent,
children, sibling are arithmetic on the index, so a "node" is just its counter
slots. Compaction pays off for arbitrary key sets; balanced assignment is precisely
the regime where it cannot. Will add this as a paragraph to the comparison
subsection.
6. Depth transitions.
Walked through in the SWIP (§Balance invariant: preservation under$N=2^D$ boundary cases for splitCount/donorCount).$d$ , no$2^{d+1}-N$ and $N-2^d$ and both$N=2^D$ example would
insertion/removal; §Counting: the
The short version: transitions are emergent, not an event — no stored
migration, the counters at the root already equal
hit the boundary values exactly at powers of two. If a worked
help, I can add one next to the existing worked examples.
7. Complexity / gas.
§Gas and performance analysis: selection O(d) reads, activation/departure O(d)$2^\ell$ hashes, $\ell \approx \log_2 N + 1$ . Benchmarks are
writes, storage O(N), registration/expiry O(1) amortized (monotonic queue head,
bounded per call — no unbounded iteration anywhere). Mining is the real cost and it
is off-chain: expected
listed as a reference-implementation deliverable; concrete numbers per depth once
there is a contract to measure.
8. The earlier notes (shrink, queue, gas of updates).
completes immediately; the donor path holds the departing node active until a
donor lands, with forfeiture + redraw on default. Withdrawal of stake itself is
the staking contract's business (strict separation in the SWIP).
block heights, head-advancing
expirebounded per call.point 2 for why the proposed alternative doesn't beat it once selection is
accounted for.
this comment was added here #74 (comment)