Problem
sp_data_set_pruning and abandoned_data_set_sweep are two global, wallet-wide jobs sharing one singleton queue. Each walks the entire wallet's data sets via getPdpDataSets, which also decodes every referenced provider's on-chain PDP offering as part of discovery. On staging this just ran for 10+ minutes without finishing — 190+ data sets skipped across 13 failing 100-item batches — because far more providers have malformed offerings than the single-bad-actor case this was designed around. The recent bisection fix (#703) makes each failure recoverable but doesn't reduce the RPC cost, which scales with how much bad data exists on-chain, and the two jobs' shared singleton means one slow wallet-wide walk blocks the other entirely.
Proposed change
- Make
SP_CLEANUP_QUEUE a per-SP singleton queue, scheduled the same way deal/retrieval/data_set_creation jobs already are — one job per provider instead of two wallet-wide jobs.
- Discover a provider's data sets from dealbot's own subgraph instead of
getPdpDataSets, scoped to that one provider. Avoids the on-chain PDP-offering decode for discovery entirely.
- Termination/deletion logic (provider-relay, on-chain
deleteDataSet, rail settlement) stays as-is — only discovery changes.
Scheduling gap this creates
Per-SP job scheduling (ensureScheduleRows) only creates schedule rows for providers currently in storageProviderRepository.findActiveAddresses(). A provider that fully deregisters drops out of that list, so a per-SP cleanup job would never be scheduled for it again — even though it may still hold a data set that genuinely needs pruning/sweeping. The current wallet-wide walk doesn't have this gap: it enumerates every data set dealbot's wallet holds directly, regardless of the provider's registry status (// Prune every provider dealbot holds data sets with, even if it is no longer registry-active, sp-cleanup.service.ts:286).
Fix: schedule SP_CLEANUP_QUEUE's per-SP jobs from the union of findActiveAddresses() and the distinct provider addresses across dealbot's subgraph-tracked data sets, not from findActiveAddresses() alone. Deal/retrieval/data-set-creation scheduling is unaffected.
What the subgraph already covers
DataSet.isActive (live + managed), pdpPaymentEndEpoch (pdpEndEpoch), fwssServiceProvider, fwssPayer, setId map directly to what pruning/sweep read today.
Gaps to close first
- No
metadata key/value map on DataSet (only a derived withIPFSIndexing boolean) — needed for the SDK's slot-matcher.
- No rail ID (
pdpRailId/cdnRailId) on DataSet — needed for abandoned_data_set_sweep's settlement check.
- Provider
serviceURL for the relay path isn't subgraph data — already available from storageProviderRepository via the existing providers_refresh job.
Problem
sp_data_set_pruningandabandoned_data_set_sweepare two global, wallet-wide jobs sharing one singleton queue. Each walks the entire wallet's data sets viagetPdpDataSets, which also decodes every referenced provider's on-chain PDP offering as part of discovery. On staging this just ran for 10+ minutes without finishing — 190+ data sets skipped across 13 failing 100-item batches — because far more providers have malformed offerings than the single-bad-actor case this was designed around. The recent bisection fix (#703) makes each failure recoverable but doesn't reduce the RPC cost, which scales with how much bad data exists on-chain, and the two jobs' shared singleton means one slow wallet-wide walk blocks the other entirely.Proposed change
SP_CLEANUP_QUEUEa per-SP singleton queue, scheduled the same waydeal/retrieval/data_set_creationjobs already are — one job per provider instead of two wallet-wide jobs.getPdpDataSets, scoped to that one provider. Avoids the on-chain PDP-offering decode for discovery entirely.deleteDataSet, rail settlement) stays as-is — only discovery changes.Scheduling gap this creates
Per-SP job scheduling (
ensureScheduleRows) only creates schedule rows for providers currently instorageProviderRepository.findActiveAddresses(). A provider that fully deregisters drops out of that list, so a per-SP cleanup job would never be scheduled for it again — even though it may still hold a data set that genuinely needs pruning/sweeping. The current wallet-wide walk doesn't have this gap: it enumerates every data set dealbot's wallet holds directly, regardless of the provider's registry status (// Prune every provider dealbot holds data sets with, even if it is no longer registry-active,sp-cleanup.service.ts:286).Fix: schedule
SP_CLEANUP_QUEUE's per-SP jobs from the union offindActiveAddresses()and the distinct provider addresses across dealbot's subgraph-tracked data sets, not fromfindActiveAddresses()alone. Deal/retrieval/data-set-creation scheduling is unaffected.What the subgraph already covers
DataSet.isActive(live + managed),pdpPaymentEndEpoch(pdpEndEpoch),fwssServiceProvider,fwssPayer,setIdmap directly to what pruning/sweep read today.Gaps to close first
metadatakey/value map onDataSet(only a derivedwithIPFSIndexingboolean) — needed for the SDK's slot-matcher.pdpRailId/cdnRailId) onDataSet— needed forabandoned_data_set_sweep's settlement check.serviceURLfor the relay path isn't subgraph data — already available fromstorageProviderRepositoryvia the existingproviders_refreshjob.