Skip to content

Move SP cleanup to per-SP jobs backed by the subgraph #706

Description

@silent-cipher

Problem

sp_data_set_pruning and abandoned_data_set_sweep are two global, wallet-wide jobs sharing one singleton queue. Each walks the entire wallet's data sets via getPdpDataSets, which also decodes every referenced provider's on-chain PDP offering as part of discovery. On staging this just ran for 10+ minutes without finishing — 190+ data sets skipped across 13 failing 100-item batches — because far more providers have malformed offerings than the single-bad-actor case this was designed around. The recent bisection fix (#703) makes each failure recoverable but doesn't reduce the RPC cost, which scales with how much bad data exists on-chain, and the two jobs' shared singleton means one slow wallet-wide walk blocks the other entirely.

Proposed change

  • Make SP_CLEANUP_QUEUE a per-SP singleton queue, scheduled the same way deal/retrieval/data_set_creation jobs already are — one job per provider instead of two wallet-wide jobs.
  • Discover a provider's data sets from dealbot's own subgraph instead of getPdpDataSets, scoped to that one provider. Avoids the on-chain PDP-offering decode for discovery entirely.
  • Termination/deletion logic (provider-relay, on-chain deleteDataSet, rail settlement) stays as-is — only discovery changes.

Scheduling gap this creates

Per-SP job scheduling (ensureScheduleRows) only creates schedule rows for providers currently in storageProviderRepository.findActiveAddresses(). A provider that fully deregisters drops out of that list, so a per-SP cleanup job would never be scheduled for it again — even though it may still hold a data set that genuinely needs pruning/sweeping. The current wallet-wide walk doesn't have this gap: it enumerates every data set dealbot's wallet holds directly, regardless of the provider's registry status (// Prune every provider dealbot holds data sets with, even if it is no longer registry-active, sp-cleanup.service.ts:286).

Fix: schedule SP_CLEANUP_QUEUE's per-SP jobs from the union of findActiveAddresses() and the distinct provider addresses across dealbot's subgraph-tracked data sets, not from findActiveAddresses() alone. Deal/retrieval/data-set-creation scheduling is unaffected.

What the subgraph already covers

DataSet.isActive (live + managed), pdpPaymentEndEpoch (pdpEndEpoch), fwssServiceProvider, fwssPayer, setId map directly to what pruning/sweep read today.

Gaps to close first

  • No metadata key/value map on DataSet (only a derived withIPFSIndexing boolean) — needed for the SDK's slot-matcher.
  • No rail ID (pdpRailId/cdnRailId) on DataSet — needed for abandoned_data_set_sweep's settlement check.
  • Provider serviceURL for the relay path isn't subgraph data — already available from storageProviderRepository via the existing providers_refresh job.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    • Status
      📌 Triage

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions