Solution 03
Publish the filtering methodology and re-screen what stays live
01 — Root cause addressed
Downstream trainers inherited contamination with no way to assess or bound it, and hash lists continue to grow after a dataset is released.
02 — Implementation
Ship a written description of exactly what was screened and how — which lists, which hashing methods, which stages, and what the screening explicitly does not cover — so consumers can judge residual risk instead of assuming none. Then re-run screening periodically against updated lists for as long as the dataset is distributed, and publish the updated counts, since a list that grows after release means a corpus that was screened once is not screened now.
03 — Prevention
Treating methodology disclosure and periodic re-screening as conditions of continued distribution keeps the residual risk visible to consumers and keeps a live dataset current with the lists.
04 — Trade-offs
Publishing methodology also tells a bad actor what the filter does not catch. Periodic re-screening is an ongoing operational commitment rather than a one-off, which volunteer-run projects may not be able to sustain. And it cannot retroactively clean models already trained on an earlier release, so it should not be offered as a remedy for anything already downstream.