Skip to content
Submit a case

Solution 03

Publish the filtering methodology and re-screen what stays live

01 — Root cause addressed

Downstream trainers inherited contamination with no way to assess or bound it, and hash lists continue to grow after a dataset is released.

02 — Implementation

Ship a written description of exactly what was screened and how — which lists, which hashing methods, which stages, and what the screening explicitly does not cover — so consumers can judge residual risk instead of assuming none. Then re-run screening periodically against updated lists for as long as the dataset is distributed, and publish the updated counts, since a list that grows after release means a corpus that was screened once is not screened now.

03 — Prevention

Treating methodology disclosure and periodic re-screening as conditions of continued distribution keeps the residual risk visible to consumers and keeps a live dataset current with the lists.

04 — Trade-offs

Publishing methodology also tells a bad actor what the filter does not catch. Periodic re-screening is an ongoing operational commitment rather than a one-off, which volunteer-run projects may not be able to sustain. And it cannot retroactively clean models already trained on an earlier release, so it should not be offered as a remedy for anything already downstream.