
Sanctions Screening Software: A Buyer's Framework Built From Real Enforcement Actions
A buyer's guide to sanctions screening software built from the failure modes that produced 2025's largest AML penalties — Block's misconfigured Bitcoin thresholds, GVA Capital's $216M ownership-resolution failure, and a 169,000-alert backlog — so compliance teams can evaluate vendors against what regulators actually penalize.
On April 10, 2025, the New York Department of Financial Services issued a $40 million consent order against Block, Inc. that contained a finding most sanctions screening vendor pitches cannot accommodate: Block had blockchain analytics tools in place, used two vendors, and still produced a configuration so far outside regulatory expectation that NYDFS stated in the order that "any amount of funds transferred to terrorism-connected wallets is illegal," and that setting alert thresholds above zero without documented risk-based justification "falls short of the regulatory requirement."
Block's system was generating alerts only when a wallet exceeded 1% exposure to terrorism-connected funds and it was blocking transactions only above 10%. Simply put: they had a tool in place, but the configuration was seriously flawed.
Two months later, OFAC issued a $215,988,868 penalty (the statutory maximum) against GVA Capital, a San Francisco venture capital firm. GVA Capital had continued managing a sanctioned Russian oligarch's investment after his OFAC designation in April 2018, working through a nephew it knew to be the sanctioned party's representative on investment matters. The firm had sought a legal opinion.
The opinion concluded, incorrectly as OFAC later established, that the holding structure did not cross the 50% ownership threshold. The screening problem was not identifying a name on a list, but resolving beneficial ownership through a proxy and a layered offshore vehicle to the underlying sanctioned interest. There is no standard name-matching tool that addresses that.
Sanctions screening software identifies, with regulatory specificity, the failure modes that actually produce penalties. They are more useful than any feature comparison because they describe the gap between what a tool does and what compliance requires, a gap that exists even when a tool is deployed, running, and processing transactions every day. Whether the buyer is a bank replacing an incumbent screening vendor, a FinTech building its first programme, or a crypto exchange responding to regulatory pressure, the failure modes are the same and the questions that follow from them are the same.
The Failure Modes Regulators Penalize
The Block and GVA Capital cases between them illustrate three distinct failure modes that the sanctions screening software market consistently underweights.
Configuration failure
The first is configuration failure. Block did not lack a screening capability. It lacked a screening configuration that reflected the regulatory standard. The consent order also identified a second configuration problem: Block had risk-rated transactions with exposure to mixing services — services designed to obscure the origin and destination of funds — as "medium" risk rather than "high," despite NYDFS publishing guidance in April 2022 identifying mixers as an elevated typology that virtual currency licensees should address in their transaction monitoring. Two years elapsed between the guidance publication and the enforcement finding. The tool had the capability to assign risk ratings, but the settings did not reflect what the regulator had published.
Resolution failure
The second is ownership resolution failure. The GVA Capital penalty notice makes clear that no name-matching system would have caught this violation because the mechanism of the violation was not a transaction with a named SDN but the continued management of property in which a sanctioned person retained a beneficial interest, channelled through structures that did not place his name on the face of the investment vehicle.
OFAC's 50 percent rule applies not just to nominal ownership but to beneficial interest, and the investigation found that the sanctioned party retained a property interest through a trust structure that a formalistic ownership check had assessed as below the threshold. A screening tool that does not resolve ownership through holding structures and proxies will not catch this category of violation.
Scalability failure
The third is scalability failure. Block's transaction monitoring alert backlog grew from approximately 18,000 unprocessed alerts in 2018 to over 169,000 by 2020. The consent order found that between February 2021 and September 2022, suspicious activity reports were filed on average 129 days after the triggering alert was first generated and in some cases more than a year later.
The statutory SAR filing window is 30 days from the point at which suspicious activity is identified. The backlog produced by alert volumes outpacing review capacity turned a monitoring programme into a documentation trail of delayed violations. The tool was generating alerts, but the architecture was not designed to manage what it generated.
{{snippets-guide}}
Why Coverage Is the Wrong Starting Metric
The most common organising principle of sanctions screening software evaluations is list coverage. How many lists does the vendor cover? Is OFAC included? The EU consolidated list? UN Security Council designations? Does it cover domestic lists in the buyer's operating markets?
. A tool that does not cover the sanctions lists relevant to a business's regulatory obligations is not fit for purpose. But list coverage is also the metric that did not produce a single named enforcement action in 2025. The Block consent order does not cite insufficient list coverage. The GVA Capital penalty notice does not cite insufficient list coverage. The enforcement actions that generated hundreds of millions of dollars in penalties in the past 12 months cite configuration settings, ownership resolution logic, detection latency, and alert queue management.
This matters because coverage is the metric vendors compete on and publish prominently, which means it is the metric that dominates procurement conversations even though it is the least predictive of whether a tool will prevent the failures that regulators penalise. A tool covering 75 lists with a 1% terrorism-exposure threshold produces the same outcome as the one that generated Block's $40 million fine. A tool covering 50 lists with zero-tolerance threshold configuration, ownership resolution through beneficial interest, and a case management system that prevents alert accumulation is the tool that addresses the actual failure modes.
The procurement implication is direct: lead with the questions that map to enforcement failures, not with the questions vendors have prepared polished answers for. For banks considering switching from an incumbent vendor, this framing is especially relevant. Large established tools with broad list coverage frequently receive scrutiny for exactly these gaps, because configuration flexibility and ownership resolution depth are not what legacy platforms were designed to optimize for.
The Questions That Separate Vendors
On configuration
Ask the vendor to demonstrate the interface for alert threshold configuration. Specifically: what is the default threshold for terrorism-connected wallet exposure, and can it be set to zero tolerance at the category level? What is the default risk rating for mixer exposure, and how does the tool respond when a regulator publishes new typology guidance? Who can change these settings, and what does the change log look like?
A vendor that cannot set terrorism-connected exposure thresholds at or below 1% — or that requires a professional services engagement to change a risk rating — is selling a tool whose configuration gap produced an enforcement action. The answer to look for is a configurable threshold per risk category, with a documented, auditable change log that records what was set, when, and by whom.
On ownership resolution
Ask the vendor what happens when a screened entity is not directly named on a sanctions list but is majority-owned or controlled by a named SDN. How does the system handle the OFAC 50 percent rule for aggregated ownership across multiple sanctioned parties? Does it have beneficial ownership data integrated into the match, or is that a separate product?
Many vendors say name matching is the core capability and ownership resolution is either absent or an upsell. That is not a disqualifying answer if the gap is acknowledged and a documented procedure exists to address it. It is a disqualifying answer if the vendor claims their tool handles the 50 percent rule and cannot demonstrate how it traces ownership through offshore vehicles and proxy structures to the underlying beneficial interest.
On detection latency
Ask the vendor what happens between the moment a new designation is published and the moment a customer in the existing book who matches that designation generates an alert. Is it minutes, hours, or the next scheduled batch run? What is the maximum window between an OFAC list publication and a match notification for a monitored entity?
This is the question that separates real-time architectures from batch systems. A batch system that runs nightly creates a measurable window during which a customer designated at midday can transact until the following morning. For businesses with regulatory obligations that require screening on an ongoing basis, that window is a documented compliance gap. Continuous monitoring that propagates list updates to the monitored portfolio within hours of publication closes this window. A polling system that checks at a fixed interval does not.
On alert volume and case management
Ask the vendor what happens to alerts when the compliance team is at capacity. Does the system queue them with timestamps? How does it prioritise the queue? What is the documented false positive rate for customers with a comparable profile to your business, and what does alert volume look like at two and five times current transaction volume?
Block's backlog did not appear suddenly. It grew over three years because the alert generation rate outpaced the review capacity, and no architectural safeguard prevented accumulation. A vendor that cannot model expected alert volume at growth-stage transaction volumes, or that cannot produce false positive rates from production environments at comparable scale, is selling a tool that may replicate that dynamic. For a practical treatment of how to structure this evaluation as part of a formal vendor selection process, the sanctions.io Vendor Selection Guide covers the full procurement framework.
What the Audit Trail Has to Prove
Across all three failure modes, the connecting thread in the regulatory findings is documentation. The Block consent order's backlog finding is partly about late SAR filings, but it is also about what the system could and could not demonstrate to examiners about when alerts were generated, when they were reviewed, and what the outcome was. The GVA Capital finding rests on what the firm knew and when, a question answered partly by documentary evidence of its communications with a known proxy.
A screening system that runs and generates outputs but does not produce a retrievable, timestamped record of every screening event (what entity was screened, against which list version, at what point in time, with what result, and what the compliance team's disposition was) is not audit-ready regardless of how accurate its matching logic is. The audit trail is not a reporting feature. It is the evidence layer that determines whether a compliance programme can be defended when a regulator asks to see it.
Ask your vendors: can you retrieve a complete screening record for a specific customer, on a specific date, showing the list version active at that time and the outcome? If the answer requires a database query by the vendor's engineering team, the audit infrastructure is not fit for purpose.
{{snippets-case}}
Questions to Take Into Your Next Vendor Demo
These five questions are derived directly from the enforcement findings above:
- Show me the threshold configuration interface for terrorism-connected wallet exposure. Set it to zero and show me the audit log entry that records the change.
- Walk me through what happens when an entity I screened clean at onboarding is added to the OFAC SDN list at 2pm today. At what point does my compliance team receive an alert, and what does that alert contain?
- Show me a production false positive rate for a customer base similar to mine. A rate from a live deployment at comparable transaction volume.
- If my alert queue reaches 500 unreviewed cases, what does the system do? Show me the queue management interface and explain how it prevents accumulation.
- Pull the complete screening record for a test entity. Show me the list version, the timestamp, the match result, and the disposition field. How long does that retrieval take?
A vendor that answers all five with working product demonstrations is addressing the failure modes that produced 2025's largest penalties.
For a broader treatment of the screening concepts underlying these criteria, the Unified Guide to Screening and the Sanctions Screening Guide cover the architecture and data decisions that determine whether a programme is structurally sound before any vendor evaluation begins.
sanctions.io is a highly reliable and cost-effective solution for real-time screening. AI-powered and with an enterprise-grade API with 99.99% uptime are reasons why customers globally trust us with their compliance efforts and sanctions screening needs.
To learn more about how our sanctions, PEP, and criminal watchlist screening service can support your organisation's compliance program: Book a free Discovery Call.
We also encourage you to take advantage of our free 7-day trial to get started with your sanctions and AML screening (no credit card is required).
