Reading lab results

Independent testing is the only way to compare endpoint security products on a level field. This page explains how to interpret the three lab sources Endpoint Index uses, with references to the academic and industry literature that underpins their methodology.


AV-Test protection score

AV-Test evaluates endpoint products every two months using two test sets: a real-world test (zero-day malware collected in the preceding weeks) and a reference set (prevalent malware from the AV-Test database). Each test contributes up to 3 points for a maximum protection score of 6.

A protection score of 6 out of 6 means the product blocked 100% of samples in both rounds. Scores of 5.5 indicate one or two misses across roughly 1,400 zero-day samples and 19,000 reference samples. The scoring methodology is documented in AV-Test's published test plan.

Why it matters

Detection-rate benchmarks remain the primary objective measure of endpoint efficacy. Research on antivirus evasion consistently finds that detection rates vary meaningfully across products, particularly on zero-day samples that rely on behavioural analysis rather than signature matching.

Sources


MITRE ATT&CK analytic coverage

MITRE ATT&CK Evaluations test endpoint products against emulated adversary campaigns (for example, ransomware operators or nation-state actors). The evaluation records whether each attack substep was detected and, if so, at what level: Technique (highest, maps to a specific ATT&CK technique), Tactic (maps to a broader category), General (detected but not classified), None, or Not Applicable.

Endpoint Index reports analytic coverage as the percentage of substeps that received any detection (Technique, Tactic, or General). A product that detects 80 out of 80 substeps achieves 100% analytic coverage. The distribution across levels matters: a product with mostly Technique-level detections provides more actionable information for incident responders.

Why it matters

ATT&CK-based evaluation reflects real-world adversary behaviour better than static malware samples. Research on detection engineering shows that detection coverage mapped to the ATT&CK framework correlates with shorter dwell time and faster incident containment.

Sources


AV-Comparatives false positives

AV-Comparatives tests business endpoint products in a real-world protection test and a malware protection test, each running for several months. A critical dimension is the false positive rate — how many legitimate files, URLs, or applications the product incorrectly blocks.

For businesses, false positives cause workflow disruption: blocked installers, quarantined internal tools, or inaccessible websites. The AV-Comparatives test uses both common business software (where false positives are most disruptive) and other clean files as test sets.

A product with high protection (99% or higher) but many false positives (more than 10) may cause more operational harm than one with slightly lower protection and zero false positives. Endpoint Index displays both metrics side by side so the reader can make this trade-off.

Why it matters

The cost of false positives in small-business environments is well-documented. A single false positive on a line-of-business application can halt operations and erode trust in the security tool, leading to policy overrides that weaken protection.

Sources