Regulation
NIS2 and penetration testing: what the directive actually requires
NIS2 does not contain the phrase penetration testing. It contains something harder: an obligation to assess whether your risk-management measures actually work. Testing is how most organisations produce that evidence — but only if it is scoped against the obligation rather than against a tool.

Almost every conversation we have about NIS2 starts the same way: someone forwards a slide claiming the directive mandates annual penetration testing, and asks for a quote. It does not. Directive (EU) 2022/2555 does not use the term at all. Understanding what it does say is the difference between buying a test that satisfies an auditor and buying one that satisfies a salesperson.
What Article 21 actually obliges you to do
Article 21 requires essential and important entities to take appropriate and proportionate technical, operational and organisational measures to manage the risks posed to the security of network and information systems, and to prevent or minimise the impact of incidents. It then lists a minimum set of areas those measures must cover.
- Risk analysis and information system security policies.
- Incident handling.
- Business continuity, including backup management and crisis management.
- Supply chain security, including security-related aspects of relationships with direct suppliers and service providers.
- Security in network and information systems acquisition, development and maintenance, including vulnerability handling and disclosure.
- Policies and procedures to assess the effectiveness of cybersecurity risk-management measures.
- Basic cyber hygiene practices and cybersecurity training.
- Policies on the use of cryptography and, where appropriate, encryption.
- Human resources security, access control policies and asset management.
- Multi-factor authentication or continuous authentication, secured voice, video and text communications, and secured emergency communication systems.
Two of those items are where testing lives. The vulnerability handling clause obliges you to have a working process for finding and fixing weaknesses. The effectiveness assessment clause obliges you to check whether your measures actually do what your policy claims. Neither prescribes a method — and that is deliberate. The directive is technology-neutral and proportionate, which means a 60-person SaaS company and a national energy operator are not expected to produce the same evidence.
Why management liability changes the buying decision
Article 20 is the clause that changed procurement behaviour across Europe. Management bodies must approve the cybersecurity risk-management measures, oversee their implementation, and can be held liable for infringements. Members of management are also required to follow training, and entities are encouraged to offer similar training to employees.
This matters for testing because it shifts the audience for your report. Before NIS2, a penetration test report was read by an engineering lead who wanted a ticket list. Now a second reader exists: a board member who has to sign an approval and needs to understand residual risk in language that survives a legal review. If your test report cannot serve both readers, you will end up paying a consultant to translate it — and that translation is where accuracy goes to die.
Scoping a test that produces regulatory evidence
The failure mode we see most often is a scope written around an asset list rather than around a claim. An asset-scoped test tells you what was probed. A claim-scoped test tells you whether a stated control holds. Only the second one answers a supervisory authority.
| Article 21 area | Weak scope | Evidence-producing scope |
|---|---|---|
| Access control and MFA | External network scan of the perimeter range | Attempt authenticated privilege escalation and MFA bypass against the identity provider and its recovery flows |
| Vulnerability handling | One-off scan report handed over as a PDF | Test, then retest after remediation, with dated evidence of both states |
| Supply chain security | Out of scope | Test the integrations and delegated access your direct suppliers actually hold in your environment |
| Incident handling | Not assessed | Adversary simulation with an agreed detection window, measuring what your SOC saw and escalated |
| Business continuity | Backup policy document reviewed | Attempt to reach, alter or delete backups from a compromised position |
Backups deserve particular attention. A very large share of the incidents that become reportable under NIS2 become reportable precisely because recovery failed. If your test scope stops at the perimeter, you have not tested the control that determines whether an incident is an inconvenience or a significant incident.
Reporting deadlines shape what you should rehearse
NIS2 sets a staged reporting timeline for significant incidents, and the first stage is short enough that it must be rehearsed rather than improvised.
- Within 24 hours of becoming aware of a significant incident: an early warning to the CSIRT or competent authority, indicating whether the incident is suspected of being caused by unlawful or malicious acts, or could have cross-border impact.
- Within 72 hours: an incident notification updating the early warning, with an initial assessment of severity, impact and indicators of compromise.
- On request: an intermediate report on relevant status updates.
- Within one month of the incident notification: a final report covering a detailed description, the type of threat or root cause, applied and ongoing mitigation measures, and any cross-border impact.
The 24-hour clock starts at awareness, not at containment. That means the deciding factor is usually whether your monitoring produced a defensible awareness timestamp — which is a detection and escalation question, not a legal one. If you want to know whether you can meet it, the honest way to find out is a tabletop or a red-team exercise with an agreed reporting objective, not a policy review.
Where organisations get NIS2 testing wrong
- Treating scope as coverage. A test that reaches 12 of 400 assets tells you very little about the other 388, and a regulator will ask how the 12 were chosen. Document the selection rationale.
- Buying an annual cadence with no trigger-based testing. NIS2 emphasises continuous risk management. Significant architectural change is a stronger trigger for testing than the calendar.
- Excluding the identity provider. It is usually the single highest-impact system in scope and the one most often carved out for operational convenience.
- Never retesting. An unverified fix is an assertion. Article 21 asks for effectiveness, and effectiveness without verification is not evidence.
- Ignoring the supply chain clause because suppliers are 'someone else's risk'. The directive explicitly makes their access your problem.
A proportionate starting point
If you are an important entity with a small security function, a defensible first-year programme looks like this: one external and one authenticated internal assessment covering the identity provider and the systems that carry your most sensitive processing; a documented remediation window; a retest of everything rated high or critical; and one incident exercise that produces a timed awareness-to-notification record. That is a modest budget and it produces evidence against four of the ten Article 21 areas.
If you are an essential entity in a sector with a national supervisory regime, expect a higher bar: broader scope, evidence of a repeatable process rather than a one-off exercise, and testing that includes detection and response rather than vulnerability discovery alone.
Frequently asked questions
Does NIS2 require an annual penetration test?
No. Directive (EU) 2022/2555 does not mention penetration testing and sets no testing frequency. It requires policies and procedures to assess the effectiveness of cybersecurity risk-management measures, and a working vulnerability handling process. Testing is the most common way to evidence both, and some national transpositions or sector regulators do set more specific expectations.
Who is in scope for NIS2?
NIS2 applies to essential and important entities in the sectors listed in its annexes, generally at medium size and above, with some entities in scope regardless of size — including certain DNS providers, TLD registries, trust service providers and public administration entities. Because size and sector thresholds are applied through national law, confirm your classification against your national transposition.
What is the NIS2 incident reporting deadline?
An early warning is due within 24 hours of becoming aware of a significant incident, an incident notification within 72 hours, and a final report within one month of that notification. Intermediate reports may be requested. The 24-hour clock starts at awareness, which makes detection and escalation quality the deciding factor.
Can management be held personally liable under NIS2?
Article 20 requires management bodies to approve and oversee cybersecurity risk-management measures and provides that they can be held liable for infringements. The specific consequences — including any temporary prohibition from management functions for essential entities — are set out in national transposing law.
Talk to the team
Need this tested rather than described?
QuietHours is a European cybersecurity practice operated by the securIT team. Send us the system, service, or control you are concerned about and we will help turn it into a workable scope.
Discuss your scope