Continuous Compliance


But, compliance! Somehow, it has become a welcome justification for superfluous gates in our IT delivery process, and the go-to excuse for embarrassingly long delays. Compliance is one of those severely under-documented topics in the IT industry that triggers massive overcompensation. Contrary to popular belief, compliance is not a good reason to go slow. The research is clear: going slow for safety is a mistake (Forsgren et al., 2018). It causes friction, decelerates feedback, drives down quality, and brings a team to a grinding halt. On the contrary, Compliance is an excellent foundation for better quality. It enables faster delivery, accelerates feedback, and ultimately produces better, more compliant outcomes.


Governance, Risk, and Compliance (GRC)

But what is Compliance, anyway? The industry tends to make an amalgam between Governance, Compliance and Risk Management. Frequently used interchangeably, although they mean different things.

Governance is all the activities an organisation performs to be compliant with external regulations (ECB, EMA, …), certifications (PCI-DSS, ISO 27001, ISAE 3000, …), frameworks (ITIL, COBIT, …), or legally binding contracts and internal regulations (the standards, policies, procedures which are translations of external regulations, certifications or frameworks in internal ways of working). (Humble et al., 2014)

IT governance defines the structure, processes, and mechanisms by which an organisation’s IT activities are directed, monitored, and controlled to achieve business objectives. The IT governance framework is essential for organisations seeking to effectively manage their IT activities, manage risks, ensure compliance, and deliver value to the organisation. (COBIT 5, ISO 38500, Humble et al., 2014)

But Governance involves more than the steps to be compliant. It also involves keeping the organisation on track while balancing the interests of all the organisation’s stakeholders.

Compliance is to obey relevant laws, regulations, legally binding contracts or even cultural norms. When imposed by law, regulation or contract, Compliance is generally not optional. In these cases, not conforming increases the risk of fines, reputational damage or even a business shutdown. (Humble et al., 2014)

Risk Management is one of the Governance activities to support Compliance. Regulations and certain certifications (e.g. ISO 270001) require an organisation to manage its risks. But good leadership also drives the need for Risk Management to build a successful organisation. (Humble et al., 2014)

A Risk is a horrible thing that could happen that can negatively impact an organisation’s goals, reputation, or operations, resulting from ineffective leadership, ethical lapses, inadequate controls, compliance failures, or strategic errors that threaten stakeholder interests. But risks are pervasive. We can never eliminate all risks. Therefore, we need Risk Management to mitigate risks.

A good car has the best brakes.

An organisation that is not compliant, or cheats on compliance, will not make it. We are compliant to be fast and better.

– a bank CIO

Speed, fast feedback, innovation, and Compliance are not a zero-sum game. We can have both. But it requires building Compliance requirements into the delivery process from the start, not at the end. Much like the lean principle of Building Quality In to the product, instead of testing it later. High performance requires Continuous Compliance. It happens all the time.

When it hurts, do it more often.

– Dave Farley, Continuous Delivery

Where do we start?

Continuous Compliance

Compliance starts by understanding the regulation.

Organisations often conflate “their approach to regulation” with regulation. Not the same thing at all. Oftentimes, regulation is about “do we do what we say we do” (ISO 27001). The most rigorous regulation says “get two people to look at it” and “have an audit trail of what happened” (PCI-DSS).

Therefore, teams have to read and understand the regulation. That is of utmost importance. When I say the team, I mean everyone. This is not limited to the Product Manager. That also includes all Engineers. This ensures that product implementations align with the regulation. Because now, Engineers understand the implications of their code. “To my understanding, if we do this, we are good”. From then on, the team is responsible and accountable for implementing and satisfying the compliance requirements. Having that in place already gets us a long way forward towards Continuous Compliance.

Once we understand the regulation, we can extract the compliance requirements, which in turn define controls to be implemented. The controls go onto the product backlog and are prioritised alongside functionality. It receives the necessary priority alongside the required functionality. Be aware that Compliance, much like functionality, has market value. Not being compliant can shut us out of the market.

The other part of Governance, and consequently of product management, is Risk Management. We continuously identify risks: business operations risks as well as IT operations risks and security risks. That is another reason to have the business, i.e. the people doing the business operations, as part of the product team to achieve Continuous Compliance. They are in the best position to identify business risks. Of course, when Risk Management becomes a team activity, naturally Engineers also start identifying business risks as they internalise the business. The same is true for Product Managers and Business Operations, who now start to identify IT operations and security risks.

A predominant Risk Management mode in IT is a “Wouldn’t It Be Horrible“-approach (Hubbard, 1985; Humble et al., 2014). We imagine a particularly catastrophic event occurring. Regardless of its likelihood, we have to avoid it at all costs. There is no sense of prioritisation. The question to answer in managing risks is: “Which risks are we willing to accept and which ones not?”. As we are taking steps to mitigate risk in one area, we inevitably introduce more risks, or new risks, in another area (Humble et al., 2014).

Risk mitigation work does not get a free pass to jump to the front of the line. Instead, we should quantify risks. A battle-tested approach is the classic ISO 27001 Risk Assessment model (Clauses 6.1.2 and 8.2), which calculates the risk value as the product of its impact and likelihood. The team establishes its risk tolerance threshold to determine whether a risk must be treated — via risk reduction (reduce the risk level), risk transfer (transfer the risk to another party) or risk avoidance (avoid the activity resulting in the high risk) — or can be accepted when the risk is below the threshold.

  Likelihood Rare (1) Unlikely (2) Possible (3) Likely (4) Certain (5)
Impact            
Low (1)   1 2 3 4 5
Medium (2)   2 4 6 8 10
High (3)   3 6 9 12 15
Severe (4)   4 8 12 16 20
Critical (5)   5 10 15 20 25

To make this assessment actionable in a product team, we can use Impact Mapping to discover how an operational or security threat impacts business goals, and Cost of Delay to calculate the actual monetary impact (e.g., potential fines, revenue loss, or downtime costs per week if the risk materialises). Quantifying the impact in monetary terms (such as Low: < 50.000 EUR to Critical: > 1.000.000 EUR) allows us to prioritise risk-mitigation controls in the backlog alongside functional features using economic value rather than gut feeling.

Every single day, at the start of the day, the team reviews the risks. Did we identify new risks? What is the impact of the risk when it happens? What is the likelihood of the risk happening? What is the cost to mitigate the risk? Should we mitigate the risk, or can we accept the risk? All of this is documented as Decision Records (much like Architecture Decision Records; yet, this time for non-architecture matters).

As with regulation, identified risks to be mitigated create compliance requirements, which in turn define controls. These controls go once again to the backlog and are prioritised along with all other controls and functionality. Whether to prioritise a feature before a risk mitigation control or a regulation requirement control is, in the end, a business decision and falls under the Product Manager’s authority. Observe that in this case, the Product Manager is also responsible and accountable for business operations, i.e. the turnover and costs generated by the business, as well as its compliance.

This upfront risk assessment will prevent a lot of pain when going to production. By identifying these risks early, we can obtain the most appropriate controls to mitigate risk and comply with the regulation. However, the challenge will be finding the right balance of controls within the scope of the organisation and the applicable regulation, while still allowing teams to deliver value quickly and keeping risks at acceptable levels.

Operating as above is already a considerable step towards Compliance. But it does not stop there.

Get Two People to Look At It

The most demanding regulations require us to get two people to look at it, also commonly known as the “four-eyes principle”. The main reason for this principle is to avoid financial fraud or introduce faults that could kill someone in healthcare. Therefore, have someone different from the code author look at the code.

The classic implementation is formal code reviews via Pull Requests. However, Pull-Requests are one of the biggest bottlenecks in IT delivery.

With PRs we created a world where the amount of time we spend reviewing code has become disproportionate to the time spent creating value.

– Seb Rose during a SoCraTes session, Aug 29, 2022

As discussed before (But Compliance!, The Problem, The Good and the Dysfunctional of Pull Requests, and The Illusion of Compliance), Pull Requests come with the disadvantage of blocking delivery, delaying feedback, consequently decreasing quality and creating an illusion of safety. Additionally, it tends to instigate isolated, solo work, disabling collaboration and diversity of views, ultimately resulting in lower-quality products. While Pull Requests do provide strong evidence that two people looked at the code, they tell us nothing about whether it was truly reviewed — a classic example of the inspection fallacy (Laforgia, 2025).

Furthermore, security and audit teams often put far too much faith in Pull Requests to detect malicious intent or subtle flaws.

Information security, auditors, and regulators often put too much reliance on code reviews to detect fraud. Instead, they should be relying on production monitoring controls in addition to using automated testing, code reviews, and approvals, to effectively mitigate the risks associated with errors and fraud.

[…] we had a developer who planted a backdoor in the code that we deploy to our ATM cash machines. They were able to put the ATMs in maintenance mode at certain times, allowing them to take cash out of the machines. We were able to detect the fraud quite quickly, and it wasn’t through code reviews. These types of backdoors are difficult, or even impossible, to detect when the perpetrators have sufficient means, motive, and opportunity.

– someone leading the DevOps initiative at a large US financial services organisation, The DevOps Handbook, Chapter 23. Protecting the Deployment Pipeline, Case Study: Relying on Production Telemetry for ATM Systems, p.344

The fraud was quickly detected by employing detective controls using proactive production monitoring, not code reviews.

Note: No single regulation mandates Pull Requests. Only PCI-DSS explicitly requires a four-eyes principle. How to implement that is up to us.

A more efficient approach to getting two people to look at it is Pair Programming or Team Programming. Code is reviewed while it is being written, even before hitting mainline, as part of the delivery flow.

However, Pair Programming or Team Programming can be a cultural stretch for teams. In these situations, Non-Blocking Code Reviews is a better choice that works fine for non-regulated industries. With the right tooling, we record who reviewed what. This could even work for regulated industries if we have a process that prevents production deployments unless all code reviews are completed. It comes with the additional advantage that the Release Candidate could already be deployed to a test environment before the code review happened, thus reducing testing delays while still accelerating feedback compared to Pull Requests.

Have an Audit Trail of What Happened

Many compliance requirements are about providing evidence: have an audit trail of what happened. That is where Continuous Delivery enters with its central pattern: the Deployment Pipeline. The Deployment Pipeline, together with Version Control, acts as an audit trail of everything that happened when getting code out of version control into production and into the user’s hands.

The Version Control System tells us what changed, why it changed, when, and who was involved with the change. That could be a single person (isolated programming), two people (Pair Programming) or many people (Team Programming). All of this can be recorded. Either using a combination of authentication and signing keys (see But Compliance! for details), commit messages, or the Co-authored-by Git trailers. The reason why something changed is recorded by including a ticket number in the commit message. Note that the “why” is often overlooked, though it is one of the most important pieces of information, as it provides us with the context and the reason for the change.

The Deployment Pipeline tells us which commit triggered the pipeline run, which actions happened, and when they happened. It links the binary artefact that gets deployed to production to the commit that triggered the pipeline run. For manual tasks, such as Exploratory Testing or triggering the production deployment, it additionally records who performed the manual task.

The pipeline collects all the evidence: linting and code quality scanning results, Unit Test execution results, secrets in plain text detection, Software Composition Analysis (SCA, commonly known as vulnerability scanning), Static Application Security Testing (SAST, commonly known as security scanning), Automated Acceptance Testing, API Security Testing, Dynamic Application Security Testing (DAST), load and performance testing, etc. and finally, health checks and Smoke Testing.

Anytime any of these tests or checks fail, they fail the Deployment Pipeline. Whenever the pipeline fails, the team stops the line, stops all work, and they Do not Push any More Code to the Broken Pipeline; they own the failure and fix the problem with the highest priority. Only once the pipeline is fixed does the team resume any ongoing work and move on. This is paramount to prevent vulnerable software in production.

For this to work, critically, all checks, and especially automated tests, must be deterministic. When they fail, they fail every time. Tests should not fail and pass again on a rerun. These tests are useless.

What about Segregation of Duties?

This is arguably the single most misunderstood requirement in enterprise IT compliance. It is the number one excuse used to decelerate delivery, justify manual handoffs, enforce infrequent release schedules, and bring teams to a standstill. Few regulatory concepts cause so much self-inflicted friction.

It was traditionally intended to prevent financial fraud or faults that could threaten human life by ensuring no single person controls the entire delivery process end-to-end. The rule of thumb was simple: the person authoring the code may neither release nor deploy it. In IT, this is often misinterpreted as “Engineers must not have the ability to deploy in production” or “Engineers may not have production access”. This leads to blocking code reviews with Pull Requests, lengthy Change Approval Boards (CABs), or dedicated operations teams performing deployments based on engineers’ instructions, along with inevitable back-and-forth. Consequently, it hinders feedback loops and cuts quality. In the end, the things we put in place to supposedly control quality do the exact opposite: they limit quality.

Pair and Team Programming already provide continuous peer oversight against unapproved changes, ensuring no single person writes and accepts the code.

Furthermore, when compliance controls mandate that all production deployments (including emergency interventions) must pass through a single path — the Deployment Pipeline — and that no single human can bypass the pipeline’s quality and security gates, the pipeline itself acts as the independent segregation mechanism. For this to work, vitally, the Deployment Pipeline must be repeatable, reliable, consistent, and deterministic.

The pair or team writes and reviews the code, and the deterministic pipeline independently tests, verifies and deploys the code.

Engineers do, however, need access to production, especially telemetry, to investigate outages. This access must be limited to read-only permissions to satisfy the non-repudiation requirements (more on this below in There is More) while keeping teams empowered to support their systems.

There Is More

The above is already a good start. It is the minimum to make a substantial headway. However, there is more … we have to consider Continuous Safeguards.

Compliance is not a point-in-time activity that happens before a release. Vulnerability and plain-text secrets scanning, as well as DAST, must run continuously across repositories and applications because new CVEs are discovered in existing code every day, following a release.

Runtime workloads and pipelines use secrets. This requires secret stores. However, we also need a process to provision secret values into those secret stores using Infrastructure as Code. Either we store secrets encrypted in version control using tools such as SOPS, or Infrastructure as Code connects directly to a password manager that acts as the secrets source of truth.

Audit trails require non-repudiation. Therefore, we eliminate any manual production modifications and revoke write permissions from staff members. Following the Principle of Least Privilege, staff members — including administrators — are granted read-only access. This protects against scenarios where an administrator accidentally deletes a critical infrastructure resource during routine maintenance, taking down all production traffic (a true story). Crucially, it also ensures that all cloud infrastructure changes have to be performed by a Deployment Pipeline using Infrastructure as Code, making sure infrastructure changes are traceable. Only the Deployment Pipeline receives write permissions. To further limit the blast radius, we revoke delete permissions from the Deployment Pipeline by default. We only grant it when a resource is explicitly intended to be destroyed. This practice has saved us lately from accidentally destroying a data store, despite going through multiple code reviews. Finally, following the break-the-glass principle, administrators retain identity access management permissions to elevate their permissions strictly during emergency situations.

Non-repudiation also requires immutable infrastructure and ephemeral environments. This is standard when patching cloud-native workloads such as Docker containers and Lambda functions. Similarly, virtual machine patching is executed using virtual machine images; in that case, virtual machines are replaced instead of modified at runtime. Another advantage is that it eliminates configuration-drift.

When altering permissions, especially when breaking the glass, we need a trace of those changes. Therefore, we need a cloud audit trail to track all activities on the cloud platform. The audit trail tells us who (which Deployment Pipeline, which workload, which staff member) did what, and when. It does not tell us the why. Additionally, the audit trail must alert us of risky activities, such as permission changes or network modifications.

To further shorten feedback loops and shift compliance left, we leverage automation with Compliance-as-Code or Policy-as-Code. Compliance controls are defined as code and applied by Deployment Pipelines. A compliance control can be used to detect, prevent, or correct violations (Morris, 2025). Typically, Policy-as-Code is used to comply with standards such as the NIST 800-53 framework or Center for Internet Security (CIS). Detection controls report when violations occur. This is typically security monitoring of workloads and infrastructure. Prevention controls disallow non-compliant actions. This is generally the checks and tests performed by the Deployment Pipeline that stop the pipeline (such as SCA, SAST, DAST, Automated Acceptance Tests, …). Corrective controls combine detection with an automated corrective action, such as removing unauthorised user accounts.

As mentioned in the “Case Study: Relying on Production Telemetry for ATM Systems”, Continuous Compliance relies as much on continuous production monitoring and alerting (detective controls) as it does on pre-deployment checks (preventive controls).

Lastly, to keep Mean Time to Recover under control, it is vital to have standard responses in place for anticipated situations with runbooks.

Conclusion

Compliance often creates discomfort within teams. Often, the malaise happens because the team lacks Compliance ownership; instead, another team or division tells them what to do, often the Risk & Compliance division. They are kind of at their mercy. Therefore, start reading the regulation to establish constructive dialogue with the Risk & Compliance division. Bring the auditor into the loop and explain the processes and compliance controls the team puts in place, including continuous risk management, Deployment Pipelines, monitoring, and scanning.

In the end, Compliance is just putting in place the product management and engineering practices that enable high performance.

A common mistake by regulated organisations is applying a one-size-fits-all approach to regulation and enforcing compliance controls to every single part of the organisation. A better approach is to confine the regulatory requirements to the regulated activities (see PCI-DSS and continuous deployment at Etsy and “Case Study: PCI Compliance and a Cautionary Tale of Separating Duties at Etsy” from The DevOps Handbook p339).

The industry needs this exact mindset shift. It is a catalyst for speed rather than a brake pedal!

Acknowledgement

Elizabeth Zagroba for the awesome notes during the Continuous Compliance session at SoCraTes 2026.

References