· Platform Engineering · 19 min read
The Drift That Haunts Your BTP, from SAP Inside Track Sydney 2026
The slides and the talk from my session on drift in SAP BTP.

I gave this talk at SAP Inside Track Sydney 2026 on Friday 4 September and a few people asked for the slides. Here they are, with the talk written out alongside. It is based on my earlier article The Drift That Haunts Your BTP.
Who am I

I am John Patterson, Principal Consultant with Second Phase Solutions. Almost 30 years in SAP, ex-SAP Mentor, and the last 10 years focused on BTP.
Often I have been called in when customers are having issues. It is either a war room or a bridge call, and something has gone wrong. The rooms look pretty similar. Developers, operations, architects, outsourced consulting companies and vendors, each looking into their own tools to work out where the issue is.
Coming in cold gives me a different vantage point. The others have lived the solution and, like a boiling frog, have not noticed things getting worse.

The question is why systems that were well designed, with good intentions, fall apart. Usually it is some combination of debt, decay and drift.
Technical debt is the trade-offs we make, the corners we cut, all with the good intention of coming back later to fix them.
Decay is the skills we lose, through inertia and atrophy. The disaster recovery runbook nobody has exercised in three years. The backups nobody has actually tested a restore from. The Node dependency checks nobody reads.
This talk focuses on drift. It looks at why systems drift, the hands that touch the system, six things we try and the blind spot in each, what drift costs, and how we might address it.
Change as a way of life

Picture New Year’s Day. You decide to save money or get fit. You open a new Excel spreadsheet and build a plan, a budget, a diet, a training schedule. Day 0 is perfect. Day 1 is in motion.
By Day 100 the world has moved. The cost of living rises, groceries cost more, petrol goes up, interest rates shift, you get injured, your gym relocates. Your intention did not change. The environment did.
Systems behave the same way. On Day 0 and Day 1 of go live we design well, build well, document well and hand over well. From Day 100 to Day 1000 the designed solution and the running solution diverge.
Practical drift

The first thought when things drift is that it’s a discipline problem. Somebody didn’t follow the process.
But Scott Snook, in Friendly Fire (2000), describes something else. He calls it practical drift, “the slow, steady uncoupling of local practice from written procedure.” He was studying a friendly-fire shootdown, not a cloud platform. But the shape is the same.
Highly skilled, disciplined people make local calls.
Picture the start of UAT. Something has broken. You’ve got a room full of business people who’ve come interstate to test, and an executive telling you to fix it now. The procedure says raise a change, get it approved, and move it through the landscape.
You fix it.
That’s a perfectly rational local decision. But it’s globally misaligned because the fix never makes it back into the procedure.
Governance is only as good as its execution and adaptation.
Circumstances change, and people adapt. When governance doesn’t keep pace, exceptions become local norms, and drift becomes embedded in everyday practice rather than simply reflecting a failure of discipline.
Drift

Go back to the New Year’s resolution. On Day 100 the budget spreadsheet still looks good. It just doesn’t reflect your bank balance and savings.
Our systems are the same. The architecture diagrams and runbooks still look good. But what we declared in those documents, the intention, isn’t what’s running. Drift is the distance between them.
The fortress and the river

It’s easy to get nostalgic and think what we’re dealing with now is unique. It isn’t. Managing change has always been part of the job. What’s changed is how much of it we control.
The fortress was our on-premise system. We built the server room or the data centre. We raised the floor, cabled the room and made it icy cold. We procured the racks and servers and configured them just the way we wanted. We wrote the runbook. Then we put a big padlock on the door saying, “None shall pass.”
Inertia was our friend. We had quarterly change windows and forms signed in triplicate to get anything into production.
We left the fortress deliberately, drawn by the promise of lower capital costs and the ability to deliver more features faster.
I liken BTP to setting up a living room on a cruise ship. You’ve got that nice widescreen TV, the comfortable couches and everything arranged the way you like. You go to bed, and when you wake up the next day, you’re in another port. You don’t control the weather, you don’t control the speed, and you certainly don’t control the ship’s maintenance schedule.
SAP is going to ship platform updates on a Tuesday afternoon. They aren’t going to call you and ask, “Does this patch fit nicely into your Friday night change window?” Their environment moves because a cloud ecosystem has to evolve to survive.
We remember stability

The old world was not stable. It was constrained. We used to build massive dams to hold back the river. Change happened on our timeline. We had design reviews, architecture reviews and change advisory boards. Everything needed approval and everything took time. We still have them. The difference is the cloud does not wait for them. The scope was often narrow enough that a single person could hold the system in their head and the documents could stay accurate for a year.
We remember stability. What we had was visibility.
The hands of change

First, we’ve got to recognise that code is not our system boundary. Take a CAP application. Your code is the small box in the middle. Around it sit HANA Cloud, object storage, destinations, service instances, binding credentials, certificates, identity and runtime. The code you wrote is only a small part of the solution the business relies on. You own the outcome of all of it.
Second, there isn’t just one environment. Typically, you have sandbox, development, test, training, pre-production and production. Each serves a different purpose, with different SLAs, objectives, stakeholders and owners.
One design, many realities. We have to keep those realities aligned so that what we test gives us confidence in what we run in production. But each has its own owners, priorities and pressures, and different hands are changing different parts on different timelines.
Now, the hands of change.
SAP and other cloud providers update services and runtimes on their schedule, not yours.
Developers adjust bindings and credentials because the deadline is Friday and governance is Monday.
Operations and Support fix production at three in the morning. They resolve the outage and save the weekend, but the documentation no longer reflects the system they leave behind.
Security and Compliance change what correct means. A zero-day vulnerability hits the news, the standard changes by the end of the day, and your architecture becomes non-compliant while standing perfectly still. Nothing was deployed and nothing was touched. The definition moved.
Time is the most insidious hand of change. Lock everyone out and certificates will still expire, credentials will become invalid, dependencies will be deprecated and applications will break. The clock does not care that nobody is home.
None of this requires anyone to have done anything wrong. Each team can be doing its job correctly, on a different timeline.
Two kinds of change

There are two kinds of change. The ones we decide to make, and the ones we receive. Initiated change is governed. Received change has no shared owner.
For fifty years we have refined how we manage the first kind. Waterfall gave us requirements, design, build, test, release and sign-off. Agile gave us backlogs, sprints, standups, reviews, retrospectives and a definition of done. Then pipelines, continuous integration and automated deployments. All of it answering one question. How do we govern the changes we decide to make?
Received change has signals too. SAP publishes release notes. Security raises alerts. Operations monitors availability. Developers watch deployments. But nobody joins those signals together. No board reviews the impact across the system, and no ticket tracks the response from one environment to the next. There is no shared definition of done for dealing with an expired certificate, a deprecated service plan or a policy that changed while we slept.
Initiated change is managed. Received change is fragmented. The signals exist, but too often it takes a consequence to bring them together. Something has to hurt before anybody looks across the whole system.
Each team responds locally, but nothing carries those decisions back into a shared model. So teams turn to tools and build their own ways of knowing what is true.
Each approach gives us something useful, but each also has a blind spot. The risk is treating the part it shows us as the whole truth. Six approaches, three layers, six blind spots.
The action layer

The first layer is the action layer.
ClickOps is where someone goes into the BTP cockpit and makes a manual change. Often, it is the right change to make. But the person holds the context, and no artefact records it. Why that certificate is different. Why the legacy destination is still there. Which environment has the special case and why nobody removed it.
The blind spot is invisible drift. Jimmy makes a change in the cockpit. It works and everyone is happy. Then Jimmy goes on leave, or leaves the company. The change stays, but the knowledge walks out the door with him.
Scripts take the steps out of Jimmy’s head and put them somewhere the rest of us can see. They fail predictably. The script stops on line 42 because the API it calls is no longer supported. An engineer can open it, see what happened and fix it.
The blind spot is zombie resources. The script created five things before it hit line 42. It stopped, but those resources stayed. Without cleanup or a record of what was created, they can keep running and costing money long after the failed attempt.
The declaration layer

Next is the declaration layer.
Terraform is declarative. With scripts, we write the steps to perform. With Terraform, we declare what should exist, and it works out the changes needed to get there. We can manage that intent under version control and use it as our source of truth across environments.
The blind spot is ghost resources. Someone makes a change at three in the morning to restore service, is back in bed by four, and forgets to update the declaration. The resource exists, but Terraform has no idea it is there.
GitOps uses the declaration held in Git as the source of truth and continuously reconciles the environment against it. Changes made outside Git can be undone as it brings the environment back to what was declared.
The blind spot is the vengeful automaton. GitOps is like a robot vacuum cleaner. You told it to clean the floor, but you didn’t tell it the kids have been doing a jumbo jigsaw puzzle on that floor. It vacuums up every piece, doing exactly what you told it to do.
The observation layer

Next is observability.
A control plane, like Crossplane in Kyma, can manage resources and reconcile them towards their declared state. A single pane of glass, like Cloud ALM, brings monitoring information together so we can see what is happening across the landscape. We have visibility, automation and, where supported, self-healing.
The blind spot is the time loop. It is Friday afternoon and everyone is submitting their timesheets. Your application depends on an API running elsewhere, on AWS. That API starts throttling requests and returns HTTP 429. The self-healing automation retries with exponential backoff. It waits, tries again, waits longer, tries again. We don’t see red. We see “retrying” or “reconciling” and read that as healing. But the timesheets still aren’t going through. The status tells us what the automation is doing, not whether the business outcome is being restored.
Next up is our favourite topic, AI. Why wouldn’t we give the observability problem to AI? It looks like the perfect fit. It can ingest our logs and look back over the history. It can read our code in Git and our tickets in Jira. It can read Confluence for the architecture design decisions and diagrams, as well as the functional and technical specs. It can read the emails telling us what SAP is planning to change. From all of that it can infer the root cause, identify an interim fix and propose a solution. Then it can go and make the change.
The first blind spot is cognitive surrender. It is fixing it, so we stop needing to know. Nobody reads the release note any more because something else read it. The answer is only as good as the context we gave it. This is not new. We have been giving build to one team and support to another for thirty years, and the understanding has never made the trip. AI does it faster, and at every layer at once.
The second is hallucinated infrastructure. It is not deterministic. The first time it fixes that destination, it follows the runbook or what is in Terraform. The second time, it decides the code is wrong or routes to a different region. The application works, but private data is now somewhere it should never have been. Nothing fails and nothing alerts. You only find out six months later, during an audit. By then, you could be facing fines and a loss of customer trust and goodwill.
Visibility without context is noise.
You cannot delegate understanding you do not have.
Six approaches, six blind spots

ClickOps has invisible drift. Scripts leave zombie resources. Terraform misses ghost resources. GitOps becomes the vengeful automaton. The control plane gets caught in a time loop. AI creates hallucinated infrastructure.
Every one of these approaches is useful. None gives us the whole truth. None of this means stop using them. It means know which lie yours tells.
Nobody is wrong

Back to the war room from the start. The system is down and the business is losing money by the minute.
Development is looking at Git, the CI/CD pipeline and Cloud Transport Management. The code was approved, the deployment succeeded and every test passed. Everything is good. But the code running in production hasn’t looked like the code in Git since Jimmy left six months ago.
Operations is looking at Cloud ALM and Cloud Logging. Everything is green. But their observability tools aren’t picking up the downstream API that started rate limiting ninety minutes ago.
Compliance says everything is good. But its checks were run against what was deployed to production twelve months ago. They aren’t seeing the changes made since then or whether the live system still meets those requirements.
The vendor is looking at their platform. Their services are available and their checks are passing. From their view, everything is working.
So who is lying? Nobody. Every one of them is telling the truth about what they can see. Four people with their hands on different parts of the same elephant.
Everyone is looking through their own tools, with their own blind spots. Each sees part of the system. Nobody owns the frame that brings those views together.
What it costs

Money. Those zombie resources keep spending the budget. The script stopped, but the bill didn’t.
Time. An incident becomes an archaeological dig. Someone has to piece together what changed, in which environment, when and why. And if AI made the fix, what else did it change along the way? Two hours becomes two days. Every one of those hours is an hour not spent building anything.
Trust. Users stop trusting a system that behaves differently on Monday morning than it did on Friday afternoon, especially when nobody can explain why. Development starts blaming support. Support points back at development. And security always thought the developers were cowboys anyway. Teams retreat into their silos, making the shared understanding even harder to recover.
The resource bill is only part of the cost. The expensive part is the time spent reconstructing what happened and the trust lost along the way. You can stop the spend. You cannot get those hours back, and trust takes longer to rebuild.
Drift is feedback

Back on premise, we treated drift as a bug to squash. Something had moved away from what we expected. Find what changed and put it back.
In the cloud, drift is not a bug, it is a feature of the platform. It is a signal that something changed, and feedback we can use. The platform changes. Services evolve. Standards tighten. SAP and the other providers are running a business. They have to evolve, scale and meet demand. The gap between what we declared and what now exists tells us that something needs our attention.
We already have a loop for the changes we make. Plan, code, build, test, release. We need a loop for the changes we receive. Detect, understand, respond, learn.
Detect what changed. Understand why it changed and what it means for the solution. Respond based on that understanding. Then record what we learned and update what we expect across environments.
A change is made in production. We detect the difference, understand why it was needed and assess whether it belongs in other environments. We update the shared declaration and document the reason. The fix is now part of the design or runbook, rather than another exception only Jimmy understands.
The gap tells us something changed. It does not tell us to put it back.
Drift happens silently and compounds over time. We need to listen to those signals, understand what changed and feed that understanding back into the design, the code or the runbook before we end up in a war room. The business does not care that everything looks green. They pay for a service, and right now it is costing them more money broken than it ever did working.
Where that leaves us

Drift is the gap between what we designed and what is running. On-premise thinking does not survive the river. We know how to manage the changes we choose, but we still do not know who owns the changes we inherit. Every tool we use holds a piece, and each has a blind spot. And the second loop has no owner.
Delegate execution, not understanding

As we wrap up, I am not suggesting you don’t use automation. On the contrary, use as much automation as you can. The faster we move, the faster we need to catch and fix, and Terraform, self-healing observability and AI are definitely the right choice for that. What I am saying is that you can delegate the execution to them, just not the understanding.
If we acknowledge that drift is a feature, we need to teach teams the water, not just the reel. Stop seeing the architecture diagram as a permanent blueprint and start seeing it as a temporary hypothesis. Accept change as a way of life.
With cognitive surrender comes cognitive drift. Neither is new. We get one consultancy to design, another to build and another to support. The understanding does not always make those handovers. But we still have someone to go back to, someone we can ask why a decision was made and what needs to happen next. AI can help us do more, faster. But when we need to understand why something was done, who do we go back to?
Who owns understanding

At the end of a project, you know who to hand over to. You know who will support the application, own the platform and answer the call when something breaks.
But who carries the understanding forward as the system changes?
The obvious answer is support. But support is incentivised to fix things and close tickets, not to spend weeks investigating why an API hits its rate limit on the last Friday of the month. They do not have the skills for that either. Development is the same. They are measured on delivering the next change, not on working out why the last one stopped working.
I do not have an answer for you. But I will leave you with one more question.
Next time you are in a war room, look around and ask yourself who owns the bigger picture.
Remember Jimmy, the one who made the change and then left. Jimmy at least knew why. Hand the critical and systems thinking to AI, and the next time the war room fills up, nobody will know why.
That is the drift that haunts your BTP.
If you recognised any of this, I would like to hear about it. You can find me on LinkedIn, on Bluesky at @jasper07.secondphase.com.au, or by email at john.patterson@secondphase.com.au.



