In the modern world of the digital-first company (yes, that’s you!), the ability to quickly identify and effectively respond to issues is paramount, and in the cut-throat online industry, the ability to automate remediations and introduce self-healing is a game-changer. “This is nothing new!”, I hear you say, and you would be right! We IT folk have been trying to pull this off for years, with clunky manual scripts and ‘human automation’ that sometimes, kind of looked like we’d pulled it off, that we were, in fact, Tony Stark and J.A.R.V.I.S was real, but it was all smoke and mirrors, and deep inside, we knew it!

6,000 consumers across the US and Europe were asked about their online shopping behaviour and views on website functionality. 60% of consumers say they abandon purchases due to poor user experience on websites. E-commerce companies are losing an average of 5 purchases a year per consumer, with 8% of consumers abandoning more than 10 purchases.

Online study conducted by Storyblok

Well, that has all changed, it’s now not a case of ‘can’ we do it, it’s now more a case of ‘are you ready to’? The answer? Drum roll please… AIOps! With AIOPs, we can not only detect problems in real-time but also resolve them with minimal human overhead. By integrating sophisticated monitoring tools and automated remediation scripts, we really can ensure applications remain resilient, efficient, and consistently available despite the inevitable issues that occur in any IT environment!

The essence lies in three core pillars: 1) Real-Time Monitoring and Alerting, 2) Automated Remediation Scripts or Tools, and 3) A Feedback Loop for Continuous Improvement. Each of these components plays a critical role in not only addressing immediate issues but also learning and adapting over time. Lets break them down:

Real-Time Monitoring and Alerting:

What you’ll need: Robust monitoring tools such as AppDynamics, Dynatrace, DataDog and LogicMonitor or be running the open telemetry framework, all of which can detect issues like service/component crashes, memory leaks, or performance degradation in real-time.

What to monitor: It’s vital to track key metrics like heap memory usage, thread counts, garbage collection times, and CPU usage. These metrics can then be passed to your AIOps platform, where machine learning and advanced AI algorithms run predictive analysis and anomaly detection.

When asked their main reasons for leaving an e-commerce website, 37% said limited payment options, followed by 37% because of poor navigation or layout with 33% due to timeouts or slow loading speeds.

Online study conducted by Storyblok

Automated Remediation Scripts and Tools:

What you’ll need: Develop scripts and use orchestration tools like Ansible or Puppet that can perform remedial actions when called by the AIOPs platform, such as restarting the crashed/stalled/slowed services or components, clearing cache, or if you want to get really fancy adjust resource allocation.

How to do it: Integrate these tools and repos with your AIOPs platform. For instance, when a crashed/stalled/slowed service is detected, the AIOPs platform, having diagnosed the issue for you, can reach into the repo and trigger the required remediation action to bring the service back to life.

Feedback Loop and Continuous Improvement:

What you’ll need: A process where every incident and its resolution contribute to improving the overall system. This involves analysing incidents, refining monitoring events, metrics and alerts, and updating automated healing scripts.

How to do it:  Use tools like ServiceNow or JIRA for incident management. After each automated healing action, review the incident to understand its root cause and update your monitoring and remediation strategies accordingly. 

I can’t stress this enough: this is by far the most common reason for failure! Lots of organisations that I consult with have some form of AIOps and automated remediation capability, but it was set up once and left to die a slow, painful death, eventually earning a bad reputation or, worse, turning into shelfware; these are enterprise platforms that need constant watering!

In conclusion, automated application healing is a combination of proactive monitoring, effective automated remediation, and a continuous feedback loop. It requires not just the right tools but also a mindset of ongoing improvement and adaptation.

Think your organisation’s ready? Why not take my Readiness Checker to find out?