I work with a tech company that has about 30 development teams, doing our own form of scaled agile.
Within Jira, I have developed a common implementation - same issue types, workflow, custom fields, permissions, screens, etc.
We also use Advanced Roadmaps and have a hierarchy of:
- Pillar > Initiative > Epic > Product Backlog Item (PBI) > SubTask
I have developed almost 100 Automations - automating workflow state changes, syncing workflow (between Epic, PBI, SubTask), synching fields (between Epic, PBI, SubTask), automatically setting start date/finish date (based upon workflow status), automatically propagating fields (doing copy down from Epic -> PBI -> SubTask or copy up from SubTask -> PBI -> Epic), handling special operations (splitting, cloning, moving, deletion, etc.), calculating cycle time per PBI, calculating Average cycle time and Velocity per team, propagating comments in development tickets over to service desk tickets, etc.
Every once in a while, Automation for Jira goes down. Or, there's a glitch. And the automations don't run or fails. When this occurs, we effectively have "data corruption". That is, what is tracked in Jira doesn't jibe with the state it should have been in, had the Automations run. As an example, maybe the Automation to calculate the cycle time failed (for whatever reason), so the data is now missing.
To combat this, I started developing "monitors" to identify such cases. These "monitors" are scheduled to run daily. It's somewhat arbitrary, but I look for issues modified in the past 2 weeks and look for all sort of anomalies (workflow status incorrectly set, fields between Epic and PBI not matching as expected, etc.). The challenge is that these monitors are quite resource intensive. Also, because we have so many tickets, I need to run the "monitors" for "sets" of teams. As a result, I have ~ 100 monitors running daily.
I've done my best to generalize things, to minimize maintenance cost. I suppose I spend maybe 4 hours/week maintaining this.
Just recently, there was another outage and I started to wonder - do I need even more monitors -- as the teams are expecting the Automations to run flawlessly.
Anyhow, I figured that I would reach out to the community to ask - how do you ensure flawless execution (of your Automations)? Have you built a monitoring framework (to verify that the data is getting propagated as expected)? If not, what do you do to ensure the data is correct?
Thanks,
Doug