At our company we use the prometheus/grafana stack for measuring quality and uptime of our landscape. Opsgenie is being used as our alerting mechanism. So far so good!
But... as a final check we want to have a total independent "last resort"-check, from external locations. What are best practices for this?
What we want:
- Being able to validate uptime (read: HTTP response codes) from (multiple) locations
- If something is down (!= 200 range), then alert the corresponding innovation team using opsgenie
Curious to hear what external solutions are available with embedded opsgenie support.