Skip to content

Service Health

Beyond triaging individual alerts, AIOps gives you a rolled-up view of environment health: which metrics are behaving abnormally, which business services are impacted, and how complete your discovery and service mapping actually is.

The Anomaly alerts board lists metric deviations flagged by anomaly detection: each row shows the metric, node, and business service, the observed value against its learned bound (above or below), and an anomaly score. KPI cards summarize open anomalies, how many are critical, how many distinct services are affected, and the average score across the board.

A business (or application) service is the unit alerts impact and the endpoint service mapping discovers around. It tracks operational health as its own lifecycle:

State Typical next action
Operational Mark Degraded, or Start Maintenance
Degraded Mark Operational, or Start Maintenance
Non-Operational Mark Operational, or Start Maintenance
In Maintenance Mark Operational
Retired Terminal (collects a note when retiring from any active state)

A service also carries a classification (Business / Technical / Application / Infrastructure Service) and a business criticality rank (1 – most critical, through 4 – not critical), alongside its owner, managing group, support group, and environment.

The Impact board rolls every business service up into a single row: a health band (healthy / degraded / critical), the worst open-alert severity affecting it, its open alert and open incident counts, trailing availability, a 0–100 health score, and its owner. The health score weighs a service’s worst open-alert severity most heavily, then adds a smaller penalty for each additional open alert, each open incident, and recent change activity — so a service with one critical alert reads noticeably worse than one with several minor ones. Services with any open alert are the “impacted” set called out in the summary KPIs.

The Discovery board lists every scheduled discovery run: its target (an IP range or cloud account), what kind of scan it runs, which MID server executes it, its cron schedule, and the outcome of its last run — completion status, duration, and devices found. KPI cards summarize total schedules, completed runs, runs that errored, and total devices discovered across all schedules.

Service Mapping is the top-down view of what’s actually been mapped: for each business service, how many configuration items were discovered, how many tiers and entry points the map has, when it was last mapped, and a completeness percentage. A health band mirrors the same healthy/degraded/critical classification used on the Impact board.

The MID Servers board shows each Management, Instrumentation, and Discovery server’s status (Up / Down / Warning), whether it’s validated, its version, its registered capabilities, its ECC queue depth, and its thread load and last check-in time. A queue depth above threshold is flagged so a backlogged MID server — which delays discovery and event processing behind it — is easy to spot before it becomes an outage.

Synthetic monitoring runs scheduled checks against key user journeys and reports their uptime and last result (passed/degraded/failed) per location. The Log Viewer streams raw log/event lines with a severity filter, useful for confirming what a monitoring tool actually saw before it raised an event.