Grafana Alerting Architecture :: Kloudfuse Docs

Grafana Alerting Architecture

Kloudfuse uses Grafana Alert Engine for alert execution and the Alert Manager which is from Prometheus for handling notifications. This documentation is intended to help operations teams understand the alerting architecture and how the different components work together.

Basic Workflow for Alerts

  1. Each Evaluation Group will refer to its configuration to determine how and when the Alert Rules are executed.

  2. The Alert Rules that are triggered, forward the target instances to the Alert Manager.

  3. The Alert Manager organizes the triggered alerts, deciding how to group them together. It then uses Notification Policies to determine to which Contact Points receive specific notifications.

  4. A Contact Point will make an external call to forward the details to a target endpoint (SMTP, PagerDuty, Webhook, etc…​)

Alert Rules

An Alert Rule executes a query, either PromQL, or FuseQL, etc…​ to check if a result is above a certain threshold. When the threshold is breached, it will transition an alert instance from Normal to Pending or Firing. Each rule specifies:

Alert States

Each alert rule instance moves through the following states:

State Description
Normal The condition is not met. No notification is sent.
Pending The condition has been met, but the Pending period has not yet elapsed. The alert will not fire until it has been in this state for the configured duration, reducing noise from transient spikes.
Alerting The state of an alert that has breached the threshold for longer than the pending period.
Firing The condition has been met for longer than the pending period. The alert is sent to the Alert Manager.
Recovering The state of a firing alert when the threshold is no longer breached, but for less than the keep firing period.
Error The query failed to execute. This can be from a query timeout or invalid query syntax. If it is an Error state you can configure the Rule to alert if you encounter this state, but would recommend avoiding that.
No Data The query returned no data. You can configure how the alert will behave ( Normal, Alerting, or No Data) when it encounters this situation.

Evaluation Groups

An Evaluation Group is a single process that is used to schedule one or more alert rules and control how they are run and how frequently they are run. The frequency is referred to as the evaluation interval.

Grafana-managed rules, within the same group, are executed concurrently. They are evaluated at different times over the same evaluation interval but display the same evaluation timestamp.

Data source-managed rules, within the same group, are evaluated sequentially, one after the other. This is useful to ensure that recording rules are evaluated before alert rules.

NOTE: The Kfuse UI does not support implementing the creation of Data source-managed Evaluation Rules.

Evaluation Groups are executed independently of each other, which allows Alert Rules to be executed concurrently. If you have 100 Groups, you can have 100 rules be executed at the same time.

Another advantage of having separate groups is when there is a problem within one group, for example: a query taking too long, the impact is limited to that one group. The rest of the groups can continue without interruption. The limitation on this is the more groups you have, the more resources it will require, specifically memory, CPU and network bandwidth.

Key properties of an Evaluation Group:

Alert Manager

The Alert Manager is based on the architecture of the Prometheus alerting system. It receives firing and resolved alert results from alert rules and sends notifications for those alerts.

It will use Notification Policies to decide how to route alerts based on labels passed from the alerts. A single alert can match multiple notification policies.

Notification Policies

A Notification Policy defines how an alert is routed to a Contact Point. Notification Policies are structured as a tree with a default root Default Policy with child policies nested underneath it.

Each policy contains:

Grouping Alerts

The Alert Manager allows you to bundle alerts into a smaller number of notifications. This is useful when you want to avoid spamming a contact with too many notifications that trigger at once. This can be configured in the Notification Policy using the Group by labels (for example, grouping all alerts with the same cluster and severity labels into a single message).

Default Policy

The Default Policy handles alerts that do not match a specific notification policy. By using a Default Policy with a catch-all contact point, you ensure that no alert is silently overlooked.

Policy Matching

Policies are evaluated from the most specific (deepest) child outward. The first policy whose label matchers all match the incoming alert handles the notification. If no child policy matches, the alert falls through to the default policy.