Products

Custom Alert Rules in Arcfra AECP: Provide Tailored Monitoring for Every Virtual Machine

Published on by Arcfra Team
Last edited on

In traditional virtualization, alerting capabilities often focus on hosts or clusters, making it difficult to detect performance bottlenecks, resource pressure, and sustained abnormal resource usage within a VM. Operations teams have to rely on manual troubleshooting and third-party monitoring tools, leading to delayed responses, more manual effort, and monitoring costs that increase with VM growth.

To address these challenges, Arcfra Enterprise Cloud Platform (AECP) upgrades alerting capabilities by allowing users to create custom alert rules with virtual machines as alert objects. This provides fine-grained, VM-level monitoring and alerting, helping enterprises establish an intelligent, business-centered alerting system that adapts to different workload requirements.

Why Fine-Grained Custom Alerting Matters

Lacking VM-level visibility and alerting can often lead to the following recurring challenges:

1. Single-Condition Alerts Make Persistent VM Anomalies Hard to Detect

When a VM experiences sustained high memory usage or prolonged abnormal disk I/O, traditional platforms often fail to correlate data across multiple dimensions to issue an early warning. Operations teams are left reacting only after application users notice the impact, rather than detecting persistent anomalies proactively.

Customer Story 1: Six Days at Full Disk I/O Load, Missed by Traditional Alerts

In a financial services customer’s VMware environment with third-party distributed storage, a misconfigured antivirus policy caused several endpoint agents to repeatedly scan files on VM disks. As a result, disk I/O remained at full load for six consecutive days. The traffic appeared normal at the storage and virtualization layers, but it consumed substantial bandwidth and increased the risk of read/write latency and service disruption.

Because traditional alerts relied only on point-in-time thresholds, they could not determine how long the high I/O load had persisted or compare fluctuating traffic against historical baselines. This was a typical case of “abnormal behavior hidden inside seemingly normal traffic”, and one of the key problems Arcfra AECP aims to solve with its alerting system.

Applying the same alert thresholds for all workloads also makes it hard to adjust to each business system’s tolerance for resource changes. The platform cannot tell which workloads are more critical, and operations teams cannot tune alerts to match service needs.

2. Coarse Alert Details Slow Down Root Cause Analysis

Traditional alerts often stop at the host or cluster level, covering a broad scope but offering limited detail. As a result, teams struggle to identify the affected component from the alert itself, slowing root-cause analysis and remediation.

Customer Story 2: A Vague Disk-Full Alert Delayed Troubleshooting Until Business Disruption

In a financial services customer’s production environment, a core business VM suddenly failed to write data, triggering repeated application errors. The operations team suspected a disk space issue, but the traditional platform showed only the VM’s overall disk utilization without partition-level data. Administrators had to log in to the VM to investigate manually, delaying troubleshooting and leading to service disruption.

This incident exposed the limitations of coarse-grained alerts and delayed response in traditional platforms, reinforcing AECP’s need to upgrade custom alerting with file system partition visualization capabilities.

3. Per-VM Licensing Drives Up Monitoring Costs as VM Counts Grow

Enterprise IT resources are dynamic assets that evolve with the business demands. When monitoring tools are licensed per VM, every expansion also increases monitoring costs, turning the licensing model into a hidden constraint on business growth. As VM counts grow from dozens to hundreds, monitoring spend can rise sharply.

Customer Story 3: Scaling to Hundreds of VMs Made Rising Monitoring Costs Seem Inevitable

A financial services customer initially used a third-party tool to monitor and manage VMs after deploying AECP. As the business expanded, the VM count quickly grew from dozens to hundreds. The customer then realized that the tool was licensed per VM, causing monitoring costs to rise in step with resource expansion.

As the business continued to grow, the rising expenses increasingly conflicted with the company’s goal of reducing costs and improving efficiency. It even made business teams hesitate when planning further resource expansion.

Moreover, with most monitoring tools on the market using similar per-node licensing models, the team initially assumed that rising monitoring costs were simply an unavoidable consequence of VM growth.

These challenges become more serious in large-scale, multi-tenant environments, especially across financial services, public-sector IT, and development and testing. Operations teams need alerting that combines accurate VM-level monitoring with flexible rule and notification policies, so they can improve operational efficiency and service quality.

Arcfra AECP: Supporting VM-Level Custom Alerting

The Arcfra Operation Center (AOC)’s Observability Platform provides custom alert rules with VM-level monitoring at its core, enabling fine-grained, multi-dimensional alert configuration.

  • Create metric-based alert rules for a single VM or multiple VMs.
  • Define the severity level (Info, Note, or Critical), trigger threshold, duration, and notification policy for each rule.

Alert Dimensions

CategoryMetricsPurpose
CPUHigh CPU usage; high CPU ready time (%)Detect resource pressure and performance fluctuations
MemoryHigh memory usageDetect resource pressure and performance fluctuations
DiskHigh file system partition utilization; high average I/O latency; high I/O bandwidth; high disk IOPSMonitor both storage performance and capacity
NetworkTotal network bandwidthDetect abnormal traffic spikes

These alert rules can all be configured graphically and managed centrally in AOC. Alerts can be delivered through in-platform notifications, email, and webhooks.

Innovations and Product Comparison

CapabilityArcfra AECP Custom Alert RulesTraditional Virtualization Platforms
Alert Object Support✅ Fine-grained rules for individual VMsPrimarily host- or cluster-level alerts
Metric Coverage✅ Covers CPU, memory, I/O, NICs, and file system partitionsTypically covers fewer monitoring dimensions
Customization✅ Configurable thresholds, levels, duration, and notificationsUsually limited to default levels or threshold changes


Business Value and Customer Benefits

With custom alert rules in AOC, enterprises can strengthen IT operations with more precise monitoring, earlier response, and better-informed decisions:

  • Greater efficiency and lower risk: Shift from reactive troubleshooting to early warning. This reduces manual inspection, speeds up response, and helps prevent service interruptions caused by undetected resource bottlenecks, such as high memory usage, high I/O, or full partitions.
  • Better risk visibility: Monitor CPU, memory, disk I/O, file system partition usage, network bandwidth, and other resource metrics to detect risks earlier and shorten troubleshooting time.
  • More flexible monitoring: Configure alert rules by VM, VM group, or cluster. Customize thresholds, duration, severity levels, and notification policies to match different service needs.
  • Connect Alerting with Configuration and Remediation: After an alert is triggered, use AOC capabilities such as partition visualization, remote configuration, and snapshot rollback to quickly locate and resolve the issue.

Customer Story 1 Follow-Up: Multi-Dimension Monitoring Enables Early Detection of Hidden Issues

After adopting Arcfra Observability, the customer created a VM-level custom alert rule tailored to its environment: trigger an alert when a VM’s I/O bandwidth reaches twice its average over the previous seven days and remains at that level for five minutes. The platform can then detect the anomaly immediately, notify the team, and help operations teams quickly identify and address the affected VM.

By evaluating threshold, duration, and historical baseline together, AECP can detect hidden issues without prolonged manual observation. Resource anomalies hidden behind seemingly normal traffic become visible automatically, moving operations from reactive troubleshooting to proactive alerting.

Customer Story 2 Follow-Up: Alerts and Visualization Catch Risks Before Service Disruption

The customer enabled custom alerting in AOC and configured a threshold rule for VM file system partition usage: trigger a warning when any partition exceeds 90% usage for five minutes. With file system partition visualization from Arcfra VMTools, administrators can view usage details for each mount point without logging in to the VM and quickly identify that one software partition was running low on space.

Because this mount point was used for business data writes, continued data growth would have caused write failures. By combining alerting from Arcfra Observability with visualization from Arcfra VMTools, the team identified and addressed the risk before service was affected. They expanded the partition within minutes, preventing service disruption and avoiding broader business impact.

Customer Story 3 Follow-Up: Arcfra AECP Keeps Monitoring Costs from Growing with VM Count

After deploying AECP, the customer broke the link between VM growth and monitoring costs. Arcfra Observability is integrated with AOC, requires no additional licensing fees, and covers resources across the entire AECP cluster. No matter how many VMs the customer creates, monitoring remains available across the environment in real time.

This supports rapid business expansion. In the first year alone, the customer reduced monitoring spend by more than 90%, with savings continuing to grow as the VM footprint expands.

References:

About Arcfra

Arcfra simplifies enterprise cloud infrastructure with a full-stack, software-defined platform built for the AI era. We deliver computing, storage, networking, security, Kubernetes, and more — all in one streamlined solution. Supporting VMs, containers, and AI workloads, Arcfra offers future-proof infrastructure trusted by enterprises across e-commerce, finance, and manufacturing. Arcfra is recognized by Gartner as a Representative Vendor in full-stack hyperconverged infrastructure. Learn more at www.arcfra.com.