In traditional virtualization, alerting capabilities often focus on hosts or clusters, making it difficult to detect performance bottlenecks, resource pressure, and sustained abnormal resource usage within a VM. Operations teams have to rely on manual troubleshooting and third-party monitoring tools, leading to delayed responses, more manual effort, and monitoring costs that increase with VM growth.
To address these challenges, Arcfra Enterprise Cloud Platform (AECP) upgrades alerting capabilities by allowing users to create custom alert rules with virtual machines as alert objects. This provides fine-grained, VM-level monitoring and alerting, helping enterprises establish an intelligent, business-centered alerting system that adapts to different workload requirements.
Lacking VM-level visibility and alerting can often lead to the following recurring challenges:
When a VM experiences sustained high memory usage or prolonged abnormal disk I/O, traditional platforms often fail to correlate data across multiple dimensions to issue an early warning. Operations teams are left reacting only after application users notice the impact, rather than detecting persistent anomalies proactively.
Customer Story 1: Six Days at Full Disk I/O Load, Missed by Traditional Alerts
In a financial services customer’s VMware environment with third-party distributed storage, a misconfigured antivirus policy caused several endpoint agents to repeatedly scan files on VM disks. As a result, disk I/O remained at full load for six consecutive days. The traffic appeared normal at the storage and virtualization layers, but it consumed substantial bandwidth and increased the risk of read/write latency and service disruption.
Because traditional alerts relied only on point-in-time thresholds, they could not determine how long the high I/O load had persisted or compare fluctuating traffic against historical baselines. This was a typical case of “abnormal behavior hidden inside seemingly normal traffic”, and one of the key problems Arcfra AECP aims to solve with its alerting system.
Applying the same alert thresholds for all workloads also makes it hard to adjust to each business system’s tolerance for resource changes. The platform cannot tell which workloads are more critical, and operations teams cannot tune alerts to match service needs.
Traditional alerts often stop at the host or cluster level, covering a broad scope but offering limited detail. As a result, teams struggle to identify the affected component from the alert itself, slowing root-cause analysis and remediation.
Customer Story 2: A Vague Disk-Full Alert Delayed Troubleshooting Until Business Disruption
In a financial services customer’s production environment, a core business VM suddenly failed to write data, triggering repeated application errors. The operations team suspected a disk space issue, but the traditional platform showed only the VM’s overall disk utilization without partition-level data. Administrators had to log in to the VM to investigate manually, delaying troubleshooting and leading to service disruption.
This incident exposed the limitations of coarse-grained alerts and delayed response in traditional platforms, reinforcing AECP’s need to upgrade custom alerting with file system partition visualization capabilities.
Enterprise IT resources are dynamic assets that evolve with the business demands. When monitoring tools are licensed per VM, every expansion also increases monitoring costs, turning the licensing model into a hidden constraint on business growth. As VM counts grow from dozens to hundreds, monitoring spend can rise sharply.
Customer Story 3: Scaling to Hundreds of VMs Made Rising Monitoring Costs Seem Inevitable
A financial services customer initially used a third-party tool to monitor and manage VMs after deploying AECP. As the business expanded, the VM count quickly grew from dozens to hundreds. The customer then realized that the tool was licensed per VM, causing monitoring costs to rise in step with resource expansion.
As the business continued to grow, the rising expenses increasingly conflicted with the company’s goal of reducing costs and improving efficiency. It even made business teams hesitate when planning further resource expansion.
Moreover, with most monitoring tools on the market using similar per-node licensing models, the team initially assumed that rising monitoring costs were simply an unavoidable consequence of VM growth.
These challenges become more serious in large-scale, multi-tenant environments, especially across financial services, public-sector IT, and development and testing. Operations teams need alerting that combines accurate VM-level monitoring with flexible rule and notification policies, so they can improve operational efficiency and service quality.
The Arcfra Operation Center (AOC)’s Observability Platform provides custom alert rules with VM-level monitoring at its core, enabling fine-grained, multi-dimensional alert configuration.
| Category | Metrics | Purpose |
|---|---|---|
| CPU | High CPU usage; high CPU ready time (%) | Detect resource pressure and performance fluctuations |
| Memory | High memory usage | Detect resource pressure and performance fluctuations |
| Disk | High file system partition utilization; high average I/O latency; high I/O bandwidth; high disk IOPS | Monitor both storage performance and capacity |
| Network | Total network bandwidth | Detect abnormal traffic spikes |
These alert rules can all be configured graphically and managed centrally in AOC. Alerts can be delivered through in-platform notifications, email, and webhooks.
| Capability | Arcfra AECP Custom Alert Rules | Traditional Virtualization Platforms |
|---|---|---|
| Alert Object Support | ✅ Fine-grained rules for individual VMs | Primarily host- or cluster-level alerts |
| Metric Coverage | ✅ Covers CPU, memory, I/O, NICs, and file system partitions | Typically covers fewer monitoring dimensions |
| Customization | ✅ Configurable thresholds, levels, duration, and notifications | Usually limited to default levels or threshold changes |
With custom alert rules in AOC, enterprises can strengthen IT operations with more precise monitoring, earlier response, and better-informed decisions:
After adopting Arcfra Observability, the customer created a VM-level custom alert rule tailored to its environment: trigger an alert when a VM’s I/O bandwidth reaches twice its average over the previous seven days and remains at that level for five minutes. The platform can then detect the anomaly immediately, notify the team, and help operations teams quickly identify and address the affected VM.
By evaluating threshold, duration, and historical baseline together, AECP can detect hidden issues without prolonged manual observation. Resource anomalies hidden behind seemingly normal traffic become visible automatically, moving operations from reactive troubleshooting to proactive alerting.
The customer enabled custom alerting in AOC and configured a threshold rule for VM file system partition usage: trigger a warning when any partition exceeds 90% usage for five minutes. With file system partition visualization from Arcfra VMTools, administrators can view usage details for each mount point without logging in to the VM and quickly identify that one software partition was running low on space.
Because this mount point was used for business data writes, continued data growth would have caused write failures. By combining alerting from Arcfra Observability with visualization from Arcfra VMTools, the team identified and addressed the risk before service was affected. They expanded the partition within minutes, preventing service disruption and avoiding broader business impact.



After deploying AECP, the customer broke the link between VM growth and monitoring costs. Arcfra Observability is integrated with AOC, requires no additional licensing fees, and covers resources across the entire AECP cluster. No matter how many VMs the customer creates, monitoring remains available across the environment in real time.
This supports rapid business expansion. In the first year alone, the customer reduced monitoring spend by more than 90%, with savings continuing to grow as the VM footprint expands.
Arcfra simplifies enterprise cloud infrastructure with a full-stack, software-defined platform built for the AI era. We deliver computing, storage, networking, security, Kubernetes, and more — all in one streamlined solution. Supporting VMs, containers, and AI workloads, Arcfra offers future-proof infrastructure trusted by enterprises across e-commerce, finance, and manufacturing. Arcfra is recognized by Gartner as a Representative Vendor in full-stack hyperconverged infrastructure. Learn more at www.arcfra.com.