< Back to Careers
Singapore
Full-time · Arcfra
Customer Support Engineer
How to apply
Email your resume to:
careers@arcfra.com
To apply
To apply, please submit your application online at https://www.arcfra.com/careers.
Who We Are
Arcfra is a Singapore-headquartered IT innovator that simplifies on-premises enterprise cloud infrastructure with its full-stack, software-defined platform. In the cloud and AI era, we help enterprises effortlessly build robust on-premises cloud infrastructure from bare metal, offering computing, storage, networking, security, backup, disaster recovery, Kubernetes service, and more in one stack. Our streamlined design supports various workloads like traditional VM, container and K8S, AI/ML as well as VDI desktops.
To learn more, please visit: https://www.arcfra.com/
About the Role
We are looking for a highly capable Customer Support Engineer to provide advanced technical support for customers operating Arcfra’s enterprise cloud and hyperconverged infrastructure platform. This role owns complex technical issues across virtualization, distributed storage, networking, security, backup, disaster recovery, and platform operations. The successful candidate combines Professional Services delivery capability with deeper troubleshooting expertise, strong incident ownership, and the ability to manage high-priority production issues.
Key Responsibilities
  • Provide advanced remote and on-site technical support for Arcfra customers across Singapore and the wider region.
  • Take technical ownership of complex customer issues from investigation through resolution, root-cause analysis, and closure.
  • Diagnose issues across compute, virtualization, distributed storage, networking, security, management plane, backup, replication, disaster recovery, and hardware integration layers.
  • Handle high-severity production incidents involving service degradation, platform unavailability, data availability, performance, migration, or upgrade failures.
  • Perform incident triage, impact assessment, evidence collection, containment, recovery planning, escalation, and customer communications.
  • Analyse logs, alerts, metrics, configuration data, topology, version details, and operational history to isolate fault domains and identify root causes.
  • Work with Engineering and Product teams to reproduce defects, validate fixes, assess product behaviour, and track resolution progress.
  • Provide guidance on configuration, upgrades, patching, capacity planning, performance optimisation, security hardening, and operational best practices.
  • Support complex upgrade, expansion, migration, and DR activities including pre-checks, risk reviews, change planning, validation, and post-change verification.
  • Produce detailed case notes, RCA reports, knowledge-base articles, troubleshooting guides, and customer-facing technical summaries.
  • Mentor Resident Engineers and contribute to the technical development of Professional Services and partner teams.
  • Identify recurring product, process, documentation, and usability issues and provide actionable feedback internally.
Technical Requirements

Core Platform and Infrastructure Expertise

  • Strong hands-on expertise in virtualization, hyperconverged infrastructure, private cloud, distributed storage, and enterprise data-centre operations.
  • Deep understanding of KVM, VMware vSphere, OpenStack, HCI, or similar enterprise virtualization and cloud infrastructure platforms.
  • Strong Linux troubleshooting capability across logs, services, CPU, memory, file systems, disk and network diagnostics, kernel messages, and resource contention.
  • Proven ability to isolate faults across hosts, hypervisors, virtual machines, storage, networks, management services, external integrations, and hardware components.

Storage, Performance, and Resilience

  • Advanced knowledge of replication, striping, consistency, failure domains, rebuild processes, snapshots, cloning, thin provisioning, and capacity management.
  • Strong capability in diagnosing storage performance using IOPS, throughput, latency, queue depth, disk utilisation, workload patterns, and network behaviour.
  • Experience troubleshooting cluster health, node failures, storage availability, virtual volumes, capacity anomalies, and performance degradation.
  • Strong understanding of backup, restore, replication, failover, failback, DR testing, and business-continuity processes.
  • Experience with synchronous or asynchronous replication, stretched clusters, active-active architectures, or VM-level DR is highly desirable.

Networking, Security, and Support Engineering

  • Advanced troubleshooting across TCP/IP, VLANs, routing, MTU, bonding/LACP, DNS, firewalls, virtual switching, traffic isolation, and network connectivity.
  • Familiarity with RDMA, RoCE, SR-IOV, vGPU, GPU virtualization, multi-path networking, NIC bonding, and high-performance network configuration is highly desirable.
  • Understanding of secure access, SSH hardening, encryption in transit and at rest, key management, auditability, traffic monitoring, and security baseline enforcement.
  • Strong technical case-management discipline covering issue classification, priority assessment, escalation, evidence collection, customer updates, and resolution tracking.
  • Proven ability to prepare complete and actionable escalation packages for Engineering teams.
  • Strong Shell scripting; Python, Ansible, APIs, log analytics, monitoring automation, ITSM, incident management, problem management, and change management experience is preferred.
Experience and Qualifications
  • Bachelor’s Degree in Computer Science, Information Technology, Engineering, or a related discipline; equivalent practical experience will also be considered.
  • 5-8+ years of experience in enterprise technical support, infrastructure engineering, virtualization, private cloud, HCI, storage, cloud platforms, or data-centre operations.
  • At least 3 years of hands-on experience resolving complex issues in production enterprise environments.
  • Demonstrated experience managing critical incidents involving availability, performance, storage, networking, virtualization, or disaster recovery.
  • Prior Professional Services, implementation, systems integration, or technical consulting experience is strongly preferred.
  • Experience supporting enterprise customers under defined SLAs and incident-management processes is highly desirable.
  • VMware VCP/VCAP, RHCE, CCNP, HCIP, cloud infrastructure, or storage certifications are advantageous.
  • Working proficiency in English, including clear customer communication and professional technical documentation; additional Asian language capability is advantageous.
Desired Competencies
  • Strong ownership mindset and commitment to customer outcomes.
  • Excellent analytical and troubleshooting skills with an evidence-based decision-making approach.
  • Ability to communicate calmly, clearly, and professionally during high-pressure incidents.
  • Ability to explain complex technical risks and remediation options to varied stakeholder groups.
  • Strong cross-functional collaboration with Engineering, Product, QA, Professional Services, Sales Engineering, and external partners.
  • Ability to mentor junior engineers and contribute to technical excellence and knowledge sharing.
  • Willingness to participate in support rotations and respond to critical customer issues outside standard business hours when required.