< Back to Careers
Singapore
Full-time ยท Arcfra
Sustaining Software Engineer - Distributed Storage
How to apply
Email your resume to:
careers@arcfra.com
About the Role
Arcfra is looking for a Sustaining Software Engineer with a strong background in distributed systems, storage technologies, and C/C++ development.
In this role, you will be responsible for maintaining and improving our production-grade storage products. You will investigate complex software defects, troubleshoot customer-reported issues, develop reliable fixes, and work closely with support, QA, and engineering teams to ensure product stability.
This is a hands-on engineering position. It combines software development, production troubleshooting, root-cause analysis, and customer issue resolution. The ideal candidate enjoys working on mature and complex systems, understands how software behaves in real-world production environments, and is comfortable debugging problems across multiple layers of the system.
Key Responsibilities
  • Maintain and improve distributed storage and data infrastructure products.
  • Investigate and resolve customer-reported product issues, including performance degradation, service interruptions, data-path failures, and unexpected system behavior.
  • Reproduce production issues in laboratory or test environments and identify their root causes.
  • Develop, review, test, and deliver high-quality bug fixes using C or C++.
  • Analyze system logs, core dumps, stack traces, performance metrics, and storage-related diagnostic information.
  • Troubleshoot issues involving distributed systems, storage engines, networking, operating systems, concurrency, and hardware interactions.
  • Work closely with customer support and field engineering teams to collect technical information and provide troubleshooting guidance.
  • Collaborate with product development teams on complex defects, architectural improvements, and product reliability.
  • Create diagnostic tools, scripts, test cases, and internal documentation to improve troubleshooting efficiency.
  • Participate in product release validation and ensure fixes are safely backported to supported product versions.
  • Contribute to improving product observability, serviceability, stability, and maintainability.
Required Qualifications
  • Bachelor's degree or above in Computer Science, Computer Engineering, Software Engineering, or a related field.
  • Strong programming experience in C or C++.
  • Solid understanding of data structures, algorithms, multithreading, memory management, and network programming.
  • Practical experience with Linux systems and Linux debugging tools.
  • Good understanding of distributed systems concepts, such as replication, consistency, consensus, fault tolerance, distributed state management, and failure recovery.
  • Experience with one or more storage technologies, such as:
    • Distributed storage systems
    • Block, file, or object storage
    • Storage virtualization
    • RAID, snapshots, replication, or data protection
    • Local file systems or storage engines
    • NVMe, SCSI, iSCSI, Fibre Channel, or NVMe over Fabrics
  • Strong troubleshooting and root-cause analysis skills.
  • Ability to read and understand a large and mature codebase.
  • Ability to communicate technical findings clearly in written and spoken English and Chinese.
  • Willingness to work directly with customer-facing teams on complex production issues.
  • Must be a Singapore citizen, permanent resident, or hold a valid Singapore work permit.
Preferred Qualifications
  • Experience developing or maintaining distributed storage products.
  • Experience with Ceph, SPDK, RocksDB, LevelDB, distributed databases, or similar infrastructure software.
  • Familiarity with Linux kernel storage, device drivers, file systems, or networking subsystems.
  • Experience analyzing core dumps with GDB and diagnosing memory corruption, deadlocks, race conditions, and performance bottlenecks.
  • Experience with observability and diagnostic tools such as perf, eBPF, Valgrind, AddressSanitizer, strace, SystemTap, or similar tools.
  • Familiarity with Python, Bash, or other scripting languages for testing and automation.
  • Experience in enterprise infrastructure, cloud platforms, hyper-converged infrastructure, or data centre products.
  • Previous experience in sustaining engineering, escalation engineering, product support engineering, or site reliability engineering.
What We Value
  • A strong sense of ownership and responsibility for product quality.
  • Patience and persistence when investigating difficult technical problems.
  • The ability to distinguish symptoms from root causes.
  • A practical engineering mindset focused on reliable and maintainable solutions.
  • The ability to work effectively across engineering, QA, support, and customer-facing teams.
  • A willingness to understand both the product source code and the customer's production environment.