HPC High Performance Computing IT Infra Engineer
D L RESOURCES PTE LTD
Client: Research & Education Sector
Key Responsibilities
Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure
Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments
* Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)
* Administer physical and virtualized server environments (x86 architecture, VMware/KVM)
* Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation
* Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)
* Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)
* Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules
* Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives
* Maintain system documentation, operational procedures, and runbooks
Required Skills & Experience
1–3 years of experience in system administration, infrastructure support, or IT operations
Hands-on experience with Linux/Unix systems administration and command-line environments
* Basic exposure to HPC, distributed systems, or parallel computing environments
* Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)
* Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)
* Scripting and automation using Bash/Shell and/or Python
* Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)
* Strong troubleshooting, analytical thinking, and incident resolution skills
Preferred Skills
Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness)
Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)
* Awareness of container technologies (Docker) and basic DevOps practices
Project / Environment Tech Stack. (Mostly On-Prem)
The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads.
80% On-Premise HPC Infrastructure
Dell servers, Huawei servers, physical data centre infrastructure
x86 server architecture, CPU-based compute infrastructure, bare-metal servers
Virtualized server environments using VMware / KVM where applicable
Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX
HPC cluster components across compute, storage, and networking layers
High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable
Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred
HPC workload scheduling and batch processing using Slurm, PBS, or LSF
Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk
20% AWS HPC / GPU-Related Environment
AWS cloud infrastructure supporting HPC / GPU-related workloads
AWS GPU server exposure, including NVIDIA GPU-based compute instances
GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable
Cloud HPC workload support involving compute, storage, networking, and security configurations
Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups
Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices
Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources