HPC High Performance Computing IT Infra Engineer

D L RESOURCES PTE LTD

Client: Research & Education Sector

Key Responsibilities

Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure

Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments

* Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)

* Administer physical and virtualized server environments (x86 architecture, VMware/KVM)

* Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation

* Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)

* Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)

* Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules

* Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives

* Maintain system documentation, operational procedures, and runbooks

Required Skills & Experience

1–3 years of experience in system administration, infrastructure support, or IT operations

Hands-on experience with Linux/Unix systems administration and command-line environments

* Basic exposure to HPC, distributed systems, or parallel computing environments

* Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)

* Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)

* Scripting and automation using Bash/Shell and/or Python

* Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)

* Strong troubleshooting, analytical thinking, and incident resolution skills

Preferred Skills

Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness)

Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)

* Awareness of container technologies (Docker) and basic DevOps practices

Project / Environment Tech Stack. (Mostly On-Prem)

The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads.

80% On-Premise HPC Infrastructure

  • Dell servers, Huawei servers, physical data centre infrastructure

  • x86 server architecture, CPU-based compute infrastructure, bare-metal servers

  • Virtualized server environments using VMware / KVM where applicable

  • Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX

  • HPC cluster components across compute, storage, and networking layers

  • High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable

  • Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred

  • HPC workload scheduling and batch processing using Slurm, PBS, or LSF

  • Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk

20% AWS HPC / GPU-Related Environment

  • AWS cloud infrastructure supporting HPC / GPU-related workloads

  • AWS GPU server exposure, including NVIDIA GPU-based compute instances

  • GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable

  • Cloud HPC workload support involving compute, storage, networking, and security configurations

  • Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups

  • Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices

  • Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.