Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE PTE. LTD.

Role Overview

  • We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
  • The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems.
  • Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.

Key Responsibilities

  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments.
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform detailed root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics.
  • Support Kubernetes upgrades, configuration changes, patching and platform improvements.
  • Work closely with application engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
  • Participate in production readiness reviews and ensure applications meet operational standards.

Mandatory Skills

  • Strong hands-on experience with Kubernetes administration and troubleshooting.
  • Strong understanding of Kubernetes architecture, including:
  • Pods
  • Deployments
  • StatefulSets
  • Services
  • Ingress
  • ConfigMaps / Secrets
  • RBAC
  • Storage
  • Networking
  • Strong Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
  • Strong understanding of Site Reliability Engineering principles.
  • Experience supporting large-scale, high-availability production systems.
  • Strong incident management and RCA experience.
  • Hands-on scripting/automation experience using Python, Bash/Shell or similar.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent.
  • Experience working with CI/CD and automated deployment environments.
  • Strong debugging and problem-solving capabilities.

Preferred Skills

  • Infrastructure-as-Code experience using Terraform, Ansible or equivalent.
  • Helm or similar Kubernetes package/deployment management tools.
  • GitOps experience using tools such as Argo CD or Flux.
  • Knowledge of service mesh concepts.
  • Experience with container security and Kubernetes security practices.
  • Experience with cloud or private-cloud infrastructure.
  • Familiarity with distributed systems and microservices architectures.
  • Exposure to performance engineering and capacity management.
  • Experience working in globally distributed engineering environments.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.