Platform Ops Engineer - #1640
JOBSTER PRIVATE LTD.
Role Overview
We are seeking a skilled Platform Operations Engineer to support the operation, availability, performance, security, and continuous improvement of technology platforms supporting the agency's digital services and business applications.
The successful candidate will be responsible for maintaining reliable and secure production environments, monitoring platform health, managing incidents and service requests, supporting application deployments, and implementing automation to improve operational efficiency.
The role requires close collaboration with application development, DevOps, cybersecurity, infrastructure, cloud, network, and business teams. The successful candidate should be comfortable working in a highly governed environment where system availability, cybersecurity, operational resilience, and service quality are critical.
Key Responsibilities
Platform Operations
Operate, monitor, and maintain enterprise technology platforms across development, test, staging, and production environments.
Ensure platforms meet agreed availability, performance, capacity, and service-level requirements.
Perform daily health checks and operational monitoring of infrastructure, applications, databases, middleware, and supporting services.
Manage platform configuration, patching, upgrades, maintenance, and lifecycle activities.
Support the deployment and release of applications and platform components into production environments.
Maintain accurate operational documentation, configuration records, and runbooks.
Identify opportunities to improve platform stability, resilience, and operational efficiency.
Incident & Problem Management
Monitor platform alerts and respond to incidents in a timely manner.
Perform troubleshooting and root-cause analysis for infrastructure and application platform issues.
Coordinate with application, network, database, cybersecurity, and vendor teams to resolve complex incidents.
Participate in major incident management and service restoration activities.
Conduct post-incident reviews and implement preventive and corrective actions.
Support problem management and identify recurring issues requiring permanent remediation.
Cloud & Infrastructure Operations
Operate and support cloud and/or on-premises infrastructure environments.
Manage compute, storage, networking, containers, virtual machines, and platform services.
Support cloud environments such as AWS, Microsoft Azure, or Google Cloud, where applicable.
Monitor resource utilisation, capacity, availability, and performance.
Support backup, recovery, disaster recovery, and business continuity activities.
Assist with infrastructure provisioning and configuration using Infrastructure as Code where applicable.
Automation & DevOps
Develop scripts and automation to reduce manual operational activities.
Support CI/CD pipelines and automated deployment processes.
Implement Infrastructure as Code using technologies such as Terraform, Ansible, or equivalent tools.
Automate platform monitoring, health checks, patching, configuration management, and operational reporting.
Work with development and DevOps teams to improve deployment reliability and operational processes.
Promote standardisation and repeatable operational practices across environments.
Security & Compliance
Operate platforms in accordance with applicable Singapore Government cybersecurity, technology, data protection, and operational policies and standards.
Implement and maintain appropriate access controls, privileged access, logging, monitoring, and security configurations.
Support vulnerability remediation, security patching, hardening, and security assessments.
Monitor and investigate security-related alerts in collaboration with cybersecurity teams.
Ensure operational activities are properly documented and auditable.
Support compliance reviews, security audits, and technology risk assessments.
Monitoring & Performance Management
Configure and maintain monitoring, alerting, logging, and observability solutions.
Monitor system availability, performance, capacity, and resource utilisation.
Develop operational dashboards and reports to provide visibility into platform health.
Analyse trends and proactively identify potential performance or capacity issues.
Support application performance monitoring and end-to-end service monitoring.
Stakeholder & Vendor Management
Work closely with application teams, developers, cybersecurity teams, network engineers, database administrators, and other technology teams.
Coordinate with external vendors and managed service providers for operational support and issue resolution.
Participate in technical discussions, change reviews, maintenance planning, and service improvement initiatives.
Communicate operational issues, risks, and recommendations clearly to technical and non-technical stakeholders.
Support procurement, technical evaluation, and vendor performance management where required.
Requirements
Essential
Degree or diploma in Computer Science, Information Technology, Engineering, or a related discipline.
3–6 years of relevant experience in platform operations, infrastructure operations, DevOps, systems administration, cloud operations, or production support.
Hands-on experience supporting production IT environments.
Strong understanding of Linux and/or Windows server environments.
Good understanding of networking concepts, including TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls.
Experience with monitoring, logging, incident management, and troubleshooting.
Experience with scripting or automation using Python, PowerShell, Bash, or similar technologies.
Understanding of IT service management practices, including incident, problem, change, and service request management.
Strong analytical and troubleshooting skills.
Good communication and stakeholder management skills.
Ability to participate in operational support and scheduled maintenance activities where required.
Good to Have
Experience operating AWS, Microsoft Azure, or Google Cloud environments.
Experience with Kubernetes, Docker, or other container technologies.
Experience with CI/CD tools such as Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, or equivalent.
Experience with Infrastructure as Code, such as Terraform or Ansible.
Experience with observability tools such as Prometheus, Grafana, ELK/OpenSearch, Splunk, or equivalent.
Experience with API gateways, application servers, web servers, or middleware platforms.
Experience with high-availability and disaster-recovery architectures.
Knowledge of cybersecurity principles, vulnerability management, system hardening, and privileged access management.
ITIL certification or equivalent experience.
Relevant cloud or infrastructure certifications.