Cloud Support

ARISTON SERVICES PTE. LTD.

Senior Ll /L2 Support Engineer

Role Overview

  • We are seeking a highly motivated and experienced Senior Ll/L2 Support Engineer tojoin the TechOps team and provide operational support for digital platforms, e-commerce applications, and customer-facing systems.The successful candidate will serve as the primary point of contact for productionincidents, service requests, and operational issues, ensuring timely resolution and highplatform availability. This role requires strong troubleshooting skills, cloud operationsknowledge, and an inquisitive mindsetto investigate incidents beyond surfacesymptoms and identify root causes.The candidate will work closely with developers, infrastructure teams, externalvendors, and business stakeholders across multiple regions and time zones tomaintain stable, secure, and reliable digital services in a 24x7 support environment.Key ResponsibilitiesProduction Support & Incident ManagementAct as the primary Ll /L2 support contact for digital platforms, e-commercesystems, websites, and customer-facingMonitor incident queues, service requests, alerts, and support tickets, ensuringadherence to SLAS and operational procedures.Lead incident triage, troubleshooting, escalation, and resolution activities.Perform impact assessment and coordinate with relevant stakeholders duringservice disruptions.Support incident management activities and facilitate communication duringcritical outages.Conduct post-incident reviews and root cause analysis (RCA) to preventrecurrence.Develop and maintain operational runbooks, support procedures, andknowledge base articles.System Monitoring & ReliabilityMonitor application, infrastructure, and business service health usingobservability and monitoring tools.Analyse system performance, availability, error trends, and capacity utilization.Configure and tune alerts to reduce noise and improve operational visibility.Collaborate with engineering teams to improve system reliability andoperational resilience.Support Site Reliability Engineering (SRE) practices, including reliability metrics,incident reduction, and service availability improvements.Cloud & Infrastructure SupportProvide operational support for cloud-hosted applications and infrastructure,primarily on AWS.Perform first-level troubleshooting on:oooooCompute services (EC2, ECS, Lambda)NetworkingLoad BalancersCDN servicesStorage servicesInvestigate infrastructure-related issues affecting application perforrnancN oravailability.Support deployment verification and post-release monitoring activities.Application & Integration SupportTroubleshoot application issues across web, mobile, APIs, and middlewareplatforms.Analyse application logs, monitoring data, and system traces to identify rootcauses.
  • Support integrations with external systems, partners, payment gateways, andother enterprise platforms.Work closely with L3 to reproduce issues and validate fixes.Support release and deployment activities, including late-night and weekendimplementations when required.Continuous ImprovementIdentify recurring incidents and operational inefficiencies.Drive automation opportunities to reduce manual effort and repetitive supportactivities.Recommend improvements to monitoring, alerting, deployment processes, andsupport workflows.Contribute to operational excellence initiatives and service reliabilityimprovements.Stakeholder & Vendor ManagementCollaborate with internal teams, external vendors, and partners across differentgeographies and time zones.Communicate effectively with technical and non-technical stakeholders.Provide timely updates during incidents and service disruptions.Participate in operational reviews, governance meetings, and serviceimprovement discussions.Required Skills & ExperienceTechnical Skills5+ years of experience in Application Support, Production Support, TechnicalOperations, TechOps, or related roles.Strong knowledge of AWS cloud services and operational support.Good understanding of cloud-native application architecture.Experience supporting:o Digital platformso E-commerce systemso Customer-facing web applicationso Mobile applicationsKnowledge Of CDN technologies and content delivery architecture.Experience with monitoring and observability platforms such as:o Data dogo New Relico Dynatraceo AppDynamicso Grafanao CloudWatchFamiliaritywith APM (Application Performance Monitoring) concepts.Understanding Of SRE principles and operational best practices.Strong knowledge Of:oooooAPIsMicroservicesWeb servicesSystem integrationsAuthentication and authorization flowsExperience using ticketing and ITSM platforms (Jira Service Management,ServiceNow, etc.).Understanding of Incident, Problem, and Change Management processes.Familiaritywith log analysis tools such as Elasticsearch, Kibana, Splunk, orCloudWatch Logs.Soft SkillsStrong troubleshooting and analytical thinking abilities.Naturally curious and investigative, with a desire to understand the full contextbehind incidents and operational events.Excellent problem-solving and root cause analysis skills.Strong ownership mindset and accountability.Ability to work independently in a fast-paced operational environment.Good communication and stakeholder management skills.Ability to remain calm and methodical during high-severity incidents.Strong documentation and knowledge-sharingContinuous improvement mindsetwith a focus on automation and operationalefficiency.Good to Have SkillsKnowledge Of WeChat Mini Program ecosystem and integrations.
  • Experience supporting SAP Commerce, Adobe Experience Manager (AEM),Magento, Shopify, or similar e-commerce platforms.
  • Basic scripting skills (Python, Shell, Bash, PowerShell).Experience with API Gateway and event-drivenAWS Certifications (Cloud Practitioner, Solutions Architect Associate, SysopsAdministrator).ITIL Foundation certification.Experience supporting payment gateways and digital commerce ecosystems.Ability to communicate in Mandarin.Working ConditionsParticipate in a 24x7 support and on-call rotation model.
  • Support late-night, weekend, and public holiday deployments whereWork closely with internal teams, vendors, and stakeholders across multipletime zones.Respond to critical production incidents outside business hours whennecessary.Success ProfileThe ideal candidate is someone who:Loves troubleshooting and uncovering the root cause of complex issues.Thinks logically and analytically under pressure.Takes ownership from ticket creation through resolution.Continuously looks for ways to improve systems and eliminate recurringoperational issues.Balances technical depth with strong stakeholder communication.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.