AI Engineer (AWS, Splunk ITSI, Datadog, Python, New Relic, Azure, Dynatrace, ML, Service Now, Ansible, Rest API, ITIL)
EXASOFT CONSULTING PTE. LTD.
- Design, implement, and optimize enterprise observability solutions using Splunk ITSI.
- Architect and maintain observability solutions using Datadog for infrastructure, application, log, and service monitoring.
- Implement and optimize Dynatrace for application performance monitoring, infrastructure visibility, and service health analysis.
- Utilize New Relic for application monitoring, performance analysis, and proactive issue detection.
- Implement AppDynamics for application performance monitoring, transaction analysis, and application health management.
- Develop and enhance dashboards, KPIs, service views, alerts, and service-health monitoring for critical applications and infrastructure.
- Analyze logs, events, metrics, and telemetry to investigate incidents, identify anomalies, and support root-cause analysis.
- Implement intelligent event correlation, alert optimization, and noise-reduction techniques to improve operational efficiency.
- Apply AIOps, machine learning, and predictive analytics for anomaly detection, predictive monitoring, and proactive incident management.
- Integrate observability platforms with ServiceNow, CMDB, CI relationships, and service mapping to improve service visibility and dependency mapping.
- Design and implement observability solutions across AWS, Azure, and GCP cloud environments.
- Establish and monitor SLI/SLO metrics to improve service reliability and operational performance.
- Automate monitoring and operational processes using Python, Linux, Ansible, and REST APIs.
- Collaborate with infrastructure, application, cloud, and operations teams to improve service reliability and observability standards.
- 14+ years of experience in AIOps, observability, infrastructure monitoring, systems engineering, or related IT operations roles.
- Bachelor’s degree in computer science, Engineering, Information Technology, or equivalent.
- Strong hands-on experience with Splunk ITSI and enterprise observability platforms.
- Hands-on experience with Datadog, Dynatrace, New Relic, and/or AppDynamics.
- Experience in ServiceNow, CMDB, CI relationships, and service mapping.
- Experience with AIOps, machine learning, anomaly detection, event correlation, and predictive analytics.
- Experience in Python, Linux, Ansible, and REST APIs for automation and integration.
- Experience working with AWS, Azure, and/or GCP cloud environments.
- Strong understanding of ITIL, SRE, incident management, and SLI/SLO concepts.
- Strong analytical and troubleshooting skills using logs, metrics, events, and telemetry.
- Experience in designing and supporting enterprise-scale observability and monitoring platforms.
- Strong stakeholder management, technical leadership, and mentoring skills.