Site Reliability Engineer(Kubernetes, Dynatrace, Splunk, Docker, ELK, Terraform, JMeter, Prometheus, Grafana)
NEPTUNEZ SINGAPORE PTE. LTD.
Responsibilities
- Lead Site Reliability Engineering (SRE) activities to ensure high availability, stability, and reliability of enterprise applications and distributed systems.
- Provide technical leadership for L2/L3 production support, incident management, and operational excellence across mission-critical environments.
- Investigate and resolve complex production issues by analyzing application behavior, infrastructure, databases, and system performance.
- Monitor application and platform health using Dynatrace, Splunk, Prometheus, Grafana, and ELK, while building dashboards, alerts, and operational visibility.
- Perform root cause analysis (RCA) for critical incidents and drive permanent corrective and preventive actions.
- Troubleshoot Java/Spring Boot applications, APIs, Linux-based systems, and SQL-related issues to ensure rapid incident resolution.
- Support cloud-native environments across AWS and GCP, including Kubernetes-based container platforms and cloud infrastructure components.
- Develop and maintain automation solutions using Shell scripting, Python, and Terraform to improve operational efficiency and reduce manual effort.
- Support CI/CD and deployment activities using Azure DevOps, Jenkins, and modern DevOps practices.
- Collaborate with development, infrastructure, and operations teams to improve platform reliability, performance, and service availability.
- Mentor junior engineers, provide technical guidance, and promote SRE best practices, operational standards, and knowledge sharing.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- 12+ years of experience in Site Reliability Engineering (SRE), Production Support, or DevOps within enterprise environments.
- Strong experience in L2/L3 Production Support, Incident Management, RCA, and supporting distributed systems.
- Hands-on expertise with Dynatrace, Splunk, Prometheus, Grafana, and ELK for monitoring, observability, log analysis, and alerting.
- Strong proficiency in Linux administration, troubleshooting, and automation using Shell scripting or Python.
- Experience troubleshooting Java (Java 8+), Spring Boot applications, REST APIs, and application performance issues.
- Strong knowledge of SQL, database troubleshooting, query optimization, and performance tuning.
- Hands-on experience with AWS, GCP, Kubernetes, Docker, and cloud-native infrastructure.
- Experience implementing Infrastructure as Code (IaC) using Terraform.
- Good understanding of CI/CD pipelines, build and release management using Azure DevOps and Jenkins.
- Strong analytical, troubleshooting, and problem-solving skills with the ability to manage high-severity production incidents.
- Excellent communication, stakeholder management, and technical leadership skills, with experience mentoring engineering teams.
For employers only
Is this your company's job post? Verify ownership to manage this listing and receive applications directly.
Claim this listingLooking to apply for this job? Use the Apply button above.
See more jobs in Singapore, Singapore