Talent.com
OPENSOURCE TECHNOLOGIES PTE. LTD.
Kubernetes & Site Reliability Engineer (SRE)OPENSOURCE TECHNOLOGIES PTE. LTD. • Islandwide, SG
Search for other jobs
Kubernetes & Site Reliability Engineer (SRE)

Kubernetes & Site Reliability Engineer (SRE)

OPENSOURCE TECHNOLOGIES PTE. LTD. • Islandwide, SG
8 days ago
Job description

Roles & Responsibilities

Role Overview

  • We are looking for experienced Kubernetes & Site Reliability Engineers to support highly scalable, business-critical production platforms for a global technology customer in Singapore.
  • The role requires strong hands-on expertise in Kubernetes, Linux, production reliability, automation, observability, incident management and troubleshooting of distributed systems
  • Candidates should be comfortable operating large-scale production environments where availability, performance, automation and operational excellence are critical.

Key Responsibilities

  • Operate, maintain and troubleshoot large-scale Kubernetes-based production environments
  • Ensure reliability, scalability, availability and performance of critical services.
  • Investigate complex production issues and perform detailed root-cause analysis.
  • Participate in incident response and drive permanent corrective actions.
  • Automate repetitive operational activities and improve platform reliability.
  • Build and improve monitoring, alerting, logging and observability frameworks.
  • Define and track SLIs, SLOs and operational reliability metrics
  • Support Kubernetes upgrades, configuration changes, patching and platform improvements.
  • Work closely with application engineering, infrastructure, platform, security and DevOps teams.
  • Perform capacity planning, performance tuning and reliability improvements.
  • Develop and maintain operational runbooks, automation scripts and troubleshooting documentation.
  • Participate in production readiness reviews and ensure applications meet operational standards.

Mandatory Skills

  • Strong hands-on experience with
  • Kubernetes administration and troubleshooting
  • Strong understanding of Kubernetes architecture, including:
  • Pods
  • Deployments
  • StatefulSets
  • Services
  • Ingress
  • ConfigMaps / Secrets
  • RBAC
  • Storage
  • Networking
  • Strong
  • Linux systems administration and troubleshooting skills.
  • Good understanding of networking concepts such as DNS, TCP/IP, load balancing and service connectivity.
  • Strong understanding of
  • Site Reliability Engineering principles
  • Experience supporting large-scale, high-availability production systems.
  • Strong incident management and RCA experience.
  • Hands-on scripting/automation experience using
  • Python, Bash/Shell or similar
  • Experience with monitoring and observability tools such as
  • Prometheus, Grafana, Splunk, ELK/OpenSearch, Datadog or equivalent
  • Helm or similar Kubernetes package/deployment management tools.
  • GitOps experience using tools such as Argo CD or Flux.
  • Knowledge of service mesh concepts.
  • Experience with container security and Kubernetes security practices.
  • Experience with cloud or private-cloud infrastructure.
  • Familiarity with distributed systems and microservices architectures.
  • Exposure to performance engineering and capacity management.
  • Experience working in globally distributed engineering environments.


Tell employers what skills you have

Operational Efficiency
RDS
systems reliability
Oracle Alerts
Scalability
Automated Operation Monitoring
Ubuntu
Patch Management
Reliability Requirements
Root Cause Analysis
Availability Management
Technical Consultation
Routing Protocols
Incident Handling
Configuration Changes
C++

Create a job alert for this search

Kubernetes & Site Reliability Engineer (SRE) • Islandwide, SG