Talent.com
GENESIS NETWORKS PTE LTD
AI Infrastructure EngineerGENESIS NETWORKS PTE LTD • D14 Geylang, Eunos, SG
AI Infrastructure Engineer

AI Infrastructure Engineer

GENESIS NETWORKS PTE LTD • D14 Geylang, Eunos, SG
15 days ago
Job description

Roles & Responsibilities

1. Compute & ClusterManagement

  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g., NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.

2. High-Performance Networking& Storage

  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g., Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.

3. Automation &Infrastructure as Code (IaC)

  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.

4. Operations, Observability& Performance

  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.

Qualifications &Requirements

Technical Competencies

  • Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
  • High-Speed Networking: Proven experience with RDMA (RoCE v2 / InfiniBand), PFC (Priority Flow Control), and ECN configurations.
  • Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).

Experience & Education

  • Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.
  • 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).


Tell employers what skills you have

Monitor Network Performance
Timely Execution
Kubernetes
Utilization Management
Architect
Root Cause Analysis
Container Orchestration
Disaster Recovery Plans
Mental Ray
Online Monitoring
Configuration Control
Telemetry
Distributed Algorithms
Grafana
Metrics Dashboard

Create a job alert for this search

AI Infrastructure Engineer • D14 Geylang, Eunos, SG

Similar jobs

Lead AI Engineer

MasterCardSingapore, Central Singapore District, SG

Mastercard powers economies and empowers people in 200+ countries and territories worldwide.Together with our customers, we’re helping build a sustainable economy where everyone can prosper.We supp... Show more

Senior AI Engineer

MasterCardSingapore, Central Singapore District, SG

Mastercard powers economies and empowers people in 200+ countries and territories worldwide.Together with our customers, we’re helping build a sustainable economy where everyone can prosper.We supp... Show more