Job description
Roles & Responsibilities
Observability & Monitoring
- Implement, configure, and maintain enterprise monitoring and observability solutions.
- Configure monitoring for servers, network devices, applications, databases, storage, virtualization, and cloud infrastructure.
- Configure events, alerts, thresholds, dashboards, reports, and service health views.
- Work with monitoring platforms such as BMC, SolarWinds, Dynatrace, Datadog, Grafana, Prometheus, ELK, SCOM, or similar.
- Troubleshoot monitoring, event, performance, and integration issues.
ITSM & Service Management
- Implement and configure ITSM processes
- Configure workflows, forms, notifications, assignments, priorities, SLAs, and escalations.
- Integrate monitoring/event-management platforms with ITSM solutions for automated incident creation and updates.
- Support ITSM dashboards, reports, and operational processes.
Integration & Automation
- Support integrations using REST APIs, webhooks, SNMP, Syslog, agents, and other standard protocols.
- Develop basic automation scripts using Python, Bash, PowerShell, or similar languages.
- Support automation and configuration management using Ansible, Terraform, Jenkins, or similar tools.
- Assist with integrations between monitoring, ITSM, CMDB, cloud, automation, and third-party platforms.
Job Requirements
- 3–4 years of relevant experience in observability, enterprise monitoring, ITOM, ITSM, infrastructure operations, or a related field.
- Hands-on experience with at least one or more enterprise monitoring/observability platforms.
- Good understanding of monitoring, logging, alerting, event management, incident management, and service availability.
- Experience or exposure to ITSM, ITOM, CMDB, Discovery, Asset Management, and Service Mapping.
- Experience with one or more platforms such as BMC/BMC Helix, SolarWinds, Dynatrace, Datadog, Grafana, Prometheus, ELK/Elastic, SCOM, or ServiceNow.
- Good understanding of Linux/Windows operating systems and networking concepts.
- Basic knowledge of REST APIs and system integration.
- Basic scripting experience in Python, Bash, PowerShell, or similar.
- Exposure to AWS, Azure, or GCP and their monitoring capabilities.
- Exposure to Docker, Kubernetes, Ansible, Terraform, or Jenkins is an advantage.
- Strong troubleshooting, analytical, and problem-solving skills.
- Good communication and teamwork skills, with the ability to work in a customer-facing environment.
Tell employers what skills you have
Network Operating System
Asset Management
Splunk
Datadog
Itsm
CMDB
High Availability
Monitoring Solutions
Bash
Event Management
Bash/Shell/PowerShell
Dynatrace
Rest Apis
GCP
Grafana
Incident Management