GSPANN is hiring a Senior Dynatrace Platform Administrator to manage enterprise observability, covering platform administration, monitoring configuration, and reliability engineering across cloud environments.
Description
Roles and Responsibilities
- Manage administration, configuration, governance, and optimization of the Dynatrace SaaS platform across enterprise applications, infrastructure, cloud, and containerized environments.
- Support deployment, upgrades, configuration, and troubleshooting of Dynatrace OneAgent and ActiveGate to ensure comprehensive monitoring coverage.
- Implement Alerting Profiles, Problem Notifications, Metric Events, Anomaly Detection, and Davis AI-based alerting capabilities.
- Design and maintain Browser and HTTP Synthetic Monitoring solutions to proactively identify application performance issues, and extend Real User Monitoring (RUM) capabilities for visibility into end-user experience.
- Build Management Zones, tagging strategies, auto-tagging policies, naming conventions, and monitoring governance standards.
- Develop and maintain Dynatrace Dashboards, Notebooks, Workflows, and advanced Dynatrace Query Language (DQL)-based analytics using Grail.
- Lead onboarding of new applications, Application Programming Interfaces (APIs), cloud resources, infrastructure components, and business services into the Dynatrace monitoring ecosystem.
- Manage service detection, process grouping, dependency mapping, application topology, service flow, distributed tracing, and PurePath diagnostics to identify performance bottlenecks.
- Optimize monitoring integrations across Microsoft Azure, Amazon Web Services (AWS), and hybrid cloud environments.
- Implement log ingestion, log monitoring, custom metrics, business event monitoring, and enterprise observability solutions.
- Support troubleshooting, performance optimization, capacity planning, and operational insights using Data Explorer, Grail, and DQL.
- Drive platform health assessments and adoption of modern Dynatrace capabilities, including Grail, DQL, Notebooks, Workflows, and next-generation dashboard frameworks, to advance enterprise monitoring and observability best practices.
- Ensure monitoring coverage and observability readiness for new releases, deployments, infrastructure changes, and cloud migrations.
- Manage maintenance windows and monitoring policies to support planned outages and maintenance activities.
- Ensure secure user access through Role-Based Access Control (RBAC) roles, permissions, licensing, compliance requirements, and platform security standards.
- Build integrations between Dynatrace and ServiceNow, Microsoft Teams, and other enterprise Information Technology Service Management (ITSM), collaboration, and notification platforms.
- Support incident management, service restoration, Major Incident processes, Root Cause Analysis (RCA), and continuous reliability improvement initiatives.
- Ensure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) are defined, implemented, and monitored to improve platform reliability and operational excellence.
- Develop automation solutions and collaborate closely with Application, Infrastructure, Cloud, Security, and DevOps teams to improve operational efficiency.
- Drive observability, monitoring maturity, and Site Reliability Engineering (SRE) initiatives across the organization.
Skills and Experience
- Bring hands-on expertise administering the Dynatrace SaaS platform, with a background in Observability Engineering, Application Performance Monitoring (APM), or Site Reliability Engineering (SRE).
- Demonstrate strong experience deploying, configuring, troubleshooting, and maintaining Dynatrace OneAgent and ActiveGate.
- Apply expertise across Synthetic Monitoring, RUM, Application Monitoring, Infrastructure Monitoring, and Cloud Monitoring.
- Bring strong working knowledge of Grail, DQL, Data Explorer, Dashboards, Notebooks, and Workflows for advanced analytics.
- Manage configuration of Alerting Profiles, Problem Detection, Metric Events, Anomaly Detection, and Davis AI-based alerting.
- Apply governance discipline across Management Zones, tagging strategies, naming rules, and auto-tagging policies.
- Work with distributed tracing, service flow analysis, application topology mapping, and PurePath diagnostics to resolve complex performance issues.
- Develop integrations between Dynatrace and cloud platforms such as Azure and AWS, plus enterprise tools such as ServiceNow and Microsoft Teams.
- Demonstrate strong understanding of log monitoring, log ingestion, custom metrics, business events, and enterprise observability.
- Bring strong grounding in Site Reliability Engineering (SRE) practices, including incident management, problem management, RCA, SLIs, SLOs, and SLAs.
- Utilize automation, scripting, and DevOps practices to improve monitoring efficiency and platform operations.
- Apply excellent troubleshooting, analytical, and problem-solving skills.
- Demonstrate strong communication and stakeholder management skills, working effectively across cross-functional teams.
- Bring experience supporting large-scale, mission-critical enterprise applications and infrastructure environments.
- Demonstrate knowledge of Information Technology Infrastructure Library (ITIL) processes and enterprise monitoring best practices, which is an advantage.
- Work with fast-paced production environments, supporting critical business applications.