GSPANN is hiring a Datadog Platform & Observability Architect to design enterprise monitoring solutions, reducing incidents and strengthening visibility across applications and infrastructure.
Description
Roles and Responsibilities
- Design, configure, and maintain Datadog monitors, dashboards, alerts, and integrations.
- Build observability coverage across infrastructure, application, and network layers.
- Implement observability solutions for Enterprise Resource Planning (ERP) and e-commerce applications.
- Drive adoption of Datadog's AI-driven modules for automated issue detection and prevention.
- Support Infrastructure Monitoring, APM, Log Management, Synthetic Monitoring, RUM, and API monitoring across the Datadog platform.
- Manage application and infrastructure telemetry, including metrics, logs, traces, and events.
- Optimize Datadog log pipelines, including parsing rules, indexes, retention filters, and enrichment.
- Ensure Datadog Agents are properly configured and troubleshoot issues across Windows, Linux, cloud, containers, and Kubernetes environments.
- Develop application-specific dashboards that provide end-to-end visibility into application health and performance.
- Build proactive alerts for critical applications, APIs, infrastructure components, and business services.
- Optimize existing monitors by continuously improving alert thresholds, reducing noise and duplicate alerts, and refining actionable notifications.
- Support application teams with troubleshooting using logs, traces, metrics, and dependency information.
- Drive monitoring gap assessments and recommend improvements.
- Implement integration of Datadog with IT Service Management (ITSM) and collaboration platforms such as Ivanti, ServiceNow, Microsoft Teams, and Slack.
- Support automation and event-driven incident creation and enrichment.
- Collaborate with DevOps and SRE teams to implement monitoring-as-code using Terraform or similar tools.
- Support major incident troubleshooting and RCA / problem-management activities.
- Ensure Standard Operating Procedures (SOPs), monitoring standards, technical documentation, and knowledge articles stay current.
- Drive knowledge-transfer sessions to help support teams use Datadog effectively for incident triage.
- Ensure monitoring effectiveness is tracked, demonstrating measurable improvements such as reduced alert noise, faster detection, lower Mean
- Time to Recovery (MTTR), and fewer repeat incidents.
Skills and Experience
- Bring strong hands-on experience with Datadog.
- Manage Datadog Infrastructure Monitoring and Agent configuration.
- Demonstrate expertise in APM and Distributed Tracing.
- Work with Log Management and Log Pipelines.
- Manage Datadog dashboards, monitors, alerts, and notification configuration.
- Apply Synthetic and API monitoring techniques.
- Bring experience troubleshooting Windows and Linux systems.
- Demonstrate good understanding of cloud platforms, preferably Azure and/or AWS.
- Apply knowledge of containers and Kubernetes monitoring.
- Demonstrate understanding of REST APIs, JSON, and integration concepts.
- Bring scripting knowledge in Python, PowerShell, Bash, or similar languages.
- Develop Infrastructure-as-Code using Terraform, which is an advantage.
- Demonstrate strong troubleshooting and root-cause analysis skills.
- Bring experience with Datadog RUM and Digital Experience Monitoring.
- Utilize network monitoring tools and techniques.
- Apply Azure Monitor and Log Analytics knowledge.
- Demonstrate exposure to application performance optimization and capacity planning.
- Bring Workato or other integration and automation platform experience, which is an advantage.
- Develop SRE and Observability implementation experience, which is an advantage.
- Utilize Continuous Integration / Continuous Delivery (CI/CD) monitoring experience, which is an advantage.
- Demonstrate Information Technology Infrastructure Library (ITIL) knowledge covering Incident, Problem, and Change Management, which is an advantage