Datadog Platform & Observability Architect

DatadogObservabilitySite Reliability Engineering (SRE)Application Performance Monitoring (APM)Distributed TracingLog ManagementMicrosoft AzureAmazon Web Services (AWS)KubernetesTerraformPythonPowerShell

Description

GSPANN is hiring a Datadog Platform & Observability Architect to design enterprise monitoring solutions, reducing incidents and strengthening visibility across applications and infrastructure.

Roles and Responsibilities

  • Design, configure, and maintain Datadog monitors, dashboards, alerts, and integrations.
  • Build observability coverage across infrastructure, application, and network layers.
  • Implement observability solutions for Enterprise Resource Planning (ERP) and e-commerce applications.
  • Drive adoption of Datadog's AI-driven modules for automated issue detection and prevention.
  • Support Infrastructure Monitoring, APM, Log Management, Synthetic Monitoring, RUM, and API monitoring across the Datadog platform.
  • Manage application and infrastructure telemetry, including metrics, logs, traces, and events.
  • Optimize Datadog log pipelines, including parsing rules, indexes, retention filters, and enrichment.
  • Ensure Datadog Agents are properly configured and troubleshoot issues across Windows, Linux, cloud, containers, and Kubernetes environments.
  • Develop application-specific dashboards that provide end-to-end visibility into application health and performance.
  • Build proactive alerts for critical applications, APIs, infrastructure components, and business services.
  • Optimize existing monitors by continuously improving alert thresholds, reducing noise and duplicate alerts, and refining actionable notifications.
  • Support application teams with troubleshooting using logs, traces, metrics, and dependency information.
  • Drive monitoring gap assessments and recommend improvements.
  • Implement integration of Datadog with IT Service Management (ITSM) and collaboration platforms such as Ivanti, ServiceNow, Microsoft Teams, and Slack.
  • Support automation and event-driven incident creation and enrichment.
  • Collaborate with DevOps and SRE teams to implement monitoring-as-code using Terraform or similar tools.
  • Support major incident troubleshooting and RCA / problem-management activities.
  • Ensure Standard Operating Procedures (SOPs), monitoring standards, technical documentation, and knowledge articles stay current.
  • Drive knowledge-transfer sessions to help support teams use Datadog effectively for incident triage.
  • Ensure monitoring effectiveness is tracked, demonstrating measurable improvements such as reduced alert noise, faster detection, lower Mean
  • Time to Recovery (MTTR), and fewer repeat incidents.

Skills and Experience

  • Bring strong hands-on experience with Datadog.
  • Manage Datadog Infrastructure Monitoring and Agent configuration.
  • Demonstrate expertise in APM and Distributed Tracing.
  • Work with Log Management and Log Pipelines.
  • Manage Datadog dashboards, monitors, alerts, and notification configuration.
  • Apply Synthetic and API monitoring techniques.
  • Bring experience troubleshooting Windows and Linux systems.
  • Demonstrate good understanding of cloud platforms, preferably Azure and/or AWS.
  • Apply knowledge of containers and Kubernetes monitoring.
  • Demonstrate understanding of REST APIs, JSON, and integration concepts.
  • Bring scripting knowledge in Python, PowerShell, Bash, or similar languages.
  • Develop Infrastructure-as-Code using Terraform, which is an advantage.
  • Demonstrate strong troubleshooting and root-cause analysis skills.
  • Bring experience with Datadog RUM and Digital Experience Monitoring.
  • Utilize network monitoring tools and techniques.
  • Apply Azure Monitor and Log Analytics knowledge.
  • Demonstrate exposure to application performance optimization and capacity planning.
  • Bring Workato or other integration and automation platform experience, which is an advantage.
  • Develop SRE and Observability implementation experience, which is an advantage.
  • Utilize Continuous Integration / Continuous Delivery (CI/CD) monitoring experience, which is an advantage.
  • Demonstrate Information Technology Infrastructure Library (ITIL) knowledge covering Incident, Problem, and Change Management, which is an advantage

Apply Now

PDF or DOCX up to 5MB