GSPANN is hiring a Storage/Backup & Recovery Administrator to manage enterprise storage and backup platforms, driving automation, patching, and disaster-recovery readiness across Linux environments.
Description
Roles and Responsibilities
• Drive automation as a core responsibility, not an optional add-on, with strong hands-on Shell/Bash scripting and Ansible for configuration management, patching, and repetitive operational tasks.
• Build and maintain Ansible playbooks and roles for provisioning, patching, hardening, and routine operational workflows.
• Develop Python scripts to extend automation and tooling capabilities where beneficial.
• Drive an automation and AI mindset by proactively using modern AI-assisted tools for scripting, troubleshooting, log analysis, and documentation to improve speed and quality of resolution.
• Build automated, permanent fixes for recurring issues to reduce repeat incidents and manual effort.
• Manage and support enterprise SAN/NAS storage platforms with meaningful, hands-on experience; deep expertise across every storage vendor or product is not required.
• Support storage provisioning, including creating and managing LUNs, storage pools, RAID groups, volumes, and snapshots.
• Ensure storage capacity, utilization, performance, and availability are continuously monitored.
• Drive resolution of storage connectivity issues involving Fibre Channel (FC), iSCSI, multipathing, and zoning.
• Support storage migrations and expansion activities.
• Manage enterprise backup solutions such as Veritas NetBackup, Commvault, Veeam, or similar, with practical hands-on experience on at least one major platform.
• Support backup monitoring, restoration requests, disaster recovery validations, and troubleshooting.
• Ensure compliance with backup policies and recovery objectives.
• Manage the end-to-end patching lifecycle, including planning, testing, scheduling, and execution, across Linux estates.
• Drive vulnerability remediation in coordination with security teams, tracking findings through to closure within SLA.
• Implement and maintain security hardening baselines and compliance standards, such as CIS benchmarks, across Linux systems.
• Support audit and compliance reporting for patch and vulnerability status.
• Ensure proactive monitoring of systems and storage to identify risks and degradation before they cause incidents.
• Lead root-cause analysis (RCA) for critical incidents and drive permanent, automated fixes rather than workarounds.
• Drive recurring-issue reduction through problem management and process improvement.
• Optimize monitoring, automation, and operational runbooks to raise efficiency and reliability.
• Manage shift operations as Shift Lead or senior operational point of contact for Linux and storage operations.
• Lead complex incident triage, root-cause analysis, and resolution activities.
• Collaborate with cross-functional teams, vendors, and customers during major incidents.
• Support on-call rotations and shift schedules.
• Develop incident reports, technical documentation, SOPs, and knowledge articles.
Skills and Experience
• Hold a NetApp, Dell EMC, or other storage-vendor certification, or an ITIL Foundation certification, which is an advantage.
• Apply hands-on Shell/Bash scripting and Ansible for configuration management and automation.
• Work with Git and version control for scripts and automation content; Python scripting is an advantage.
• Demonstrate hands-on, practical experience with SAN and NAS technologies, Fibre Channel and iSCSI environments, and storage provisioning and capacity management.
• Bring exposure to one or more storage platforms such as EMC Dell, NetApp, Hitachi, IBM, or Pure Storage; broad multi-vendor depth is not required.
• Apply practical experience with at least one enterprise backup tool, such as Veritas NetBackup, Commvault, Veeam, or Backup Exec.
• Bring working knowledge of patching, vulnerability remediation workflows, and security hardening practices.
• Demonstrate familiarity with compliance frameworks and benchmarks such as CIS for Linux systems.
• Apply an understanding of ITIL processes, including Incident, Problem, and Change Management.
• Bring experience supporting large-scale enterprise production environments with shift and on-call responsibilities.
• Demonstrate strong analytical, troubleshooting, and communication skills.
• Bring experience leading technical bridge calls and handling critical incidents.
• Utilize AI-assisted tools to speed up scripting, troubleshooting, log analysis, and documentation.
• Work with Docker and Kubernetes, which is an advantage.
• Bring exposure to Linux integration with Active Directory (SSSD, Kerberos, keytabs), which is an advantage.