GSPANN is hiring a Systems Engineer to manage enterprise Linux systems and storage, leading patching, backup, security hardening, and incident response as a shift lead.
Description
Roles and Responsibilities
- Manage enterprise Linux environments across Red Hat Enterprise Linux (RHEL), CentOS, Oracle Linux, and Ubuntu, including OS installation, configuration, patching, upgrades, hardening, user accounts, permissions, file systems, services, and kernel parameters.
- Drive troubleshooting of system and cross-domain issues spanning CPU, memory, storage, networking, virtualization, and application performance to isolate true root cause.
- Manage packages and repositories using tooling such as Satellite, Spacewalk, and YUM/DNF/APT repositories, and maintain scripts and automation content under Git version control.
- Build and maintain Shell/Bash scripts and Ansible playbooks and roles for configuration management, provisioning, patching, hardening, and routine operational workflows, treating automation as a core responsibility rather than an optional add-on.
- Implement AI-assisted tools and workflows for scripting, troubleshooting, log analysis, and documentation to improve the speed and quality of resolution.
- Drive the reduction of recurring issues by identifying root causes and building automated, permanent fixes that cut repeat incidents and manual effort.
- Manage enterprise SAN and NAS platforms, provisioning Logical Unit Numbers (LUNs), storage pools, Redundant Array of Independent Disks (RAID) groups, volumes, and snapshots, and monitor capacity, utilization, performance, and availability.
- Support storage migrations and expansion activities, and troubleshoot connectivity issues involving Fibre Channel (FC), Internet Small Computer Systems Interface (iSCSI), multipathing, and zoning.
- Manage enterprise backup solutions such as Veritas NetBackup, Commvault, or Veeam, with hands-on administration of at least one major platform.
- Ensure backup monitoring, restoration requests, disaster recovery validations, and compliance with backup policies and recovery objectives.
- Manage the end-to-end patching lifecycle, including planning, testing, scheduling, and execution, across Linux estates.
- Drive vulnerability remediation in coordination with security teams, tracking findings through to closure within Service Level Agreement (SLA) targets.
- Implement and maintain security hardening baselines and compliance standards such as Center for Internet Security (CIS) benchmarks across Linux systems.
- Support audit and compliance reporting for patch and vulnerability status.
- Optimize monitoring, automation, and operational runbooks to raise efficiency and reliability, proactively identifying risks and degradation before they cause incidents.
- Lead Root Cause Analysis (RCA) for critical incidents, driving permanent, automated fixes rather than workarounds and tracking recurring issues through problem management and process improvement.
- Manage shift-lead responsibilities for Linux and storage operations and support teams, leading complex incident triage, root cause analysis, and resolution activities.
- Collaborate with cross-functional teams, vendors, and customers to coordinate resolution during major incidents.
- Support on-call rotations and shift schedules to maintain continuous coverage.
- Ensure incident reports, technical documentation, Standard Operating Procedures (SOPs), and knowledge articles stay accurate and current.
Skills and Experience
- Hold certifications such as Red Hat Certified System Administrator (RHCSA), Red Hat Certified Engineer (RHCE), other Red Hat Specialist Certifications, NetApp, Dell EMC, or other storage vendor certifications, or Information Technology Infrastructure Library (ITIL) Foundation, which is an advantage.
- Bring deep hands-on expertise in RHEL, CentOS, Oracle Linux, and Ubuntu, including system performance tuning and filesystem management across XFS, EXT4, and Logical Volume Manager (LVM).
- Demonstrate proficiency with patch and repository management tooling for enterprise Linux environments.
- Apply hands-on Shell/Bash scripting and Ansible expertise for configuration management and automation, backed by Git version control for scripts and automation content.
- Develop automation using Python scripting, which is an advantage.
- Work with SAN and NAS technologies, FC and iSCSI environments, and storage provisioning and capacity management, with exposure to platforms such as Dell EMC, NetApp, Hitachi, IBM, or Pure Storage.
- Manage hands-on experience with at least one enterprise backup tool such as Veritas NetBackup, Commvault, or Veeam; experience with Backup Exec is an advantage.
- Bring working knowledge of patching, vulnerability remediation workflows, security hardening practices, and compliance frameworks such as CIS benchmarks for Linux systems.
- Demonstrate understanding of ITIL processes covering Incident, Problem, and Change Management.
- Bring experience supporting large-scale enterprise production environments with shift and on-call responsibilities.
- Demonstrate strong analytical, troubleshooting, and communication skills.
- Manage technical bridge calls and critical incident handling.
- Utilize AI-assisted tools to speed up scripting, troubleshooting, log analysis, and documentation.
- Work with Docker and Kubernetes, which is an advantage.
- Bring exposure to Linux integration with Active Directory, including System Security Services Daemon (SSSD), Kerberos, and keytabs, which is an advantage.