GSPANN is hiring a Systems Engineer with expertise in Linux Administration to manage enterprise Linux environments, driving automation, patching, and security hardening while leading incident resolution.
Description
Roles and Responsibilities
• Support enterprise Linux environments (RHEL, CentOS, Oracle Linux, Ubuntu) through administration, monitoring, and troubleshooting.
• Implement OS installation, configuration, patching, upgrades, and hardening activities.
• Manage user accounts, permissions, file systems, services, kernel parameters, and performance tuning.
• Optimize system performance by resolving issues related to CPU, memory, storage, networking, and applications.
• Drive root-cause isolation across cross-domain issues spanning Linux, storage, virtualization, network, and application dependencies.
• Manage packages and repositories using patch and repository management tooling such as Satellite, Spacewalk, and YUM/DNF/APT repos.
• Support Git and version control workflows for scripts, automation content, and configuration management.
• Drive automation as a core responsibility, not an optional add-on, with strong hands-on Shell/Bash scripting and Ansible for configuration management, patching, and repetitive operational tasks.
• Build and maintain Ansible playbooks and roles for provisioning, patching, hardening, and routine operational workflows.
• Develop Python scripts to extend automation and tooling capabilities where beneficial.
• Drive an automation and AI mindset by proactively using modern AI-assisted tools for scripting, troubleshooting, log analysis, and documentation to improve speed and quality of resolution.
• Build automated, permanent fixes for recurring issues to reduce repeat incidents and manual effort.
• Manage the end-to-end patching lifecycle, including planning, testing, scheduling, and execution, across Linux estates.
• Drive vulnerability remediation in coordination with security teams, tracking findings through to closure within Service Level Agreement (SLA) timelines.
• Implement and maintain security hardening baselines and compliance standards, such as Center for Internet Security (CIS) benchmarks, across Linux systems.
• Support audit and compliance reporting for patch and vulnerability status.
• Ensure proactive monitoring of systems and storage to identify risks and degradation before they cause incidents.
• Lead root-cause analysis (RCA) for critical incidents and drive permanent, automated fixes rather than workarounds.
• Drive recurring-issue reduction through problem management and process improvement.
• Optimize monitoring, automation, and operational runbooks to raise efficiency and reliability.
• Manage shift operations as Shift Lead or senior operational point of contact for Linux and storage support.
• Lead complex incident triage, root-cause analysis, and resolution activities.
• Collaborate with cross-functional teams, vendors, and customers during major incidents.
• Support on-call rotations and shift schedules.
• Develop incident reports, technical documentation, SOPs, and knowledge articles.
Skills and Experience
• Hold RHCSA, RHCE, a Red Hat Specialist certification, or ITIL Foundation certification, which is an advantage.
• Demonstrate strong expertise in RHEL, CentOS, Oracle Linux, Ubuntu, system performance tuning, filesystem management (XFS, EXT4, LVM), and patch and repository management tooling.
• Apply hands-on Shell/Bash scripting and Ansible for configuration management and automation.
• Work with Git and version control for scripts and automation content; Python scripting is an advantage.
• Bring working knowledge of patching, vulnerability remediation workflows, and security hardening practices.
• Demonstrate familiarity with compliance frameworks and benchmarks such as CIS for Linux systems.
• Apply an understanding of ITIL processes, including Incident, Problem, and Change Management.
• Bring experience supporting large-scale enterprise production environments with shift and on-call responsibilities.
• Demonstrate strong analytical, troubleshooting, and communication skills.
• Bring experience leading technical bridge calls and handling critical incidents.
• Utilize AI-assisted tools to speed up scripting, troubleshooting, log analysis, and documentation.
• Work with Docker and Kubernetes, which is an advantage.
• Bring exposure to Linux integration with Active Directory (System Security Services Daemon (SSSD), Kerberos, keytabs), which is an advantage.