Job Summary
We are seeking an experienced Linux SME to provide expert-level support, governance, and optimization of enterprise Linux environments within a Managed Services framework. The role involves ensuring SLA-driven operations, high availability, security, and compliance while driving automation and continual service improvement.
Key Responsibilities
● Act as the Linux Subject Matter Expert within the Managed Services team for escalations, operations, and technical advisory.
● Provide L3/L4 support for complex Linux-related incidents, problems, and changes across multiple distributions (RHEL, CentOS, Ubuntu, SUSE).
● Manage day-to-day Linux operations including monitoring, patching, OS upgrades, performance tuning, and troubleshooting.
● Ensure security hardening, compliance (CIS, PCI, SOX, ISO), and adherence to managed services baselines.
● Automate recurring administration tasks (patching, user management, monitoring) using scripting (Bash, Python) and tools (Ansible, Puppet, Chef).
● Support backup, DR, and high availability strategies for critical Linux workloads.
● Collaborate with AWS, Windows, Network, and Security SMEs to deliver integrated managed services.
● Conduct root cause analysis (RCA) for high-severity issues and present service improvement recommendations.
● Contribute to Service Improvement Plans (SIPs) and Continuous Service Improvement (CSI) initiatives.
● Provide documentation, knowledge base articles, and operational runbooks for L1/L2 teams.
● Mentor and train junior engineers to strengthen Linux operational capability in the Managed Services team.
Required Skills & Qualifications
● 8–12 years of IT experience with at least 5+ years in Linux administration in enterprise environments.
● Deep expertise in Linux OS (RHEL, CentOS, Ubuntu, SUSE) and system internals.
● Strong experience in day-to-day managed services operations (incident, problem, change, patch management).
● Hands-on automation experience using Ansible, Puppet, Chef, or similar.
● Proficiency in scripting (Bash, Python, Perl).
● Strong troubleshooting skills in multi-tier, distributed, and hybrid environments.
● Experience with monitoring and observability (Nagios, Zabbix, Prometheus, Grafana, ELK, Splunk).
● Knowledge of virtualization (VMware, KVM) and Linux on cloud platforms (AWS, Azure, GCP).
● Familiarity with ITIL processes and managed services delivery.
● Excellent communication and customer-handling skills for escalations and service reviews.
Preferred Qualifications
● RHCE / RHCSA / LFCS certifications.
● Knowledge of containers (Docker, Podman) and orchestration (Kubernetes, OpenShift).
● Experience in hybrid cloud Linux operations and multi-tenant managed environments.
● Exposure to security/compliance frameworks (PCI, SOX, ISO 27001).
● Experience with service governance, audits, and RCA preparation for customers