Role Overview
As an Operations Excellence Engineer, you will play a key role in sustaining and enhancing our data center operations. Based at our Dammam Office in KSA, you'll work in an ONSITE environment, driving operational excellence initiatives across multi-site critical infrastructure environments. You will bring expertise in robust operational governance, data center reliability principles, and continuous improvement, ensuring outstanding availability and compliance in live facility environments.
Responsibilities
- Implement and enforce Uptime Institute Tier principles including availability, redundancy, and fault tolerance across data center operations.
- Develop, review, and enhance data center operational governance encompassing policies, SOPs, MOPs, and operational control frameworks.
- Lead operations assurance activities, implement audit methodologies, and conduct compliance gap assessments and non-conformance management.
- Manage site takeover and transition activities, including readiness assessments, cutover planning, and shadowing/reverse-shadowing models in live environments.
- Coordinate incident management frameworks, escalation models, incident classification, and validate root cause analyses and corrective actions.
- Implement reliability-centered maintenance (RCM) principles, preventive/corrective maintenance governance, and maintenance risk management.
- Oversee change management and operational risk control processes including impact assessment, rollback planning, and approval workflows in live data center operations.
- Assess electrical systems (UPS, generators, switchgear, ATS/STS, power distribution) for operational and maintenance risk.
- Evaluate mechanical systems (chillers, CRAH/CRAC units, cooling systems, pumps) for operational performance and reliability.
- Oversee ELV and monitoring systems including BMS, EPMS, DCIM, fire safety, access control, and alarm management practices.
- Prepare operational performance reports, define KPIs, perform trend analysis, and report on availability, incidents, and compliance.
- Ensure audit readiness with robust evidence management, maintaining complete maintenance records, incident logs, and compliance documentation.
- Champion HSE best practices and safe systems of work, including permit-to-work, lockout/tagout, and contractor safety management in critical environments.
- Coordinate standardization and alignment of operational practices across regions to meet group operational standards.
Must have requirements
- Minimum 5 years of experience in data center operations, facility management, or operational excellence roles, with exposure to multi-site operations. Proven experience in operations assurance, audits, standardization, incident management, and governance within mission-critical environments.
- Bachelor’s or masters degree in Electrical Engineering, Mechanical Engineering, Facilities Engineering, or equivalent.
- Deep familiarity with Uptime Institute Tier principles: availability, redundancy, fault tolerance, concurrent maintainability, and avoidance of single points of failure.
- Hands-on experience in developing and implementing data center policies, SOPs, MOPs, work instructions, and operational control frameworks.
- Solid background in operations assurance and auditing, including gap analysis, compliance reviews, and non-conformance management.
- Demonstrated experience in site takeover and transition management with practical application of shadowing/reverse-shadowing, cutover planning, and post-takeover live environment stabilization.
- Proven ability to execute incident management frameworks and escalation models, and coordinate live event response with validated root cause analysis and corrective actions.
- Expertise in reliability-centered maintenance (RCM) concepts, preventive and corrective maintenance processes, and maintenance risk management.
- Thorough understanding of change management and operational risk control in live data center settings, including rollback planning and approval workflows.
- Strong knowledge of data center electrical systems (UPS, generators, MV/LV switchgear, ATS/STS, power distribution) for risk assessment.
- Solid grounding in data center mechanical systems (chillers, CRAH/CRAC, cooling systems, pumps) for performance and reliability evaluation.
- Knowledge of ELV and monitoring systems (BMS, EPMS, DCIM, fire detection/suppression, access control, alarm management).
- Experience in performance reporting, KPI definition, trend analysis, and management reporting for availability, incidents, and compliance.
- Demonstrated audit readiness and evidence management for maintenance records, incident logs, and compliance documentation.
- Comprehensive knowledge of HSE requirements and experience managing permit-to-work, lockout/tagout, and contractor safety programs in critical environments.
- Experience coordinating multi-site operations and aligning operational practices with group-wide standards.
Nice to have requirements
- Lean Six Sigma certification (Green Belt or higher).
- Industry-recognized data center certifications such as CDCP, CDCS, or equivalent.
- Relevant ISO certifications, with hands-on experience in leading ISO 9001 audit processes.