As an Operations Excellence Engineer, you will play a key role in sustaining and enhancing our data center operations. Based at our Head Office in Dubai, you'll work in an ONSITE environment, driving operational excellence initiatives across multi-site critical infrastructure environments. You will bring expertise in robust operational governance, data center reliability principles, and continuous improvement, ensuring outstanding availability and compliance in live facility environments.
Responsibilities
Implement and enforce Uptime Institute Tier principles including availability, redundancy, and fault tolerance across data center operations.
Develop, review, and enhance data center operational governance encompassing policies, SOPs, MOPs, and operational control frameworks.
Lead operations assurance activities, implement audit methodologies, and conduct compliance gap assessments and non-conformance management.
Manage site takeover and transition activities, including readiness assessments, cutover planning, and shadowing/reverse-shadowing models in live environments.
Coordinate incident management frameworks, escalation models, incident classification, and validate root cause analyses and corrective actions.
Oversee change management and operational risk control processes including impact assessment, rollback planning, and approval workflows in live data center operations.
Assess electrical systems (UPS, generators, switchgear, ATS/STS, power distribution) for operational and maintenance risk.
Evaluate mechanical systems (chillers, CRAH/CRAC units, cooling systems, pumps) for operational performance and reliability.
Oversee ELV and monitoring systems including BMS, EPMS, DCIM, fire safety, access control, and alarm management practices.
Prepare operational performance reports, define KPIs, perform trend analysis, and report on availability, incidents, and compliance.
Ensure audit readiness with robust evidence management, maintaining complete maintenance records, incident logs, and compliance documentation.
Champion HSE best practices and safe systems of work, including permit-to-work, lockout/tagout, and contractor safety management in critical environments.
Coordinate standardization and alignment of operational practices across regions to meet group operational standards.
Must have requirements
Minimum 5 years of experience in data center operations, facility management, or operational excellence roles, with exposure to multi-site operations. Proven experience in operations assurance, audits, standardization, incident management, and governance within mission-critical environments.
Bachelor’s or masters degree in Electrical Engineering, Mechanical Engineering, Facilities Engineering, or equivalent.
Deep familiarity with Uptime Institute Tier principles: availability, redundancy, fault tolerance, concurrent maintainability, and avoidance of single points of failure.
Hands-on experience in developing and implementing data center policies, SOPs, MOPs, work instructions, and operational control frameworks.
Solid background in operations assurance and auditing, including gap analysis, compliance reviews, and non-conformance management.
Demonstrated experience in site takeover and transition management with practical application of shadowing/reverse-shadowing, cutover planning, and post-takeover live environment stabilization.
Proven ability to execute incident management frameworks and escalation models, and coordinate live event response with validated root cause analysis and corrective actions.
Expertise in reliability-centered maintenance (RCM) concepts, preventive and corrective maintenance processes, and maintenance risk management.
Thorough understanding of change management and operational risk control in live data center settings, including rollback planning and approval workflows.
Strong knowledge of data center electrical systems (UPS, generators, MV/LV switchgear, ATS/STS, power distribution) for risk assessment.
Solid grounding in data center mechanical systems (chillers, CRAH/CRAC, cooling systems, pumps) for performance and reliability evaluation.
Knowledge of ELV and monitoring systems (BMS, EPMS, DCIM, fire detection/suppression, access control, alarm management).
Experience in performance reporting, KPI definition, trend analysis, and management reporting for availability, incidents, and compliance.
Demonstrated audit readiness and evidence management for maintenance records, incident logs, and compliance documentation.
Comprehensive knowledge of HSE requirements and experience managing permit-to-work, lockout/tagout, and contractor safety programs in critical environments.
Experience coordinating multi-site operations and aligning operational practices with group-wide standards.
Nice to have requirements
Lean Six Sigma certification (Green Belt or higher).
Industry-recognized data center certifications such as CDCP, CDCS, or equivalent.
Relevant ISO certifications, with hands-on experience in leading ISO 9001 audit processes.