Are you an experienced infrastructure and cloud professional looking to lead a high-performing operations team in a modern hybrid IT environment?
As our Cloud Operations Lead, you will be responsible for ensuring the stability, security, and performance of Epson Europe's cloud and on-premises infrastructure. You will lead a team of IT professionals, drive operational excellence, support our cloud transformation journey, and work closely with internal stakeholders and external partners to deliver reliable and efficient IT services.
Your mission:
Team Leadership and Daily Operations
- Lead, mentor, and coordinate the Cloud Operations team.
- Prioritize and oversee incident, problem, and service request management.
- Ensure timely resolution of operational tickets and maintain high service levels.
- Drive automation initiatives using PowerShell, Bash, Python, Azure DevOps, and GitHub Actions.
- Define and maintain engineering standards, runbooks, and operational procedures.
- Contribute to infrastructure governance and change review processes.
Cloud and On-Premises Operations
- Oversee daily operations across Azure, on-premises infrastructure, and hybrid environments.
- Design, deploy, and maintain physical and virtual server environments.
- Manage Azure Landing Zones, subscriptions, governance, and security controls.
- Develop and maintain Infrastructure as Code using Terraform and Azure Bicep/ARM.
- Manage Azure networking, virtual machines, connectivity, and monitoring solutions.
- Support and lead workload migrations from on-premises environments to Azure.
- Monitor system health, performance, capacity, and availability.
- Manage patching, backups, disaster recovery processes, and platform stability.
- Administer Windows Server, Active Directory, DNS, DHCP, Group Policy, and PKI services.
- Respond to and resolve critical infrastructure incidents within agreed SLAs.
Incident Management and On-Call Support
- Lead troubleshooting and root cause analysis for complex infrastructure issues.
- Coordinate and participate in on-call support activities.
- Drive continuous improvements to reduce recurring operational issues.
Service Reliability and Optimization
- Implement monitoring, alerting, and operational dashboards.
- Optimize performance, resource utilization, and cloud costs.
- Support automation initiatives that improve operational efficiency.
Vendor and Partner Coordination
- Manage relationships with third-party providers and support partners.
- Review service performance and SLAs, ensuring corrective actions when required.
- Coordinate infrastructure changes and implementations delivered by external vendors.
Customer and Stakeholder Engagement
- Collaborate closely with Delivery, Security, Applications, and Service Desk teams.
- Communicate incidents, operational updates, and planned changes effectively.
- Translate business and operational needs into service improvements.
Governance, Documentation and Compliance
- Maintain operational documentation, procedures, and runbooks.
- Ensure compliance with IT policies, security standards, and change management processes.
- Support audits and operational readiness reviews.
Backup, Recovery & Business Continuity
- Own and evolve the organization’s backup and recovery strategy.
- Manage Veeam Backup & Replication and Azure Backup solutions.
- Maintain backup policies, repositories, and recovery services.
- Conduct restore testing, disaster recovery exercises, and continuity planning activities.
- Ensure compliance with retention, audit, and business continuity requirements.
What we ask for:
- Bachelor’s degree in Computer Science, Information Technology, or a related field. Additional training in cloud operations, Azure, or IT infrastructure is an advantage.
- Strong experience managing cloud and on-premises infrastructure operations in medium to large-scale environments.
- Solid knowledge of Microsoft Azure services, virtual machines, storage, networking, and monitoring.
- Hands-on experience with VMware environments supporting on-premises workloads.
- Good understanding of Microsoft 365 administration and operational support.
- Experience using Jira for ticket management and GitHub for documentation, scripting, and automation.
- Knowledge of identity and access management concepts, including Entra ID, Active Directory, role-based access control, and secure authentication methods.
- Strong troubleshooting and problem-solving skills across complex infrastructure environments.
- Experience with patching, backup, disaster recovery, and operational maintenance activities.
- Ability to participate in on-call rotations and support priority incident response.
- Understanding of hybrid cloud environments and cloud migration initiatives.
- Experience with monitoring platforms, alerting solutions, dashboards, and automation tools.
- Hands-on experience with Infrastructure as Code, including Terraform and/or Azure Bicep/ARM.
- Experience with Veeam Backup & Replication and/or Azure Backup in enterprise environments.
- Proven experience administering VMware vSphere and/or Microsoft Hyper-V production environments.
- Familiarity with monitoring tools such as Azure Monitor, Grafana, Prometheus, and SCOM.
- Exposure to infrastructure CI/CD pipelines using Azure DevOps and/or GitHub Actions.
- Proven ability to lead and coordinate infrastructure or operations teams.
- Strong customer-focused mindset with excellent communication and collaboration skills.
- Ability to work effectively with internal stakeholders and external service providers.
- Structured, proactive, and organized approach to problem-solving and documentation.
- Good written and spoken English communication skills.
Certifications
- Microsoft Azure Administrator Associate (AZ-104).
- Microsoft Azure Solutions Architect Expert (AZ-305).
We are keen to hear from you even if you don't match all listed requirements, but you identify with our brand and passion for innovation.