Title: Infrastructure Disaster Recovery (DR) Service Delivery Manager Location: Tempe, AZ- 3 days per week Duration: 12+ months
Call notes from Hiring Manager: Position Overview
Lead and coordinate enterprise Disaster Recovery (DR) activities.
Support DR failovers, Point-in-Time Recovery, and BCP plans.
Act as a technical coordinator between operations, engineers, and application teams.
Ensure DR exercises are planned, tested, and executed successfully.
Improve processes, runbooks, and automation.
Support teams across the US, APAC, and EMEA.
Responsibilities
Plan DR exercises, typically up to 10 weeks in advance.
Validate CMDB data and confirm assets in scope.
Coordinate failover steps and technical dependencies.
Prepare change requests and support CAB approvals.
Coordinate technical teams during DR events.
Identify missing participants and resolve issues quickly.
Make or support go/no-go decisions.
Review and improve DR runbooks and sequencing.
Identify tasks that can run in parallel.
Identify automation opportunities.
Support Point-in-Time Recovery exercises.
Maintain required audit and compliance records.
Support annual DR plan testing and certification.
Raise and track required DR incident tickets.
Capture lessons learned and improve future exercises.
Support additional resilience, compliance, and technology risk projects.
Participate in occasional weekend DR exercises.
Must-Have Skills
Strong infrastructure engineering or operations background.
Hands-on Disaster Recovery experience.
Experience with DR failovers and recovery activities.
Knowledge of Point-in-Time Recovery and BCP.
Understanding of application and infrastructure dependencies.
Experience with enterprise DR planning and execution.
CMDB, runbook, change management, and CAB knowledge.
Strong technical problem-solving skills.
Ability to coordinate multiple technical teams.
Strong communication and stakeholder management skills.
Confident and able to speak up during critical events.
Ability to understand and validate technical information.
Ability to work independently and make informed decisions.
Nice-to-Have Skills
Automation experience.
Experience improving manual processes.
Developer or engineering automation background.
Interest in AI-enabled automation.
Large enterprise DR experience.
Financial services experience.
Experience in regulated industries.
Healthcare or similar regulated industry experience.
Experience working with global teams.
Education
Degree in Computer Science, IT, Engineering, or related field preferred.
Equivalent technical experience may be considered.
Experience
7+ years in technical infrastructure engineering or operations.
4+ years of hands-on Disaster Recovery experience.
Enterprise-scale DR experience preferred.
Must be able to explain personal contributions to DR activities.
Pure Project Management experience is not sufficient.
Pure infrastructure experience without DR experience is not sufficient.
Additional Insights Ideal Candidate
Strong mix of infrastructure and DR experience.
Former engineers moving into a DR-focused role may be a good fit.
DR does not need to be the candidate's only responsibility.
Must understand the technology behind the DR process.
Must know potential technical risks and failure points.
Must have genuine, hands-on experience.
Technical Screening
Focus on what the candidate personally did.
Do not accept team-level answers only.
Validate technologies listed on the resume.
Ask specific questions about middleware and databases.
Assess the size and complexity of past DR environments.
Enterprise experience is preferred over limited branch-level DR exposure.
Assess both technical depth and communication skills.
Automation
Automation is a strong plus, not a mandatory requirement.
The team expects automation to increase over time.
Candidates should identify repetitive work and improvement opportunities.
AI and engineering tools may be used to improve efficiency.
Industry
Financial services experience is preferred.
Experience in another regulated industry is also valuable.
Healthcare experience may be considered transferable
Client is seeking an Infrastructure Disaster Recovery (DR) Service Delivery Manager that will work with the Engineering and Operations Managers across the Infrastructure towers to assess, test, and help to implement various infrastructure solutions with the Resiliency Team. This diverse position will require the individual to work with application and infrastructure owners, DR program managers, automation engineers, and business function testers.
A successful candidate will help to understand the overlap / connections of the existing applications with the underlying infrastructure mechanisms, to aid in cloning the solution in an isolated lab environment, configuring, and testing recovery solutions with DR automation experts, and validating failover testing with business function testers. This person will have hands-on experience across multiple on-premises and cloud technologies including Control-M, database, ETL, platform (mainframe, iSeries, VMWare, Unix/Linux, Windows) messaging, middleware, monitoring, operating systems, networking, storage, and virtualization.
Responsibilities • Serve as key liaison during DR events for Infrastructure, in a spokesperson role for the infra towers if a gap has been identified, facilitating communications between Clients groups • Have or acquire an in-depth understanding of the various moving parts and functions involved in the DR process especially as involves infrastructure components • Represent the infrastructure towers in DR planning meetings with IT Leadership and operational committees to communicate the status, capabilities, and infrastructure gaps • Apply in depth or broad technical knowledge to provide maintenance / automation solutions across one or more technology areas (e.g., database administration) • Integrate technical expertise and business understanding to coordinate / envision superior solutions for the infrastructure components • Assess and coordinate high availability and seamless failover resiliency mechanisms across multiple layers of the application and infrastructure stacks • Manage the update and testing of DR Plans, Point In Time Recovery (PiTR) Plans, and DR Testing of those plans within SLAs without extension • Demonstrate technical expertise and exert influence outside of immediate Infrastructure tower • Coordinate and drive innovative team solutions to complex problems • Coordinate DR for new Infrastructure technologies • Service as the POC for DR Test / Fusion incident management events across the infrastructure towers • Assist in integration of solutions into the DR automation program, reducing overall recovery time for Infrastructure services • Coordinate innovation and technology experiments in private, public and hybrid cloud at enterprise scale • Maintain an effective technical network across technical SMEs and architects for multiple service areas • Analyze high volumes of technical and operational data during a disaster recovery event • Ingest and communicate program metrics / issues with the greater Infrastructure team to allow for data-driven risk management and process improvements • Sign off on Infrastructure procedural documentation, Disaster Recovery Plans, training modules, and other disaster recovery artifacts for quality assurance • Coordinate and help to guide the implementation of a comprehensive testing strategy enabled through data analytics and automation • Manage the Infrastructure remediation review program to ensure gaps identified by other functions and stakeholders are appropriately addressed • Identify and remediate Infrastructure gaps identified by the control teams or audit, or through internal disaster recovery assessments • Share key insights and learnings from participation in Infrastructure meetings to share with management
Qualifications: • Bachelor’s degree in Information Technology, Management Information Systems, Computer Science or a related discipline • 7+ years as Infrastructure technical engineer operating at enterprise scale • 4+ years of DR experience • Advanced technical problem-solving skills • Practical experience in both technology infrastructure and application development architectures • Infrastructure and service architecture/engineering experience, including functional and technical requirements gathering, and solution development • Strong system analysis and service development experience • Practical experience operating in an Agile development environment • Automation experience • Financial or Regulatory domain experience a plus • VMWare vRO experience a plus • Service Now experience • Familiarity with Fusion a plus
Knowledge/Skills: • Experience in creating Disaster Recovery, BCP, PiTR and DR Plans • Demonstrated understanding and knowledge of security concepts such as Replication and Backups, Enterprise Architecture, Authentication, Access Management, and Network Segmentation • Experienced in conducting testing exercises for Disaster/BCP/Cyber Plans • Well-versed in disaster security industry best practices and regulatory and compliance frameworks • Highly motivated, energetic self-starter who takes ownership of issues and drives them to resolution • Good organizational skills - manage and prioritize multiple tasks across different time horizons within deadlines • Cooperative, and collegial manner to be able to flex to wide range of stakeholders and in high-stakes situations • Strong decision-making skills • Strong analytical, problem solving and process re-engineering skills • Excellent project management skills; oral and written communication • Skills in translating broad strategic intent into tactical plans and directions are essential • Applies broad industry knowledge and awareness • Negotiate with senior leaders across the business
Applicant Notices & Disclaimers
For information on benefits, equal opportunity employment, and location-specific applicant notices, click here
At SPECTRAFORCE, we are committed to maintaining a workplace that ensures fair compensation and wage transparency in adherence with all applicable state and local laws.This position's pay range is $61.88/hr - $65.00/hr.
Infrastructure Disaster Recovery (DR) Service Delivery Manager