About the role
Manager, Site Reliability Engineering
62 M
Overview
Responsible for leading the Site Reliability Engineering Center of Excellence and the Forward Deployed SRE program supporting critical banking platforms, applications, and technology services. Manages an organization of employees and contingent resources through direct reports, program managers, SRE leaders, technical leads, and matrixed delivery relationships.
Establishes the enterprise SRE strategy, operating model, engineering standards, governance, talent model, and adoption roadmap. Accountable for improving service reliability, availability, scalability, performance, resiliency, deployment safety, and operational maturity across supported technology domains.
Leads the deployment of SRE capabilities into application and platform teams through a Forward Deployed SRE model. Partners with senior leaders across application development, infrastructure, cloud engineering, architecture, cybersecurity, technology operations, risk, and business-aligned technology organizations to prioritize engagements and deliver measurable reliability improvements.
Balances strategic leadership, people management, program execution, technical governance, and operational accountability. Ensures SRE practices are implemented consistently and that reliability investments produce measurable improvements in customer experience, operational risk, engineering productivity, release quality, and service performance.
Primary Responsibilities
SRE Strategy and Center of Excellence Leadership
Establish and execute the vision, strategy, operating model, service offerings, and multiyear roadmap for the Site Reliability Engineering Center of Excellence.
Define enterprise SRE standards, engineering practices, governance processes, engagement models, and maturity expectations.
Lead the adoption of reliability engineering practices across application development, infrastructure, platform engineering, cloud engineering, and technology operations.
Translate enterprise technology, business, resiliency, and risk priorities into an actionable SRE portfolio and delivery roadmap.
Establish a scalable SRE service model that includes consulting, enablement, embedded engineering, Forward Deployed SRE engagements, reusable capabilities, and sustained ownership by application and platform teams.
Define intake, prioritization, engagement, transition, and exit criteria for SRE services.
Develop and maintain an SRE maturity model used to assess service teams, identify reliability gaps, and guide improvement plans.
Establish communities of practice, technical forums, training programs, playbooks, reference architectures, and reusable engineering patterns that expand SRE capabilities across the organization.
Ensure the SRE Center of Excellence remains aligned with enterprise engineering standards, cloud strategy, operational risk requirements, and evolving industry practices.
Represent the SRE organization in senior leadership forums, architecture reviews, operational governance meetings, and enterprise transformation initiatives.
Forward Deployed SRE Program Leadership
Lead the Forward Deployed SRE program, placing SRE professionals into high-priority application and platform teams to address complex reliability challenges and improve operational maturity.
Manage program managers, SRE leaders, and technical leads responsible for coordinating engagements across multiple technology domains.
Establish a transparent intake and prioritization process based on customer impact, service criticality, operational risk, incident history, reliability maturity, strategic importance, and anticipated business value.
Partner with application and platform leaders to define engagement objectives, scope, deliverables, staffing, success measures, dependencies, and duration.
Ensure Forward Deployed SRE teams deliver sustainable engineering improvements rather than becoming long-term substitutes for application support or production operations.
Establish shared accountability for participation, knowledge transfer, remediation activities, and long-term ownership of implemented reliability capabilities.
Develop transition and exit plans that enable application and platform teams to sustain SRE practices after engagements conclude.
Evaluate engagement effectiveness using measurable outcomes such as availability, SLO attainment, incident frequency, restoration time, alert quality, automation adoption, toil reduction, change-failure rate, deployment reliability, and engineering maturity.
Convert common findings and lessons learned into reusable standards, automation, tools, training, and engineering patterns.
Continuously optimize the Forward Deployed SRE operating model based on demand, capacity, outcomes, stakeholder feedback, and changes in technology strategy.
Organizational and People Leadership
Lead an organization of employees and contingent resources through direct and indirect management relationships.
Manage and develop program managers, SRE managers, technical leaders, and senior engineering professionals.
Establish clear roles, responsibilities, decision rights, performance expectations, and accountability across the SRE organization.
Build a high-performing organization with expertise in reliability engineering, observability, cloud platforms, release engineering, deployment orchestration, automation, Infrastructure as Code, incident management, resiliency engineering, testing, and program delivery.
Recruit, retain, coach, and develop diverse engineering and program management talent.
Conduct workforce, capacity, and succession planning to ensure the organization has the leadership and technical capabilities required to meet current and future demand.
Define career paths and skill-development plans for SRE professionals in partnership with engineering and human resources leaders.
Provide regular feedback, performance management, recognition, coaching, and development opportunities.
Create an environment that encourages engineering discipline, continuous learning, constructive challenge, accountability, collaboration, and knowledge sharing.
Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
Manage staffing and sourcing strategies, including the appropriate use of employees, contingent labor, managed services, and specialized partners.
Ensure third-party resources and service providers meet applicable engineering, security, risk, performance, financial, and contractual expectations.
Reliability Engineering Governance
Establish standards and governance for Service Level Indicators, Service Level Objectives, error budgets, availability targets, and reliability reporting.
Partner with service owners and business stakeholders to align reliability objectives with customer expectations, service criticality, business impact, risk tolerance, and cost.
Define common reliability metrics and executive-level reporting for service health, operational performance, deployment health, risk, and improvement outcomes.
Ensure reliability metrics are accurate, actionable, consistently defined, and tied to accountable service owners.
Establish governance for error-budget decisions, including remediation, release-risk evaluation, delivery tradeoffs, escalation, and investment prioritization.
Drive adoption of reliability-by-design practices throughout the Software Development Lifecycle.
Establish operational readiness requirements for applications and platforms entering production or undergoing material change.
Ensure readiness assessments address availability, scalability, observability, performance, supportability, recoverability, security, testing, capacity, release and rollback readiness, documentation, and operational ownership.
Partner with architecture and engineering leaders to incorporate reliability requirements into solution designs, architecture reviews, and engineering standards.
Identify systemic reliability risks and sponsor cross-organizational remediation programs.
Release Engineering and Deployment Orchestration
Establish enterprise SRE standards for safe, repeatable, observable, and recoverable application and infrastructure releases.
Partner with application development, platform engineering, cloud engineering, quality engineering, change management, and technology operations teams to improve release and deployment practices.
Define approved deployment patterns based on service criticality, architecture, customer impact, technical capability, and risk.
Promote progressive delivery practices, including blue-green deployments, canary releases, rolling deployments, ring-based deployments, feature flags, traffic splitting, and controlled production experimentation where appropriate.
Establish requirements for automated pre-deployment, in-deployment, and post-deployment validation using technical health signals, service-level indicators, business metrics, and customer-experience measures.
Promote deployment orchestration that integrates CI/CD pipelines, Infrastructure as Code, automated testing, observability, approval controls, and policy enforcement.
Establish standards for automated rollback, roll-forward, traffic evacuation, deployment pausing, and recovery when release health thresholds are breached.
Define release health criteria and quality gates based on error rates, latency, saturation, availability, dependency health, business transactions, and customer-impact indicators.
Promote small, incremental, independently deployable changes that reduce release complexity and limit failure impact.
Partner with platform teams to provide reusable deployment templates, pipeline capabilities, policy-as-code controls, and self-service release patterns.
Establish traceability between changes, deployments, configuration updates, incidents, service telemetry, and business outcomes.
Improve release observability through deployment markers, version-aware dashboards, automated change correlation, and real-time health analysis.
Use error budgets, service criticality, testing evidence, deployment history, and current service health to inform release-risk decisions.
Measure and improve deployment frequency, lead time for changes, change-failure rate, rollback effectiveness, failed deployment recovery time, and release-related customer impact.
Ensure deployment practices comply with applicable technology risk, cybersecurity, change-management, segregation-of-duties, and regulatory requirements.
Observability, Automation, and Cloud Engineering
Define the enterprise observability strategy for supported services, including standards for telemetry, distributed tracing, metrics, logging, dashboards, alerting, dependency mapping, and customer-experience monitoring.
Govern the effective implementation and use of Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, Log Analytics, and other approved enterprise tools.
Drive standardization of telemetry and observability patterns across cloud-native, hybrid, distributed, and legacy environments.
Establish expectations for actionable alerts, signal quality, service-health visibility, release observability, and real-time operational insight.
Reduce alert fatigue and operational noise by improving monitoring coverage, alert thresholds, routing, correlation, suppression, ownership, and automation.
Sponsor reusable dashboards, instrumentation patterns, automation libraries, reference architectures, and reliability controls.
Develop and execute a strategy to reduce operational toil through automation, self-service capabilities, automated recovery, and self-healing solutions.
Establish measurable toil-reduction goals and ensure reclaimed engineering capacity is redirected toward reliability, automation, and product improvements.
Promote Infrastructure as Code using Terraform and other approved technologies to improve repeatability, control, recoverability, and environment consistency.
Partner with Azure and enterprise platform teams to improve scalability, resiliency, deployment safety, lifecycle management, and operational controls.
Incident, Problem, and Resiliency Management
Provide leadership and executive coordination during significant technology incidents affecting customers, critical business services, or enterprise operations.
Ensure appropriate technical leadership, stakeholder communication, escalation, decision-making, and recovery focus during high-severity events.
Partner with incident management, technology operations, application teams, infrastructure teams, and business leaders to minimize customer impact and restore services safely.
Establish expectations for timely, objective, and technically rigorous Root Cause Analysis.
Ensure corrective and preventive actions address underlying technical, process, monitoring, testing, release, deployment, architecture, and organizational causes.
Track material reliability actions to completion and escalate overdue or inadequately addressed risks.
Analyze incident, change, deployment, and operational telemetry to identify recurring failure patterns and enterprise-level improvement opportunities.
Sponsor proactive problem-management efforts that reduce repeat incidents and improve service stability.
Establish standards for resiliency testing, fault-tolerance validation, performance testing, capacity planning, disaster recovery, and recovery automation.
Promote game days, controlled failure testing, and other approved validation methods to test service behavior and organizational readiness.
Partner with technology resiliency and continuity teams to ensure recovery capabilities meet applicable business and regulatory requirements.
Ensure lessons from incidents, failed changes, deployment events, and resiliency exercises are incorporated into engineering standards, training, tooling, and future designs.
Communicate material reliability risks, customer impacts, remediation strategies, and progress to senior management and governance bodies.
Program and Portfolio Management
Manage a portfolio of SRE initiatives and Forward Deployed SRE engagements spanning multiple lines of business, platforms, applications, and technology domains.
Establish portfolio governance, delivery cadences, health reporting, dependency management, escalation paths, and outcome tracking.
Provide direction and oversight to program managers responsible for planning, sequencing, coordinating, and reporting SRE initiatives.
Define measurable objectives, key results, milestones, and success criteria for the SRE Center of Excellence and Forward Deployed SRE program.
Monitor execution against approved roadmaps, budgets, commitments, capacity, risk tolerances, and benefit expectations.
Resolve delivery barriers that span organizational boundaries and escalate decisions requiring senior leadership involvement.
Establish demand-management and capacity-planning processes for SRE services.
Prioritize investments based on customer impact, operational risk, service criticality, strategic alignment, regulatory considerations, and measurable reliability benefit.
Develop business cases for staffing, tooling, platform capabilities, release engineering, automation, observability, and engineering transformation.
Manage applicable budgets, forecasts, vendor relationships, contracts, licensing requirements, and financial commitments.
Demonstrate the value of SRE investments through measurable improvements in operational performance, customer experience, risk reduction, release quality, and engineering productivity.
Stakeholder and Executive Leadership
Build trusted partnerships with senior leaders across application engineering, infrastructure, cloud, platform engineering, architecture, cybersecurity, technology operations, risk, audit, and business-aligned technology organizations.
Advise technology and business leaders on reliability risks, service-health trends, engineering priorities, deployment risks, operational readiness, and investment decisions.
Communicate complex technical issues, tradeoffs, risks, and recommendations in language appropriate for technical, business, risk, and executive audiences.
Present SRE strategy, operating metrics, program status, investment needs, material risks, and reliability outcomes to leadership and governance forums.
Facilitate decisions where reliability, delivery speed, cost, customer impact, and operational risk must be balanced.
Influence teams outside the direct reporting structure to adopt enterprise SRE standards and resolve cross-functional reliability concerns.
Establish clear accountability among the SRE organization, product teams, service owners, platform teams, release teams, and technology operations.
Use structured stakeholder feedback to continuously improve SRE services and engagement practices.
Risk, Controls, and Regulatory Responsibilities
Understand and adhere to the Company’s risk and regulatory standards, policies, and controls in accordance with the Company’s Risk Appetite.
Ensure SRE practices, tools, automation, release processes, and operating models comply with applicable technology, cybersecurity, data, privacy, resiliency, change-management, and third-party risk requirements.
Identify, assess, document, and escalate material reliability, operational, technology, program, release, and control risks.
Ensure appropriate preventive and detective controls are incorporated into SRE standards, automation, CI/CD pipelines, deployment processes, observability solutions, and operational workflows.
Maintain M&T internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
Support examinations, audits, risk assessments, control testing, and remediation activities related to the SRE organization and supported services.
Ensure risk acceptances, exceptions, and material reliability or deployment decisions are documented, approved, monitored, and escalated in accordance with Company requirements.
Establish governance that provides transparency into unresolved reliability risks and associated remediation commitments.
Complete other related duties as assigned.
Scope of Responsibilities
Supervisory and Managerial Responsibilities
Manages an organization employees and contingent resources through direct reports, managers, program managers, technical leads, and matrixed delivery relationships.
Directly manages a combination of people leaders, program managers, senior technical professionals, and other SRE personnel.
Accountable for organizational design, staffing, recruiting, performance management, compensation recommendations, talent development, succession planning, employee engagement, and workforce capacity.
Establishes priorities and allocates people and funding across the SRE Center of Excellence and Forward Deployed SRE portfolio.
Provides leadership across multiple concurrent programs and technical engagements.
Accountable for the quality, timeliness, control effectiveness, financial management, and measurable outcomes of work delivered by the organization.
May manage employees and teams across multiple locations, technology domains, and organizational reporting structures.
Organizational Impact
Influences reliability strategy, release engineering, engineering practices, and operational performance across multiple technology organizations.
Decisions affect critical applications, business services, customer experience, operational risk, technology investment, delivery performance, and engineering productivity.
Operates with significant autonomy while escalating material business, technology, customer, financial, regulatory, and operational risks.
Regularly interacts with senior technology leaders, business stakeholders, risk partners, auditors, vendors, and enterprise governance bodies.
Education and Experience Required
A combined minimum of 11 years’ higher education and/or work experience, including a minimum of 4 years’ engineering and/or architecture experience and 5 years’ leadership experience, including people management.
Complete understanding of the Software Development Life Cycle and modern software delivery practices.
Demonstrated leadership delivering enterprise-wide SRE, reliability engineering, production engineering, or comparable technology capabilities, including operating models and adoption mechanisms.
Proven ownership of SDLC reliability and operational readiness controls, enforceable standards, templates, quality gates, and supporting evidence requirements.
Experience establishing governance for SLOs, SLIs, error budgets, observability, release readiness, deployment safety, resiliency, and production operations.
Experience with release and deployment orchestration, including progressive delivery practices such as blue-green deployments, canary releases, rolling deployments, feature flags, automated health validation, and rollback.
Strong execution management skills, including prioritization, dependency management, resource planning, and institutionalizing solutions beyond direct involvement.
Experience managing complex initiatives, large system enhancements, cloud or platform transformations, application conversions, and production issue resolution.
Experience leading multidisciplinary engineering teams and coordinating delivery across application, infrastructure, platform, cloud, operations, risk, and business organizations.
Excellent verbal and written communication skills, with strong analytical, decision-making, negotiation, and organizational capabilities.
Confidence leading multiple teams across geographies and time zones.
Prior experience presenting technology strategy, delivery progress, operational performance, investment needs, and material risks to senior management.
Core Technical Requirements
Strong understanding of SRE practices, including SLOs, SLIs, error budgets, observability, incident and problem management, Root Cause Analysis, operational readiness, resiliency, capacity planning, and toil reduction.
Strong understanding of cloud-native and hybrid architectures, distributed systems, APIs, containers, service dependencies, and Microsoft Azure.
Experience leading observability capabilities using Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, Log Analytics, or comparable technologies.
Experience governing Infrastructure as Code, preferably using Terraform, and CI/CD pipeline practices.
Understanding of automated regression, performance, resiliency, recovery, and post-deployment testing.
Knowledge of release quality gates, traffic management, automated rollback, self-healing, high-availability, and disaster-recovery patterns.
Ability to provide credible leadership to senior engineers and evaluate architectural proposals for reliability, release, and operational risks.
Education and Experience Preferred
Bachelor’s or Master’s degree in engineering, computer science, information systems, business administration, or a related discipline.
Minimum of 12 years of technology management or large program leadership experience.
Experience leading an SRE Center of Excellence, production engineering organization, platform engineering function, or Forward Deployed SRE operating model.
Hands-on familiarity with enterprise observability, incident management, service management, deployment orchestration, and engineering workflow platforms.
Experience building reusable SRE capabilities, automation frameworks, deployment patterns, reference architectures, and self-service engineering solutions with multi-team and vendor coordination.
Exposure to AI-assisted operations, AIOps, automated incident correlation, predictive analytics, performance engineering, or chaos engineering capabilities.
Subject matter expertise in business-critical applications, distributed platforms, and integrated systems.
Strong understanding of the Bank’s application framework, business plan, risk environment, and strategic objectives.
Proven mentoring, coaching, organizational development, and enterprise leadership capabilities.
Self-motivated leader with the ability to motivate and influence others across organizational boundaries.
#LI-JB3
M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.Location
Buffalo, New York, United States of AmericaAplyr's read
M&T Bank is a regional financial institution with a strong community focus, attracting professionals in commercial banking and financial services.
What's promising
- •M&T Bank has a strong regional presence in the Northeastern U.S.
- •The bank offers diverse career opportunities in commercial and retail banking.
- •M&T Bank is known for its community-oriented approach and local involvement.
What to watch
- •M&T Bank faces competition from larger national banks.
- •Limited public information about its technological innovation initiatives.
- •The bank's regional focus may limit geographic career mobility.
Why M&T Bank
- •M&T Bank emphasizes a community-focused banking model.
- •The company offers specialized roles in commercial credit and equipment financing.
- •M&T Bank provides hybrid work opportunities for certain roles, enhancing work-life balance.
Aplyr’s read is generated by AI from public sources. Was it useful?
About M&T Bank
M&T Bank is a regional bank holding company that provides a wide range of financial services, including commercial and retail banking, investment services, and mortgage banking.
Similar roles
Lead Site Reliability Network Security Engineer, Fortinet
Jabil
Director, Site Reliability Engineering
Anduril Industries
Lead Site Reliability Engineer - Interface & Connectivity
SimCorp
Manager - Production Operations & Site Reliability Engineering
Alcon
Head of DevOps & Site Reliability Engineering
Finastra
Platform Engineering & Site Reliability Engineering- Vice President
Deutsche Bank