Engineering Manager, Reliability Engineering Flywheel - EDA Infrastructure — Seattle
NVIDIA · US, WA, Redmond
Posted Posted 40 mins ago
Refyne take
This Engineering Manager, Reliability Engineering Flywheel - EDA Infrastructure role at NVIDIA was posted in the last 5 days. Before applying, run your resume through the checker on the right — most rejections here are keyword and formatting mismatches, not qualifications.
NVIDIA’s EDA Infrastructure organization builds and operates the systems that support chip development. We are looking for an engineering manager to lead the team responsible for operational processes and platforms across incident management, maintenance, on-call, issue management, and customer-serving readiness.
You will own the roadmap and delivery, from defining how teams work to building the tools they use. You will partner with infrastructure and service owners to improve reliability, reduce manual work, and ensure services are ready to support customers. Your team will use automation, AI, and lessons from operational events to drive improvements.
What you’ll be doing:
• Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results.
• Set technical direction, prioritize work, and guide execution across engineering and operational disciplines.
• Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness.
• Hire and develop engineers and technical leads, building a team with clear ownership and accountability.
• Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents.
What we need to see:
• BS degree or equivalent experience with 10+ overall years of software engineering or related experience, including 5+ years of engineering leadership managing teams or complex technical programs.
• Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement.
• Strong technical judgment in software architecture, platform integration, and engineering tradeoffs.
• Clear communication with engineers, cross-functional partners, and executive stakeholders.
• A record of developing engineers, growing teams, and delivering results under pressure.
Ways to stand out from the crowd:
• Established readiness standards covering service ownership, support coverage, and reliability objectives.
• Built, integrated, and scaled platforms pertaining the incident, maintenance, customer experience management, along with on-call and production readiness
• Applied AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation.
• Supported EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements.
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous?
Do you love a challenge? If so, we want to hear from you. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. If you're creative and self-motivated, we want to hear from you! NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 3, and 272,000 USD - 431,250 USD for Level 4.
You will also be eligible for equity and benefits .
Applications for this job will be accepted at least until September 26, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Explore related searches
Browse the wider category, city and company pages this role belongs to.
Similar site reliability engineer jobs
Other openings posted in the last five days that match this role.
- Executive Director - Site Reliability Engineering - Retail PharmacyCVS · RI - WoonsocketonsiteleadPosted 5 mins ago
- Site Reliability Engineer 6476266Accenture · Charlotte, 1120 S Tryon St., CorponsitemidPosted 46 mins ago
- Senior Site Reliability Engineer (US Federal)Workday · USA.VA.RestononsiteseniorPosted 51 mins ago
- Site Reliability Engineer (US Federal)Workday · USA.VA.RestononsitemidPosted 51 mins ago
- Senior PDK Reliability EngineerIntel · US, Oregon, HillsborohybridseniorPosted 51 mins ago
- Staff TDI Site Reliability Engineer, Okta FederalOkta · Washington, DChybridstaffQuick applyPosted 1 hour ago
More jobs at NVIDIA
- Senior Network Security Engineer - DGX CloudNVIDIA · US, CA, Santa ClarahybridseniorPosted 40 mins ago
- Senior SOC Design EngineerNVIDIA · US, CA, Santa ClarahybridseniorPosted 40 mins ago
- Senior ASIC Design Engineer – Clocks IPNVIDIA · US, CA, Santa ClarahybridseniorPosted 40 mins ago
- Senior Software Engineer, Linux PlatformNVIDIA · US, CA, Remote · RemoteremoteseniorPosted 40 mins ago
- Senior System Firmware Engineer - SOC and GPUNVIDIA · US, CA, Santa ClaraonsiteseniorPosted 40 mins ago
- Senior Cell Modeling and Verification EngineerNVIDIA · US, CA, Santa ClarahybridseniorPosted 40 mins ago