Software Engineer- AI/ML, Amazon Neuron Training — Cupertino, California, USA
Amazon · Cupertino, California, USA
Posted Posted 7 hours ago
Refyne take
This Software Engineer- AI/ML, Amazon Neuron Training role at Amazon was posted in the last 5 days. Before applying, run your resume through the checker on the right — most rejections here are keyword and formatting mismatches, not qualifications.
The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit used to accelerate deep learning and GenAI workloads on AWS Trainium, Amazon's custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate with PyTorch and JAX, so customers can train frontier-scale models on Trainium without rewriting their stack.
The Distributed Training team enables the training of a wide range of models, from large-scale pretraining through post-training and reinforcement learning, on AWS's custom ML accelerators. As more customer workloads shift toward RLHF, PPO/GRPO, and other fine-tuning methods, we are building the distributed training infrastructure, parallelism techniques, numerics, and high-performance kernels that these methods depend on. As part of the broader Neuron organization, we work across frameworks, kernels, compiler, runtime, and collectives — a true hardware and software co-design in practice. We not only optimize current performance but also contribute to future architecture designs, since the gaps we characterize today become requirements for the next generation of Trainium.
We are looking for software engineers to help build and fine tune these distributed training solutions. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, where you will help shape the direction of AI acceleration technology.
Key job responsibilities
You’ll implement and tune components of our distributed training stack for large-scale training, post-training, and reinforcement learning workloads on the latest Trainium instances, working across PyTorch and the Neuron software stack. You'll contribute to parallelism strategies such as data, tensor, and pipeline parallelism and apply reduced-precision formats under the guidance of senior team members. You'll profile workloads to help determine whether a bottleneck sits in compute, memory, collectives, or host overhead, and work with compiler, runtime, and collectives engineers to help land the fix. You'll build and maintain internal tooling, benchmarks, and tests that keep the team's performance work reproducible, and take on increasing ownership as you grow in the role.
About the team
Inclusive Team Culture
Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 16 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust.
Work/Life Balance
Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives.
- Bachelor's degree or above in computer science or equivalent
- 3+ years of experience with full software development life cycle in production
- 3+ years of experience with at least one programming language such as Python, C/C++, or a similar language
- 2+ years of experience with system design (design patterns, reliability and scaling) of new or existing systems
- Familiarity with LLM/transformer fundamentals such as attention mechanisms, autoregressive decoding, KV-cache behavior, and various forms of parallelism
- Master's degree or above in computer science or equivalent
- Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware, or experience in computer architecture
- Experience with ML frameworks such as Pytorch/Jax, Distributed libraries and Frameworks, RL frameworks or End-to-end Model Training
- Experience with performance engineering: workload profiling, characterization (compute bound, memory bound, network bound), and optimization
- Experience with open source projects
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .
USA, CA, Cupertino - 165,200.00 - 223,600.00 USD annually
Explore related searches
Browse the wider category, city and company pages this role belongs to.
Similar software engineer jobs
Other openings posted in the last five days that match this role.
- Android Framework/Core OS Software Engineer IIZebra Technologies · Holtsville, New YorkhybridmidPosted 17 mins ago
- Firmware/Software Engineer SeniorZebra Technologies · Lincolnshire, IllinoishybridseniorPosted 17 mins ago
- Senior Software EngineerLeidos · Arlington, VAonsiteseniorPosted 20 mins ago
- Senior Software Development Engineer in Test (Senior SDET / SETI)Fiserv · Sunnyvale, CaliforniaonsiteseniorPosted 21 mins ago
- Sr. Manager, Software EngineeringFiserv · Sunnyvale, CaliforniaonsiteseniorPosted 21 mins ago
- Software Engineer III - Java/C#bofa · New YorkonsitemidPosted 22 mins ago
More jobs at Amazon
- Sr. SDE - EmbeddedAmazon · Redmond, Washington, USAonsiteseniorPosted 7 hours ago
- Mission Operations Systems Engineer, Optical Inter-Satellite LinkAmazon · Redmond, Washington, USAonsitemidPosted 7 hours ago
- Software Development Engineer, 3P Measurement TechAmazon · New York, New York, USAonsitemidPosted 7 hours ago
- Sr. Software Engineer- AI/ML, Amazon Neuron TrainingAmazon · Cupertino, California, USAonsiteseniorPosted 7 hours ago
- Software Development Engineer, Measurement, Ad Tech, and Data Science (MADS) - Tax & PromotionsAmazon · New York, New York, USAonsitemidPosted 7 hours ago
- Sr. Solutions Architect, Annapurna MLAmazon · Seattle, Washington, USAonsitestaffPosted 7 hours ago