Développeur.se principal.e de l’infrastructure / Lead Infrastructure Developer
About this role
At EA, we believe games are powerful because they bring together multiple ways people engage: play, watch, create, and connect. And increasingly, the biggest entertainment platforms aren't just places to consume content — they're places where communities build.
Creator-made content is already a proven part of EA's history — from community creation tools in Battlefield to The Gallery in The Sims 4. We believe new creative technologies and tools will expand how players engage with and contribute to our experiences, supported by thoughtful product design, safety systems, and global reach. Our focus is on enabling more players to participate in creative expression by making creation easier, safer, and more rewarding.
As Lead Infrastructure Engineer, you will own the GPU fleet our researchers train on including capacity, scheduling, diagnostics, and support. You will set technical direction for GPU operations and infrastructure architecture. You will additionally lead Infrastructure as Code setup, granting permissions, and debugging infrastructure problems.
This is a hybrid role, working three days per week in Redwood City, Montreal, or Vancouver.
You will report to the Head of Data and Infrastructure.
Responsibilities:
You will own GPU fleet operations across our AWS estate.
You will build the scheduling layer from zero.
You will diagnose GPU and node failures fast and completely and drive hardware evidence and replacement through AWS support and capacity-block channels.
You will run researcher support as a first-class product including holding office hours, owning the support channel, and driving the recurring causes out of existence with self-service tooling, preflight checks, and documentation
You will instrument the fleet including utilization, queue depth, job success rate, and cost per experiment metrics.
You will partner with our external compute and lab partnerships as a technical contact, and with EA's central infrastructure groups on shared services and escalation.
You will author runbooks, decision records, and onboarding docs.
Qualifications:
8+ years of experience operating production infrastructure, with deep, current, hands-on AWS depth — EC2 GPU fleets, EKS, IAM and cross-account security, VPC and networking, S3 and FSx
Experience scheduling, diagnosing, and managing GPUs in AWS specifically
Experience operating GPU fleets at 1000+ GPU scale
Expertise in scripting and automation with Python, PowerShell, bash, or equivalent
Expertise in infrastructure as code (Terraform or equivalent)
Familiarity with a GPU scheduling or orchestration layer (like Slurm, Kubernetes with Kueue or Volcano, Ray, dStack or SkyPilot)
Observability practice including Grafana, Prometheus, or equivalent
Similar open roles
| Title | Company | Location | Posted |
|---|---|---|---|
| Artiste de la surface d'environnement - Environment Surfacing Artist (DNEG Animation) | DNEG | Montreal, Canada | 2026-09-17 |
| Generalist Software Engineer - EA SPORTS™ FC | Electronic Arts | Vancouver, Canada | 2026-09-16 |
| Développeur.se - opérations liées à l’apprentissage automatique / MLOps Developer | Electronic Arts | Montreal, Canada | 2026-09-15 |
| Développeur.se d’outils / Tools Developer (Apex Legends) | Electronic Arts | Vancouver - Great Northern Way, Canada | 2026-09-11 |
| Senior Software Engineering Manager - EA SPORTS™ UFC | Electronic Arts | Vancouver, Canada | 2026-09-04 |
| Site Reliability Engineer | Electronic Arts | Vancouver, Canada | 2026-09-01 |
| Senior AI Software Developer | Electronic Arts | Montreal, Canada | 2026-08-31 |
| Lead Software Engineer - EA SPORTS™ NHL | Electronic Arts | Vancouver, Canada | 2026-08-31 |