Live Support Engineer
About this role
As one of the largest sports entertainment platforms in the world, EA SPORTS FC is redefining football with genre-leading interactive experiences, connecting a global community of fans to The World's Game through innovation and unrivaled authenticity.
With more opportunity than ever to design, innovate, and create new, immersive experiences that bring joy, inclusivity, and connection to fans everywhere, we invite you to join our passionate and dynamic team as we pioneer the future of football fandom.
EA SPORTS FC Mobile Shanghai is a global team devoted to developing and operating a high-quality mobile football game experience. Our quest for creativity, respect for autonomy, and emphasis on collaboration are at the heart of our team culture, which empowers us to create high-quality games and experiences worldwide.
As a team, we are passionate, innovative, and open to possibilities. We learn from past experiences and strive for progress. We value team synergy and believe a relaxed working environment can yield better results. That's why we promote and support maintaining a healthy work-life balance.
You will report to devops director
Responsibilities:
Prod and Pre-Prod Environment Monitoring
Monitor client and server stability across pre-production and production environments, including crashes, ANRs, error rates, key performance metrics, service health, core gameplay flows and infrastructure status.
Use Firebase, Grafana, Sentry, Splunk, VictoriaLogs, and VictoriaMetrics to identify abnormal trends and correlate metrics, logs, and error events.
Identify and drive improvements to alert noise, monitoring gaps, missing telemetry, and recurring issues, while continuously optimizing dashboards and alerting rules.
Provide stability support during release windows and major esports events.
Live Incident Response
You will work as a first-line responder for live incidents. Follow Incident Runbooks for investigation, communication, escalation, and recovery, providing updates on impact, investigation progress, mitigation actions, and recovery status throughout an incident.
Follow approval processes and Runbooks to perform mitigation actions such as restarts, failovers, configuration changes, and Kill Switch activation, confirming authorization, risks, and rollback plans before execution.
Maintain complete operational records for post-incident review and audit purposes.
Escalate promptly to service owners, development teams, or infrastructure teams in accordance with the Escalation Policy when an issue falls outside the authorized scope or Runbook.
AI-Powered Live Monitoring and Operations Tools
You will build AI-powered live monitoring, alert analysis, and operations support tools, including internal AI Agents, MCP tools, knowledge bases, prompts, evaluation cases, and runtime configurations.
Integrate data sources and APIs from monitoring, logging, and collaboration platforms to automatically collect and correlate metrics, logs, and error information.
Turn repetitive health checks, anomaly summaries, impact analysis, and incident context collection into automated workflows.
Evaluate the accuracy of AI-generated analysis and reduce false positives, false negatives, and non-applicable recommendations.
Knowledge and Process Development
You will develop knowledge bases covering client data flows, service architecture, key dependencies, error codes, and metric catalogs, along with dashboard guides and troubleshooting documentation.
Convert incident insights into applicable Runbooks.
Work with development teams to complete the handover and acceptance of dashboards, alerts, and Runbooks, and drive documentation updates after feature releases.
Track incident action items and lead responsible teams to deliver long-term fixes.
Qualifications:
A degree in computer science or a related field, or equivalent technical capability with three or more years of experience in SRE, operations, production support or technical support roles.
Familiarity with Kubernetes, containers, cloud platforms, and microservices architecture.
Foundational development skills in Go, Python, or another programming language, with the ability to integrate APIs, write automation scripts, and maintain internal tools.
Familiarity with the following: Prometheus, ELK, Grafana, Sentry, Splunk, VictoriaLogs, VictoriaMetrics and Firebase.
Basic knowledge of large language models, AI Agents, prompting, RAG, or MCP, with the ability to validate the reliability of AI-generated output.
Read technical documentation in English and communicate in writing with global development and operations teams.
Willing to participate in shift rotations, release support, and on-call coverage.
Similar open roles
| Title | Company | Location | Posted |
|---|---|---|---|
| Lead Server Engineer | Electronic Arts | Shanghai, China | 2026-09-14 |
| Full Stack Engineer | Electronic Arts | Shanghai, China | 2026-09-14 |
| (Senior) Full Stack Engineer | Electronic Arts | Shanghai, China | 2026-09-14 |
| Server Engineer | Electronic Arts | Shanghai, China | 2026-09-10 |
| Client Engineer - Engine | Electronic Arts | Shanghai, China | 2026-09-10 |
| Client Engineer - Pipeline/Tools | Electronic Arts | Shanghai, China | 2026-09-09 |
| Client Engineer, Unannounced 3A Project | Electronic Arts | Shanghai, China | 2026-07-22 |
| Server Engineer, Unannounced 3A Project | Electronic Arts | Shanghai, China | 2026-07-22 |