Software Engineer II Platform Data Reliability New

Sony Interactive Entertainment
Reading, PA, United States
8 days ago
Apply on www.gamesjobsdirect.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$150,100.0 - $225,100.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Build Automation Cloud Computing Configuration Management Code Review Databases Data as a Services Data Integrity Linux Distributed Systems Amazon DynamoDB Fault Tolerance
+21 more
NoSQL Redis Reliability Engineering Ansible Software Engineering Data Streaming Data Logging Aerospike Amazon ElastiCache Grafana Concurrency Software Troubleshooting Caching Infrastructure as Code (IaC) Kubernetes Information Technology Cassandra Apache Kafka Data Management Video Streaming Terraform

Job description

We are seeking a Software Engineer II (Platform Data Reliability & Automation) to help build, automate, and operate scalable data platforms using Infrastructure as Code (IaC) and cloud technologies.

This role focuses on improving the reliability and automation of NoSQL, streaming, and caching services across AWS and GCP environments. You’ll develop automation, observability, and operational tooling supporting technologies such as Cassandra, Aerospike, Kafka, and Redis.

Working alongside senior engineers, platform teams, and product teams, you’ll contribute to highly available infrastructure supporting billions of transactions and millions of players globally. By applying software engineering and database reliability engineering principles, you’ll help reduce manual work, improve system uptime, and make data services easier and safer for engineering teams to use., * Develop, maintain, and improve Infrastructure as Code and configuration-management automation using tools such as Terraform and Ansible to provision, configure, monitor, scale, and manage NoSQL, streaming, and caching platforms.

  • Build automation that enables repeatable and reliable deployment of data services across cloud and hybrid environments.
  • Contribute to the reliability, availability, scalability, performance, and resiliency of platform data services.
  • Contribute to defining, measuring, and improving service-level indicators, service-level objectives, and error budgets.
  • Develop automation for operational activities such as scaling, failover, backup, recovery, upgrades, and routine maintenance.
  • Build and enhance observability solutions using metrics, logging, tracing, dashboards, and alerts.
  • Troubleshoot issues affecting Cassandra, Aerospike, Kafka/MSK, Redis, and related platform services.
  • Participate in on-call rotations and incident response, contributing to root-cause analysis and the implementation of permanent fixes.
  • Write reliable, maintainable, and well-tested Go code for infrastructure automation, platform services, and operational tooling.
  • Collaborate with engineering, platform, security, and operations teams to integrate and deliver reliable data services.
  • Create and maintain operational documentation, procedures, runbooks, and automation playbooks.
  • Participate in code reviews, technical design discussions, and continuous improvement initiatives.
  • Explore practical applications of AI-assisted automation, anomaly detection, automated remediation, and developer-productivity tooling where appropriate.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of experience in software engineering, database reliability engineering, site reliability engineering, platform engineering, or a related field.
  • Experience developing production software in Go, with an understanding of idiomatic code, testing, concurrency, error handling, and maintainability.
  • Hands-on experience developing or maintaining Infrastructure as Code or configuration-management tools such as Terraform or Ansible.
  • Experience deploying or operating workloads on Kubernetes.
  • Experience with AWS or GCP and familiarity with managed services such as MSK, DynamoDB, ElastiCache, Memorystore, or equivalent technologies.
  • Working knowledge of one or more NoSQL, caching, or streaming technologies, such as Cassandra, Aerospike, Kafka, AWS MSK, or Redis.
  • Understanding of distributed systems concepts, including availability, consistency, replication, fault tolerance, and horizontal scaling.
  • Working knowledge of Linux, networking, storage, and common system-troubleshooting techniques.
  • Familiarity with observability tools and practices, including metrics, logging, tracing, alerting, and dashboard creation.
  • Ability to independently diagnose and resolve technical problems within a defined scope, learn unfamiliar systems, and seek guidance when addressing complex or ambiguous challenges.
  • Strong written and verbal communication skills, with the ability to collaborate effectively across teams.

Benefits & conditions

At SIE, we consider several factors when setting each role’s base pay range, including the competitive benchmarking data for the market and geographic location.

Please note that the base pay range may vary in line with our hybrid working policy and individual base pay will be determined based on job-related factors which may include knowledge, skills, experience, and location.

In addition, this role is eligible for SIE’s top-tier benefits package that includes medical, dental, vision, matching 401(k), paid time off, wellness program and coveted employee discounts for Sony products. This role also may be eligible for a bonus package. Click here to learn more.

The estimated base pay range for this role is listed below.

$150,100 - $225,100 USD

Please note, Sony Interactive Entertainment conducts background checks at the offer stage for all new employees (which may include criminal background checks for some roles) and will need to process personal information to support these checks.

About the company

Why Sony Interactive Entertainment?

Sony Interactive Entertainment isn’t just the Best Place to Play - it’s also the Best Place to Work. Sony Interactive Entertainment (SIE) is the company behind the PlayStation brand. As a subsidiary of Sony Group Corporation, we’re part of a proud legacy of innovation and excellence. SIE is a dynamic technology company, delivering cutting-edge hardware and network services to more than 100 million people and an entertainment leader, home to some of the most beloved and recognizable intellectual properties (IP) in the world. Our role at SIE is to create and nurture the experiences under the PlayStation brand, a name synonymous with entertainment excellence and creativity.

Software Engineer II - Platform Data Reliability & Automation

Ready to level up your career? Join PlayStation as a Software Engineer II focused on Platform Data Reliability and Automation and help build reliable, scalable experiences for millions of players around the globe.

At PlayStation, we’re known not only for delivering exceptional gaming experiences but also for fostering an engineering environment centered on innovation, creativity, collaboration, and technical excellence. We welcome passionate engineers who enjoy solving challenging problems and are excited about shaping the future of play., Sony Interactive Entertainment pushes the boundaries of entertainment and innovation, starting from the launch of the original PlayStation in Japan in 1994. Today, we continue to deliver innovative and thrilling experiences to a global audience through our PlayStation line of products and services that include generation-defining hardware, pioneering network services, and award-winning games. Continue reading

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.gamesjobsdirect.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

Videos

See all

Related articles

See all