> Markdown version of [/jobs/ext/2866123-software-engineer-ii-platform-data-reliability-new](https://www.wearedevelopers.com/jobs/ext/2866123-software-engineer-ii-platform-data-reliability-new). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer II Platform Data Reliability New - **Company:** Sony Interactive Entertainment - **Location:** Reading, PA, United States - **Experience:** Experienced - **Salary:** $150,100.0 - $225,100.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Build Automation, Cloud Computing, Configuration Management, Code Review, Databases, Data as a Services, Data Integrity, Linux, Distributed Systems, Amazon DynamoDB, Fault Tolerance, NoSQL, Redis, Reliability Engineering, Ansible, Software Engineering, Data Streaming, Data Logging, Aerospike, Amazon ElastiCache, Grafana, Concurrency, Software Troubleshooting, Caching, Infrastructure as Code (IaC), Kubernetes, Information Technology, Cassandra, Apache Kafka, Data Management, Video Streaming, Terraform - **Published:** September 12, 2026 - **Apply:** https://www.gamesjobsdirect.com/job/sony-interactive-entertainment/software-engineer-ii-platform-data-reliability/358465 ## About the Role * Bachelor's or Master's degree in Computer Science or a related field, or equivalent practical experience. * 3+ years of experience in software engineering, database reliability engineering, site reliability engineering, platform engineering, or a related field. * Experience developing production software in Go, with an understanding of idiomatic code, testing, concurrency, error handling, and maintainability. * Hands-on experience developing or maintaining Infrastructure as Code or configuration-management tools such as Terraform or Ansible. * Experience deploying or operating workloads on Kubernetes. * Experience with AWS or GCP and familiarity with managed services such as MSK, DynamoDB, ElastiCache, Memorystore, or equivalent technologies. * Working knowledge of one or more NoSQL, caching, or streaming technologies, such as Cassandra, Aerospike, Kafka, AWS MSK, or Redis. * Understanding of distributed systems concepts, including availability, consistency, replication, fault tolerance, and horizontal scaling. * Working knowledge of Linux, networking, storage, and common system-troubleshooting techniques. * Familiarity with observability tools and practices, including metrics, logging, tracing, alerting, and dashboard creation. * Ability to independently diagnose and resolve technical problems within a defined scope, learn unfamiliar systems, and seek guidance when addressing complex or ambiguous challenges. * Strong written and verbal communication skills, with the ability to collaborate effectively across teams. ## Description We are seeking a Software Engineer II (Platform Data Reliability & Automation) to help build, automate, and operate scalable data platforms using Infrastructure as Code (IaC) and cloud technologies. This role focuses on improving the reliability and automation of NoSQL, streaming, and caching services across AWS and GCP environments. You'll develop automation, observability, and operational tooling supporting technologies such as Cassandra, Aerospike, Kafka, and Redis. Working alongside senior engineers, platform teams, and product teams, you'll contribute to highly available infrastructure supporting billions of transactions and millions of players globally. By applying software engineering and database reliability engineering principles, you'll help reduce manual work, improve system uptime, and make data services easier and safer for engineering teams to use., * Develop, maintain, and improve Infrastructure as Code and configuration-management automation using tools such as Terraform and Ansible to provision, configure, monitor, scale, and manage NoSQL, streaming, and caching platforms. * Build automation that enables repeatable and reliable deployment of data services across cloud and hybrid environments. * Contribute to the reliability, availability, scalability, performance, and resiliency of platform data services. * Contribute to defining, measuring, and improving service-level indicators, service-level objectives, and error budgets. * Develop automation for operational activities such as scaling, failover, backup, recovery, upgrades, and routine maintenance. * Build and enhance observability solutions using metrics, logging, tracing, dashboards, and alerts. * Troubleshoot issues affecting Cassandra, Aerospike, Kafka/MSK, Redis, and related platform services. * Participate in on-call rotations and incident response, contributing to root-cause analysis and the implementation of permanent fixes. * Write reliable, maintainable, and well-tested Go code for infrastructure automation, platform services, and operational tooling. * Collaborate with engineering, platform, security, and operations teams to integrate and deliver reliable data services. * Create and maintain operational documentation, procedures, runbooks, and automation playbooks. * Participate in code reviews, technical design discussions, and continuous improvement initiatives. * Explore practical applications of AI-assisted automation, anomaly detection, automated remediation, and developer-productivity tooling where appropriate. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Navigating the Corporate Jungle: Life as a Developer in a large Company](https://www.wearedevelopers.com/videos/621-navigating-the-corporate-jungle-life-as-a-developer-in-a-large-company) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering)