> Markdown version of [/jobs/ext/3417824-senior-site-reliability-engineer-data-infrastructure-san-jose-new](https://www.wearedevelopers.com/jobs/ext/3417824-senior-site-reliability-engineer-data-infrastructure-san-jose-new). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - Data Infrastructure (San Jose) New - **Company:** BYTEDANCE INC. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $218,400.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Big Data, Computer Programming, Data as a Services, Data Centers, Data Infrastructure, Data Systems, Distributed Systems, Domain Name System (DNS), Python (Programming Language), MySQL, Networking Basics, Performance Tuning, Queueing Systems, Redis, Reliability Engineering, Site Reliability Engineering Practices, Software Engineering, TCP/IP, AI Infrastructure, Scripting, System Availability, Build Management, Core Data, Kubernetes, Information Technology, Apache Flink, Apache Kafka, Golang - **Published:** September 1, 2026 - **Apply:** https://www.gamesjobsdirect.com/job/bytedance/senior-site-reliability-engineer-data-infrastructure-san-jose/356979 ## About the Role Minimum Qualifications: - Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience. - 5+ years of experience in a Site Reliability Engineering, Production Engineering, or similar role. - Strong proficiency in a programming or scripting language (e.g., Go, Python, Bash) for automation and tool development. - Deep understanding of Linux/Unix operating systems, networking fundamentals (TCP/IP, DNS), and distributed systems. Preferred Qualifications: - Extensive hands-on experience managing large-scale data infrastructure (e.g., MySQL, Redis, Kafka, Flink). - Proven experience with container orchestration technologies, particularly Kubernetes, in a production environment. - Expertise in designing, analyzing, and troubleshooting large-scale distributed systems. - A systematic problem-solving approach, coupled with strong communication skills and a sense of ownership. - Experience leading incident response for complex, high-impact events. - Experience in the operation and construction of Data Centers is a big plus. ## Description The Data Infrastructure SRE team is responsible for the reliability, scalability, and efficiency of the core data services that power our products. We manage a massive, distributed environment built on technologies like Kubernetes, Redis, MySQL, and Message Queue. Our work is not about building features, but about engineering the resilience and performance of the underlying platform that all product teams depend on. We are the guardians of production, ensuring our data systems run smoothly, nonstop. This role includes participation in a rotational on-call schedule to ensure nonstop coverage for our critical data infrastructure. You will be expected to respond to, troubleshoot, and resolve production incidents. Our team collaborates across multiple time zones, and you will engage in rigorous change management and post-incident review processes to maintain system stability. Role Summary As a Site Reliability Engineer, you will be on the front lines of keeping our large-scale data systems running reliably and efficiently. You will focus on hands-on operational work, from responding to alerts and managing production changes to automating routine tasks. This role is an excellent opportunity to develop deep expertise in modern infrastructure technologies and SRE practices while working alongside senior engineers to solve challenging problems. Responsibilities: - Incident response and postmortems: Act as an incident commander for critical production issues, guiding the team through triage and resolution. Drive deep, blameless post-incident reviews and ensure that follow-up actions are implemented to prevent recurrence. - SLO/SLA and error budgets: Define, negotiate, and maintain Service Level Objectives (SLOs) for critical data services. Champion the use of error budgets to balance reliability work with feature development. - Capacity and cost optimization: Lead initiatives in capacity planning, performance tuning, and resource management. Develop strategies and automation to ensure our infrastructure scales efficiently and stays within budget. - Pragmatic automation and AI orchestration: Design and build automation and leverage AI Agents to eliminate operational toil, improve deployment safety, and enhance overall operational efficiency. Focus on creating maintainable, robust tools and intelligent workflows that make the entire team more effective. - Operational excellence and change management: Uphold and improve our standards for production operations, including runbooks, monitoring, and alerting. Vet complex changes and deployments to ensure they meet our bar for production readiness. - Data Center and AI Infrastructure: Lead the construction, maintenance, and optimization of data centers and specialized AI infrastructure, ensuring high availability and peak performance for complex AI-driven workloads. - Cross-team influence and mentorship: Act as a subject matter expert on reliability, consulting with application development and other infrastructure teams. Mentor junior SREs, helping them develop their technical and operational skills. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)