> Markdown version of [/jobs/ext/3318746-tech-lead-data-infrastructure-site-reliability-expiring-soon](https://www.wearedevelopers.com/jobs/ext/3318746-tech-lead-data-infrastructure-site-reliability-expiring-soon). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Tech Lead - Data Infrastructure Site Reliability Expiring soon - **Company:** BYTEDANCE INC. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $254,400.0 - **Contract:** Permanent contract - **Skills:** Systems Engineering, Big Data, Cloud Computing, Databases, Data Centers, Data Infrastructure, Decision Support Systems, Distributed Systems, Machine Learning, NoSQL, Performance Tuning, Reliability Engineering, Software Engineering, SQL Databases, Data Streaming, Systems Architecture, Cloud Platform System, Kubernetes - **Published:** September 1, 2026 - **Apply:** https://www.gamesjobsdirect.com/job/bytedance/tech-lead-data-infrastructure-site-reliability/350852 ## About the Role Minimum Qualifications - 5+ years of experience in Site Reliability Engineering, Software Development, or related fields, with a strong focus on designing, building, scaling, and operating cloud-based systems. - Deep hands-on expertise in at least one of the following areas: - Databases (SQL/NoSQL) - Kubernetes or container orchestration - Big Data processing and storage systems (streaming and batch) - Strong knowledge of system architecture, distributed systems, and performance bottlenecks. - Excellent communication and collaboration skills, with experience working across engineering, product, and data science teams. Preferred Qualifications - Proven track record of driving automation, tooling, and process improvements that enhance reliability and efficiency. - Experience in cost optimization and performance tuning at scale, backed by data-driven decision making. - Thought leadership in adopting new technologies, improving operational practices, and influencing system design. ## Description Team Introduction: Our Site Reliability Engineering (SRE) team blends software and systems engineering to build and operate large-scale data infrastructure with high reliability and efficiency. We provide a dependable cloud environment that powers our global business. In this role, you will leverage your expertise in data center architecture, data infrastructure services, and systems and tools development to solve complex scaling and reliability challenges. We're looking for a Technical Lead (SRE) who can provide deep technical leadership, drive architectural improvements, and collaborate effectively across multiple organizations. You'll partner with engineering, product, data, and infrastructure teams to deliver resilient, scalable platforms. This is a highly technical, hands-on role that requires strong problem-solving ability, clear communication, and the ability to influence without formal authority. Responsibilities - Strong hands-on skills in the design, development, and operation of large-scale cloud infrastructure and distributed systems. - Collaborate with cross-functional teams (e.g., Advertising, Machine Learning, E-commerce, and Core Infra) to drive system reliability, performance, and scalability. - Lead initiatives to automate operations, eliminate toil, and improve overall system efficiency. - Troubleshoot complex production issues, perform root-cause analysis, and drive long-term reliability improvements. - Promote best practices in system design, observability, performance optimization, and cost efficiency. - Communicate complex technical concepts effectively to both technical and non-technical stakeholders. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Tomorrow's cloud data platforms - fully managed database-as-a-service (DBaaS)](https://www.wearedevelopers.com/videos/254-tomorrow-s-cloud-data-platforms-fully-managed-database-as-a-service-dbaas) - [In-Memory Computing - The Big Picture](https://www.wearedevelopers.com/videos/626-in-memory-computing-the-big-picture) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)