> Markdown version of [/jobs/ext/597839-systems-engineer](https://www.wearedevelopers.com/jobs/ext/597839-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Systems Engineer - **Company:** Recursion Technologies, Inc. - **Location:** Richardson, TX, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Data Analysis, Systems Engineering, Big Data, Computer Engineering, Data Validation, Data Governance, Extract Transform Load (ETL), Software Debugging, Document Management Systems, Distributed Data Store, Github, Apache Hadoop, Hadoop Distributed File System, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Kerberos (Protocol), Network Troubleshooting, Log Analysis, Prometheus, SQL Databases, SQLAlchemy, Data Streaming, Data Logging, Scripting, Apache Yarn, System Availability, Apache Spark, Kubernetes, Information Technology, Apache Kafka, Data Pipelines - **Published:** June 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=f69f1e8adbaaa70b ## About the Role Do you have experience in Python?, Do you have a Bachelor's degree?, Bachelor's Degree is required in Computer Science or Computer Engineering or Computer Information Systems or Information Technology or Data Science. ## Description · Design, develop, and maintain large-scale data processing pipelines using Apache Spark. · Monitor and troubleshoot Spark job failures, including driver/executor crashes and performance bottlenecks. · Manage and optimize workloads running on Hadoop (HDFS, YARN) clusters. · Provide support for onboarding new data pipelines and services into the platform. · Analyze and resolve resource allocation issues such as CPU/memory quota exceedance in Kubernetes environments. · Build and maintain ETL pipelines for ingesting, transforming, and loading large datasets. · Ensure data quality, consistency, and integrity across distributed data systems. · Implement alerting rules and thresholds using monitoring platforms (e.g., Prometheus-based systems). · Work with Kafka to manage data streaming pipelines, including topic configuration and access control. · Troubleshoot Kafka consumer/producer issues, including lag, permissions, and connectivity errors. · Implement and maintain data retention policies in Lakehouse architectures. · Perform log analysis and debugging using distributed logging tools. · Coordinate with infrastructure teams to resolve cluster-level or networking issues. · Configure and manage storage paths, table-level retention, and lifecycle policies for datasets. · Develop and execute SQL queries for data analysis, validation, and reporting. · Automate workflows and monitoring using Python scripts and APIs (e.g., GitHub API, SQLAlchemy). · Continuously improve system efficiency, scalability, and cost optimization. · Analyze alerts from monitoring systems and take proactive action to prevent outages. · Investigate production incidents (P1/P2) and perform root cause analysis (RCA). · Collaborate with cross-functional teams (developers, SREs, data engineers) to resolve system issues. · Conduct data validation and reconciliation between upstream and downstream systems. · Maintain dashboards and observability tools (e.g., Hubble, internal monitoring systems). · Optimize performance of distributed jobs by tuning configurations and execution plans. · Handle identity and access management issues across systems (Kerberos, service accounts, ACLs). · Support migration and integration of new data technologies into existing ecosystems. · Work on Kubernetes-based Spark deployments and troubleshoot pod scheduling and quota issues. · Ensure high availability and reliability of data pipelines and streaming jobs. · Participate in on-call rotations and respond to critical production alerts. · Document system architecture, troubleshooting steps, and operational procedures. · Ensure compliance with organizational data governance and security standards. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [JavaScript? No. Java Scripts! - Scripting with Java](https://www.wearedevelopers.com/videos/2094-javascript-no-java-scripts-scripting-with-java) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Software Engineer Career: Things You Should Know](https://www.wearedevelopers.com/magazine/143-software-engineer-career-things-you-should-know) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 13 Best Python Libraries for Developers in 2025](https://www.wearedevelopers.com/magazine/371-the-13-best-python-libraries-for-developers-in-2025)