Data Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+31 more
Job description
This role is for a Data Engineer to design, build, and maintain the ETL pipelines and data infrastructure that feed a data lake, anomaly detection models, and AI agents. The position focuses on constructing robust, scalable data pipelines using Spark/Scala, ensuring data quality and availability across a growing portfolio of network data sources. The goal is to enable downstream consumers such as data scientists, agents, and dashboards to access reliable, well-structured data., * Design, develop, and maintain scalable ETL pipelines using Apache Spark (Scala) to ingest, transform, and load network data into the data lake.
- Onboard new data sources (network telemetry, syslogs, SNMP traps, device configuration data, ticketing systems) by building ingestion pipelines.
- Implement monitoring and alerting solutions to ensure data pipeline reliability and performance.
- Develop and manage deployment pipelines for continuous integration and delivery of data engineering solutions.
- Manage and optimize data storage solutions, including distributed file systems, relational databases, and external sources accessed via API.
- Implement data quality checks, validation rules, and automated testing to ensure pipeline reliability and data integrity.
- Optimize pipeline performance for large-scale data processing across batch and mini-batch processing patterns.
- Manage and evolve data schemas, partitioning strategies, and storage formats.
- Support data backfills and recovery when upstream issues or schema changes require reprocessing.
- Collaborate with data scientists and agent developers to deliver datasets that support anomaly detection models and AI agent workflows.
Requirements
- Expert-level experience in building and maintaining ETL pipelines for Big Data.
- Programming experience with Python and/or Scala.
- Experience with large-scale data processing with Spark.
- Experience with AWS services: S3, Glue, Athena, and EMR, including working with EMR as the operating system.
- Strong understanding of relational databases and SQL.
- Knowledge of data architecture, data warehousing, partitioning strategies, and columnar storage formats (e.g., Parquet).
- Experience with workflow orchestration tools, with a preference for Airflow.
- Proficiency with Linux-based operating systems and shell scripting.
- Experience with Git-based version control and collaborative development workflows., * Experience with streaming or mini-batch data processing (Spark Streaming, structured streaming, or similar).
- Experience with Apache Kafka or similar messaging/streaming platforms.
- Experience with NoSQL databases.
- Experience in the telecommunications industry or other large-scale network operations environments.
- Familiarity with network data sources: telemetry, syslogs, SNMP traps, device configuration data.
- Experience with data integration via REST APIs and cloud SDKs (e.g., boto3).
- Experience writing automated tests for data pipelines.
- Knowledge of text analysis or log parsing techniques.
About the company
Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.
Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Making Data Warehouses Fast: A Developer’s Story
Data Engineer Salary UK
Top Big Data Technologies That You Need to Know
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again