> Markdown version of [/jobs/ext/1426089-software-engineer-data-acquisition](https://www.wearedevelopers.com/jobs/ext/1426089-software-engineer-data-acquisition). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Data Acquisition - **Company:** People Data Labs - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $160,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** Airflow, Amazon Web Services, Proxy Servers, Microsoft Azure, Big Data, BigQuery, Command-Line Interface, Data Architecture, Extract Transform Load (ETL), Data Warehousing, Software Debugging, Software Design Patterns, Distributed Data Store, Distributed Systems, Domain Name System (DNS), Fault Tolerance, Python (Programming Language), Network Architecture, Packet Analyzer, Software Architecture, Queueing Systems, SQL Databases, Data Streaming, Web Applications, Web Crawlers, Data Storage Technologies, Data Ingestion, Snowflake, Concurrency, Apache Spark, Parallel Computation, Indexer, Backend, Data Lakes, Information Technology, Apache Kafka, Spark Streaming, Asynchronous Programming, Software Coding, Amazon Simple Queue Service (SQS), Data Pipelines, Amazon Redshift, Databricks - **Published:** July 24, 2026 - **Apply:** https://www.builtincolorado.com/auth/login?destination=/job/senior-software-engineer-data-acquisition/10356040 ## About the Role * 7+ years of professional experience building or operating backend or infrastructure systems at scale * Solid programming experience in Python, Go, Rust, or similar, including experience with async / await, coroutines, or concurrency frameworks * Strong grasp of software architecture and backend fundamentals; you can reason clearly about concurrency, scalability, and fault tolerance * Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request / response) * Familiarity with network architecture and debugging (HTTP, DNS, proxies, packet capture and analysis) * Solid understanding of distributed systems concepts: parallelism, asynchronous programming, backpressure, and message-driven design * Experience designing or maintaining resilient data ingestion, API integration, or ETL systems * Proficiency with Linux / Unix command-line tools and system resource management * Familiarity with message queues, orchestration, and distributed task systems (Kafka, SQS, Airflow, etc.) * Experience evaluating and monitoring data quality, ensuring consistency, completeness, and reliability across releases People Thrive Here Who Can * Work independently in a fast-paced, remote-first environment, proactively unblocking themselves and collaborating asynchronously * Communicate clearly and thoughtfully in writing (Slack, docs, design proposals) * Write and maintain technical design documents, including pipeline design, schema design, and data flow diagrams * Scope and break down complex projects into deliverable milestones, and communicate progress, risks, and blockers effectively * Balance pragmatism with craftsmanship, shipping reliable systems while continuously improving them Some Nice To Haves * Degree in a quantitative field such as computer science, mathematics, or engineering * Experience as a Red Teamer * Experience working on large-scale data ingestion, crawling, or indexing systems * Experience with Apache Spark, Databricks, or other distributed data platforms * Experience with streaming data systems (Kafka, Pub/Sub, Spark Streaming, etc.) * Proficiency with SQL and data warehousing (Snowflake, Redshift, BigQuery, or similar) * Experience with cloud platforms (AWS preferred, GCP or Azure also great) * Understanding of modern data storage and design patterns (parquet, Delta Lake, partitioning, incremental updates) * Knowledge of modern data design and storage patterns (e.g., incremental updating, partitioning and segmentation, rebuilds and backfills) * Experience building and maintaining data pipelines on modern big-data or cloud platforms (Databricks, Spark, or equivalent) ## Description We are looking for individuals who can balance extreme ownership with a "one-team, one-dream" mindset. Our customers are trying to solve complex problems, and we only help them achieve their goals as a team. Our Data Engineering & Acquisition Team ensures our customers have standardized and high quality data to build upon. You will be crucial in accelerating our efforts to build standalone data products that enable data teams and independent developers to create innovative solutions at massive scale. In this role, you will be working with a team to continuously improve our existing datasets as well as pursuing new ones. If you are looking to be part of a team discovering the next frontier of data-as-a-service (DaaS) with a high level of autonomy and opportunity for direct contributions, this might be the role for you. We like our engineers to be thoughtful, quirky, and willing to fearlessly try new things. Failure is embraced at PDL as long as we continue to learn and grow from it. What You Get to Do * Contribute to the architecture and improvement of our data acquisition and processing platform, increasing reliability, throughput, and observability * Use and develop web crawling technologies to capture and catalog data on the internet * Build, operate, and evolve large-scale distributed systems that collect, process, and deliver data from across the web * Design and develop backend services that manage distributed job orchestration, data pipelines, and large-scale asynchronous workloads * Structure and model captured data, ensuring high quality and consistency across datasets * Continuously improve the speed, scalability, and fault-tolerance of our ingestion systems * Partner with data product and engineering teams to design and implement new data products powered by the data you help collect, and enhance and improve upon existing products * Learn and apply domain-specific knowledge in web crawling and data acquisition, with mentorship from experienced teammates and access to existing systems ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Enjoying SQL data pipelines with dbt](https://www.wearedevelopers.com/videos/823-enjoying-sql-data-pipelines-with-dbt) - [Dynamic Entities in .NET: Building Low-Code Systems on Top of Entity Framework Core](https://www.wearedevelopers.com/videos/100218-dynamic-entities-in-net-building-low-code-systems-on-top-of-entity-framework-core) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market)