Senior Back-End Engineer (Website Scanning)
Role details
Job location
Tech stack
Job description
Our product tackles important scanning + remediation challenges in the privacy + consent space (think of cookies + consent management, GDPR / CCPA, google tag manager, data privacy, etc.) and our product is used by stakeholders seeking to ensure compliance w/ standard data privacy laws as well as law firms who need to capture, assess and triage issues related to privacy + security., We are seeking a senior Back-End Software Engineer focused on building and scaling high-performance web scanning and data ingestion systems. You will design, implement, and operate robust distributed services that crawl, scrape, process, and store large volumes of web data while maintaining reliability, compliance, and efficiency at scale. The role combines deep systems engineering, web automation, anti-bot evasion strategies, and platform-level scalability work across cloud and containerized environments., * Design, implement, and maintain scalable microservices and distributed systems for web crawling, scraping, and data ingestion using Go or Rust or Python
- Build and operate headless browser automation and scraping pipelines using Puppeteer, Playwright, Selenium, or lightweight HTTP clients (Requests, AIOHTTP, OkHttp, Axios) with robust HTML parsing (BeautifulSoup, lxml, Cheerio) and CSS/XPath selector strategies.
- Develop and optimize REST and GraphQL APIs to serve downstream consumers
- Implement concurrency and parallelism patterns (asyncio, goroutines, multithreading, multiprocessing) and task queues (Celery, RQ) to maximize throughput while controlling resource usage and backpressure.
- Architect and tune message-driven pipelines with brokers like Kafka, RabbitMQ, AWS SQS, and task scheduling with Airflow or Cron for ETL workflows.
- Design storage and indexing solutions using relational databases (PostgreSQL, MySQL), NoSQL stores (MongoDB, Redis), and search systems (Elasticsearch) with attention to schema design, normalization, and query performance.
- Build robust proxy management, IP rotation strategies, robots.txt and sitemap parsing, rate limiting/throttling, retry/backoff patterns, and CAPTCHA handling and solving strategies as needed.
- Implement fingerprinting and anti-bot evasion techniques, caching strategies, reverse proxy configurations (Nginx), and secure handling of credentials, tokens, and sensitive data (OAuth, JWT).
- Containerize services with Docker and deploy and manage them on Kubernetes or serverless platforms (AWS Lambda, Google Cloud Functions), leveraging cloud platforms (AWS, GCP, Azure) for scalability and resilience.
- Instrument systems for observability and reliability using logging and monitoring tools (ELK, Prometheus, Grafana, Datadog), set up alerts, and conduct post-incident analysis and performance tuning.
- Collaborate with product, data, and security teams to define SLAs, compliance boundaries, and automated tests; own CI/CD pipelines using Git, GitHub Actions, Jenkins, or GitLab CI for repeatable deployments.
- Lead efforts to harden systems with input sanitization, secure coding practices, and automated unit/integration testing and test automation to ensure correctness and reduce technical debt.
- Mentor engineers, perform code reviews, and contribute to architectural decisions that improve scalability, maintainability, and developer productivity.
Requirements
- Bachelors or Masters degree in Computer Science, Engineering, or equivalent practical experience.
- 5+ years building back-end systems with production experience in Python, Go, Java, Node.js, or Rust; demonstrable projects in at least two of these languages.
- Strong experience with distributed systems, microservices architecture, concurrency (asyncio, goroutines, multithreading), and message brokers (Kafka, RabbitMQ, AWS SQS).
- Proven experience designing and operating web crawling and scraping systems, including headless browsers (Puppeteer, Playwright, Selenium), HTTP clients, HTML parsing (BeautifulSoup, lxml, Cheerio), and selector strategies (CSS/XPath).
- Familiarity with proxy management, IP rotation, CAPTCHA mitigation strategies, fingerprinting and anti-bot evasion, robots.txt and sitemap parsing, and rate limiting/throttling implementation.
- Practical knowledge of databases (PostgreSQL, MySQL, MongoDB, Redis), search/indexing (Elasticsearch), and data pipeline/ETL patterns.
- Experience with containerization (Docker), orchestration (Kubernetes), and cloud platforms (AWS, GCP, Azure); experience with serverless platforms is a plus.
- Experience with CI/CD tools (Git, Jenkins, GitLab CI, GitHub Actions), logging and observability stacks (ELK, Prometheus, Grafana, Datadog), and performance tuning at scale.
- Strong understanding of security best practices (OAuth, JWT, input sanitization), API rate compliance, error handling and retry/backoff patterns, and caching strategies.
- Comfortable with test-driven development, unit and integration testing, and automation of test suites and deployments.
- Excellent debugging, profiling, and performance optimization skills; ability to triage incidents and implement long-term fixes.
- Strong communication skills, ability to work cross-functionally, mentor teammates, and drive projects from concept to production.
Benefits & conditions
Pulled from the full job description
- 401(k)
- Health insurance
- Vision insurance
- Dental insurance
- Stock options, Fully Remote Opportunity Medical/ Dental / Vision Benefits 401k Annual Bonus Stock options (compelling equity) *, $160000 - $200000 USD per year