Web Scraping Engineer (U.S. Citizen)

Powerhouse Institute Inc
United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$95,000.0 - $105,000.0
Working hours
Regular working hours
Job source

Tech stack

HTML Java (Programming Language) JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Big Data Cascading Style Sheets (CSS) Databases Data Cleansing Information Engineering Web Scraping
+14 more
Data Mining Data Visualization Software Debugging Information Sciences JSON Python (Programming Language) Machine Learning Selenium SQL Databases Extensible Markup Language (XML) Data Storage Management Information Technology Api Design Data Pipelines

Job description

  • Position supports the development and maintenance of the agency’s web scraping infrastructure. The position is responsible for extracting data from various websites and APIs, ensuring data quality and accuracy, and optimizing the scraping process for efficiency. Duties include:
  • Develop and maintain web scraping scripts and tools to extract data from websites and APIs.
  • Collaborate with cross-functional teams to understand data requirements and implement scraping solutions accordingly.
  • Monitor and troubleshoot scraping processes to ensure data quality and accuracy.
  • Optimize scraping scripts for performance and efficiency, considering factors such as speed, scalability, and resource utilization.
  • Stay up to date with the latest web scraping techniques, tools, and best practices.
  • Conduct data analysis and validation to ensure the integrity of scraped data.
  • Collaborate with data engineering and data science teams to integrate scraped data into our data pipelines and systems.
  • Document and communicate technical solutions, processes, and best practices to team members.

Requirements

Do you have experience in XML?, NOTE: This is a full-time employment opportunity (No C2C, subcontractor or 1099 engagements, please.). The candidate MUST be a U.S. Citizen and able to complete/pass/hold at a minimum public trust investigation (active public trust is preferred). This is a remote opportunity; candidate must reside in the continental United States., * Must be a U.S. Citizen (no dual status) as mandated by our government client.

  • Must be able to complete/pass/hold at a minimum public trust investigation (active public trust is preferred).
  • Must reside in the continental United States.
  • 3+ years of professional experience in web scraping or a similar role.
  • Proficiency in Python and Java and experience with web scraping libraries such as Beautiful Soup, Scrapy, or Selenium.
  • Knowledge of AI/machine learning techniques for data extraction and classification.
  • Understanding of HTML, CSS, and JavaScript to navigate and interact with websites.
  • Experience working with APIs and handling different data formats (JSON, XML, etc.).
  • Familiarity with database systems and SQL for data storage and retrieval.
  • Familiarity of data cleaning and preprocessing techniques to ensure data quality.
  • Strong problem-solving skills and ability to troubleshoot and debug scraping issues.
  • Excellent communication and collaboration skills to work effectively in a team environment.
  • Attention to detail and ability to handle large volumes of data efficiently.
  • BS/BA degree in Computer Science, Information Sciences, or related IT discipline. Additional years of related professional experience can be substituted for a BS/BA degree.

Additional Qualifications (a PLUS):

  • Experience with cloud platforms for scalable web scraping infrastructure.
  • Familiarity with data visualization tools and techniques.
  • Understanding of legal and ethical considerations related to web scraping.

Compensation decisions depend on a wide range of factors, including but not limited to skill sets, experience and training, security clearances, licensure and certifications, location, and other business and organizational needs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:43 min

Scaling web scraping infrastructure to bypass strict security restrictions

Tim Ruscica · Coffee With Developers

2:21 min

Projecting external HTML content using default and named slots

Rowdy Rabouw Rowdy Rabouw · WWC 2022

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:35 min

Selecting Selenium for browser automation and testing

Benjamin Bischoff Benjamin Bischoff · Europe 2026 Virtual

1:19 min

Defining web scraping and recognizing proper use cases

Lars Kölker · WWC 2023

6:12 min

Streaming HTML content natively using declarative processing instructions

Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all