Data Engineer (Python / PySpark / SQL)

WilsonCTS
San Jose, CA, United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$104,000.0 - $145,600.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Big Data Computer Programming Information Engineering Data Infrastructure Database Development Software Debugging Distributed Computing Environment Python (Programming Language) Machine Learning Object-Oriented Software Development Operational Databases
+7 more
SQL Databases Feature Engineering Sql Optimization System Availability Pyspark Data Analytics Data Pipelines

Job description

  • Design, build, and maintain scalable feature engineering pipelines for analytics and machine learning applications.
  • Develop efficient, incremental data pipelines using PySpark to process enterprise-scale datasets.
  • Create and optimize data models that support reporting, analytics, and predictive modeling.
  • Monitor production data pipelines, troubleshoot issues, and ensure high availability and data quality.
  • Collaborate closely with data scientists, analysts, and engineering teams to prioritize business requirements.
  • Write clean, maintainable, object-oriented Python code following best practices.
  • Optimize SQL queries and data processing workflows for performance and scalability.
  • Support continuous improvements to data infrastructure and pipeline reliability.
  • Mentor team members and promote best practices for feature engineering and data development.

Requirements

  • 5+ years of overall experience in Data Engineering, Data Analytics, or a related field.
  • 2+ years of hands-on experience developing enterprise-scale solutions using PySpark.
  • 4+ years of experience with data modeling and designing scalable data solutions.
  • Strong programming skills in Python, including object-oriented programming concepts.
  • Advanced SQL skills with experience querying and transforming large datasets.
  • Experience building and maintaining production data pipelines.
  • Strong troubleshooting and debugging skills in production environments.
  • Excellent communication and collaboration skills with technical stakeholders., * Experience supporting machine learning feature engineering pipelines.
  • Familiarity with large-scale distributed data processing environments.
  • Experience working in Agile development environments.
  • Passion for mentoring teammates and sharing technical knowledge.

Benefits & conditions

  • Work on impactful, data-driven products used by millions of customers.
  • Collaborate with talented engineers, analysts, and data scientists.
  • Long-term contract with competitive W2 compensation.
  • Opportunity to contribute to modern data engineering and machine learning initiatives.

If you’re passionate about building scalable data solutions and enjoy solving complex data challenges with Python, PySpark, and SQL, we’d love to hear from you!

MSP

Pay: $50.00 - $70.00 per hour

Benefits:

  • Health insurance

About the company

Our client is a leading global technology company recognized for building innovative consumer software and data-driven digital products. They are seeking a Data Engineer II to join a high-impact analytics and machine learning team responsible for developing scalable data solutions that power business insights and intelligent decision-making.

This is an excellent opportunity for a data professional who enjoys building robust data pipelines, working with large-scale datasets, and collaborating with analysts and data scientists to enable advanced analytics and machine learning initiatives.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

1:14 min

Evolution of distributed SQL database architectures

Wei Hu Wei Hu · World Congress 2024

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all