Machine Learning Engineer

Vanguard
Malvern, PA, United States
2 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Training Data Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Microsoft Azure Big Data Cloud Computing Configuration Management Databases Data Discovery Extract Transform Load (ETL)
+24 more
Software Design Patterns Distributed Computing Environment Distributed Systems Python (Programming Language) Machine Learning Language Modeling NoSQL Systems Development Life Cycle Software Tools Azure Machine Learning Requirements Management Software Engineering Scripting Large Language Models Prompt Engineering Model Validation Pyspark Information Technology Data Management Api Design Software Version Control Data Pipelines Software Library Programming Languages

Job description

We are seeking an experienced Machine Learning Engineer to join our AI/ML Engineering team. You will be responsible for developing and optimizing complex data pipelines, integrating model pipelines, and building scalable AI/ML solutions, including large language models (LLMs). The ideal candidate will possess a robust background in traditional machine learning, applied GenAI, and significant experience with large datasets and AWS cloud-based AI/ML services.

Supports and performs the development and programming of machine learning integrated software algorithms to structure, analyze, and leverage data in a production environment.

Core Responsibilities

  • Leverages data pipeline designs and supports the development of data pipelines to support model development. Proficient with software tools that develop data pipelines in a distributed computing environment (PySprak, GlueETL).
  • Supports integration of model pipelines in a production environment. Develops understanding of SDLC for model production.
  • Reviews pipeline designs, makes data model design changes as needed. Documents and reviews design changes with data science teams.
  • Supports data discovery & automated ingestion for model development. Performs detailed analysis of raw data sources for data quality, applies business context, and model development needs.
  • Engages with internal stakeholders to understand and probe business processes in order to develop hypotheses. Brings structure to requests and translates requirements into an analytic approach. Participates in and influences ongoing business planning and departmental prioritization activities.
  • Runs model monitoring scripts, follows process for alerts to management as needed. Addresses issues found in data pipelines from model monitoring alerts.
  • Participates in special projects and performs other duties as assigned.

Requirements

  • Undergraduate degree or equivalent experience; a graduate degree is preferred.
  • Minimum of 5 years of relevant work experience.
  • At least 3 years of hands-on experience designing ETL pipelines using AWS services (e.g., Glue, SageMaker).
  • Proficiency in programming languages, particularly Python (including PySpark, PySQL) and familiarity with machine learning libraries and frameworks.
  • Strong understanding of cloud technologies, including AWS and Azure, and experience with NoSQL databases.
  • Familiarity with Feature Store usage, LLMs, GenAI, RAG, Prompt Engineering, and Model Evaluation.
  • Experience with API design and development is a plus.
  • Solid understanding of software engineering principles, including design patterns, testing, security, and version control.
  • Knowledge of Machine Learning Development Lifecycle (MDLC) best practices and protocols.
  • Understanding of solution architecture for building end-to-end machine learning data pipelines., Algorithms, Amazon Web Services (AWS), Application Programming Interface (API), Artificial Intelligence (AI), Best Practices, Business Model, Business Plan, Business Processes, Cloud Computing, Computer Systems, Data Analysis, Data Management, Data Modeling, Data Quality, Data Science, Database Extract Transform and Load (ETL), Design Patterns Programming Methodologies, Distributed Computing, Establish Priorities, Machine Learning, Microsoft Windows Azure, Modeling Languages, NoSQL, Problem Solving Skills, Product Lifecycle, Production Systems, Programming Languages, Python Programming/Scripting Language, Requirements Management, Scalable System Development, Scripting (Scripting Languages), Software Development Lifecycle (SDLC), Software Engineering, Source Code/Configuration Management (SCM), Team Player, Test Design, Training Data Sets

About the company

About Vanguard

At Vanguard, we don’t just have a mission-we’re on a mission.

To work for the long-term financial wellbeing of our clients. To lead through product and services that transform our clients’ lives. To learn and develop our skills as individuals and as a team. From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.

How We Work

Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

2:10 min

Why organizations combine big data and machine learning

Ayon Roy · LIVE

Videos

See all

Related articles

See all