Platform Architect

North Carolina State University
Raleigh, NC, United States
1 day ago
Apply on jobs.ncsu.edu
Prepare application

Role details

Contract type
Temporary contract
Employment type
Part-time / full-time
Experience level
Expert
Experience required
2 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Software Quality Computer Programming Continuous Integration Information Engineering Data Sharing Data Stores Decision Support Systems DevOps
+37 more
Distributed Systems Identity and Access Management Python (Programming Language) Machine Learning Network Planning and Design Network Segmentation NoSQL Open Source Technology Role-Based Access Control Cloud Services Tensorflow Zero Trust Network Access Azure Machine Learning Software Engineering Cloud Platform System Feature Engineering Data Ingestion Pytorch Large Language Models Deep Learning Reliability of Systems Generative AI Infrastructure as Code (IaC) Cloudformation Containerization AI Platforms Scikit Learn Kubernetes Infrastructure Automation Frameworks Data Lineage Enterprise Integration Machine Learning Operations Terraform Software Version Control Data Pipelines Serverless Computing Docker

Job description

The Platform Architect will be an employee in the Department of Biological and Agricultural Engineering (BAE) jointly administered by the College of Agriculture and Life Sciences (CALS) and the College of Engineering at NC State University. The department actively contributes to the academic, research and extension missions of the University. Wolfpack Perks and Benefits As a Pack member, you belong here, and can enjoy exclusive perks designed to enhance your personal and professional well-being. As you consider this opportunity, we encourage you to review our Employee Value Proposition and learn more about what makes NC State the best place to learn and work for everyone. What we offer, Disclaimer: Perks and Benefit eligibility is based on Part-Time or Full-Time Employment status. Eligibility and Employer Sponsored Plans can be found within each of the links offered. Essential Job Duties This is a repost; previous applicants need not reapply.

This position supports the REFRAME project, a multi-year, externally funded research initiative led by NC State University. The project aims to deliver a modular, AI-enabled, open-source platform for feedstock-agnostic evaluation of future biomass. It will integrate legacy open-source models, novel surrogates, shared metadata schemas, and a large language model (LLM) to enhance accessibility for varied stakeholders.

Leveraging well-characterized agricultural and food processing residues, REFRAME will generate insights into historically underfunded circular biomass streams and provide deployment-ready tools for end-to-end scenario analysis, counterfactual modeling, and decision support across the value chain. The Platform Architect will play a central role in designing and operationalizing the technical backbone of this platform in close collaboration with faculty, research staff, and external partners.

The Platform Architect is a senior technical leadership role that will lead the software and infrastructure development for the digital backbone for the REFRAME project. This digital backbone is the central nervous system for developing, deploying, monitoring, and governing all machine learning models related to the project. The platform architect will be responsible for strategic roadmapping of infrastructure, system design, tech-stack selection, devOps, security & compliance, workflow integration, and process optimization, ensuring the platform’s scalability, reliability, and efficiency. The architect will also work closely with the project manager and software development personnel, bridging gaps between project personnel and development personnel to ensure that ML models can transition from research workflows to scalable, secure, and reliable production environments efficiently.

The Platform Architect is a senior technical leadership role responsible for defining the strategy, design, and implementation of our solution-specific Machine Learning Operations (MLOps) platform. This platform is the central nervous system for developing, deploying, monitoring, and governing all machine learning models related to the project. The Architect bridges the gap between data science experimentation and production engineering rigor, ensuring that ML models can transition from research workflows to scalable, secure, and reliable production environments efficiently.

Platform Vision and Architecture :

  • Strategic Roadmapping: Define the technical vision and multi-year roadmap for the end-to-end ML platform, aligning it with project and research objectives and data science needs (including GenAI/LLM capabilities).
  • System Design: Architect, design, and document a robust, scalable, and cost-efficient platform covering the entire ML lifecycle: Data Ingestion, Feature Store, Model Training, Model Registry, Model Serving/Inference, and Monitoring.
  • Technology Selection: Evaluate, select, and integrate appropriate cloud-native services (AWS, Azure, or GCP), hybrid and on-prem computing resources, and open-source MLOps tools (e.g., Kubeflow, MLflow, Airflow) to build a cohesive ecosystem.

MLOps and Delivery Excellence :

  • Automation: Lead the implementation of MLOps pipelines using CI/CD practices to automate model training, testing, validation, deployment, and automated retraining workflows.
  • Infrastructure as Code (IaC): Design and enforce IaC standards (Terraform/CloudFormation) for provisioning and managing all underlying compute, networking, and storage resources (e.g., Kubernetes clusters, GPU instances).
  • Feature Engineering: Define shared data and feature management patterns to ensure consistency and reuse across model training and inference workflows.

Governance, Security, and Compliance:

  • Model Governance: Implement standards for model versioning, lineage tracking, and compliance to ensure models are traceable, reproducible, and meet ethical AI principles.
  • Security: Architect security measures, including network segmentation, access control (IAM/RBAC), and data encryption across the entire ML pipeline.
  • Performance & Cost Optimization: Establish comprehensive monitoring and alerting for model performance (drift, bias) and infrastructure metrics, driving continuous optimization for performance and cloud costs (FinOps).

Technical Leadership and Collaboration:

  • Cross-Functional Partnership: Act as the primary technical point of contact, collaborating closely with Data Scientists, Data Engineers, Software Engineers, Product Managers, domain scientists, industry partners, process engineering and systems modeling experts.
  • Documentation: Create and maintain high-quality architectural diagrams, reference implementations, and technical documentation.

Other Responsibilities

  • Other duties as assigned.

Requirements

Minimum Education and Experience

  • PhD or Master’s degree in relevant field with 2+ years of experience in the following areas: Software Engineering, Data Engineering, or Platform Engineering and ML/AI platform architecture / MLOps OR
  • Bachelor’s degree in relevant field with 4+ years of experience in the follow areas: Software Engineering, Data Engineering, or Platform Engineering and ML/AI platform architecture / MLOps

Other Required Qualifications Areas of technical proficiency include:

  • Cloud Expertise: Expert-level proficiency in at least one major cloud platform (AWS, Azure, or GCP) and its ML-specific services (e.g., AWS SageMaker, GCP Vertex AI, Azure ML).
  • Containerization & Orchestration: Mastery of Kubernetes and Docker for containerizing and distributed ML workloads.
  • Coding Proficiency: Strong programming skills in Python and experience with core ML/Deep Learning frameworks (e.g., TensorFlow, PyTorch, scikit-learn).
  • Data & Infrastructure: Deep knowledge of distributed systems, data pipelines, and Infrastructure as Code., * Advanced ML Experience: Experience designing infrastructure for advanced ML use cases, such as Generative AI, Large Language Models (LLMs), or Real-Time Inference systems.
  • Networking/Security: Familiarity with zero-trust architecture principles and network design for secure MLOps environments.
  • Data Store Experience: Practical experience with NoSQL/Vector Databases and Feature Stores.
  • Problem-Solving: The ability to analyze complex issues across the entire stack and devise effective solutions.
  • Adaptability & Learning: The tech landscape changes rapidly, so the ability to quickly learn new languages, frameworks, or tools is essential.
  • Communication & Collaboration: Effectively communicating technical details to both technical and non-technical stakeholders, and working seamlessly within a team.
  • Attention to Detail: Ensuring high standards for code quality, system reliability, and user-facing usability.

Required License(s) or Certification(s) N/A Valid NC Driver’s License required No Commercial Driver’s License required No Recruitment Dates and Special Instructions

Benefits & conditions

Anticipated Hiring Range Commensurate with Experience ($120,000- $150,000) Work Schedule Monday - Friday, 8am - 5pm (40 hours/week, with flexibility)

About the company

NC State University participates in E-Verify. Federal law requires all employers to verify the identity and employment eligibility of all persons hired to work in the United States.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.ncsu.edu
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all