Senior Compute Platform Engineer

GSK
Greater London, UK
2 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology Airflow Amazon Web Services Application Performance Management Computing Platforms Confluence JIRA Automation of Tests Microsoft Azure C++ (Programming Language) Cloud Computing
+28 more
CMake Code Review Continuous Integration Information Engineering DevOps Distributed Systems Github Python (Programming Language) OpenMP Open Source Technology Package Management Systems Cloud Services Software Deployment Software Engineering Workflow Management Systems Circleci Data Logging Google Cloud High Performance Computing Gitlab Git Kubernetes Infrastructure Automation Frameworks Information Technology Build Tools Hardware Infrastructure Docker Jenkins

Job description

  • Design, build, and operate tools, services, workflows, etc. that deliver high value through solutions to key business problems.
  • Responsible for development of key components of a hybrid on-prem/cloud compute platform for both interactive and scalable batch computing and establishing processes and workflows to transition existing HPC users and teams to this platform.
  • Responsible for code-driven environment, applications, and container/image builds as well as CI/CD-driven application deployments.
  • Consult science users on application scalability to petabytes of data by deeply understanding software engineering, algorithms, and underlying hardware infrastructure and their impact on performance.
  • Confidently optimise design and execution of complex solutions within large-scale distributed computing environments.
  • Produce well-engineered software, including appropriate automated test suites, technical documentation, and operational strategy.
  • Ensure consistent application of platform abstractions to ensure quality and consistency with respect to logging and lineage.
  • Fully versed in coding best practices and ways of working, and participate in code reviews and partnering to improve the team’s standards.
  • Adhere to QMS framework and CI/CD best practices and help guide improvements to them that improve ways of working.
  • Provide leadership to team members to help others get the job done right.

Requirements

  • Bachelor’s degree in data engineering, Computer Science, Software Engineering or related discipline.
  • 6+ years of professional experience.
  • Experience with Python.
  • Experience with Cloud.
  • Experience with High Performance Compute (HPC)., * Deep knowledge and use of at least one common programming language: e.g., Python, C++, Java, including toolchains for documentation, testing, and operations / observability.
  • Deep expertise in modern software development tools / ways of working (e.g., git/GitHub, DevOps tools, metrics / monitoring).
  • Deep cloud expertise (e.g., AWS, Google Cloud, Azure), including infrastructure-as-code tools and scalable compute technologies, such as Google Batch and Vertex.
  • Experience with CI/CD implementations using git and a common CI/CD stack (e.g., Azure DevOps, CloudBuild, Jenkins, CircleCI, GitLab).
  • Deep expertise with Docker, Kubernetes, and the larger CNCF ecosystem including experience with application deployment tools such as Helm.
  • Experience with low-level application build tools (make, CMake) as well as automated build systems such as Spack or Easybuild.
  • Experience with workflow orchestration with tools such as Argo Workflow, Airflow, and scientific workflow tools such as Nextflow, Snakemake, VisTrails, or Cromwell.
  • Experience with application performance tuning and optimization, including in parallel and distributed computing paradigms and communication libraries such as MPI, OpenMP, Gloo, including deep understanding of the underlying systems (hardware, networks, storage) and their impact on application performance.
  • Demonstrated excellence with agile software development environments using tools like Jira and Confluence.
  • Deep familiarity with the tools, techniques, optimizations in high-performance applications space, including engagement with the open-source community (and potentially making contributions to such tools).

About the company

At GSK, we want to supercharge our data capability to better understand our patients and accelerate our ability to discover vaccines and medicines. The Onyx Research Data Platform organization represents a major investment by GSK R&D and Digital & Tech, designed to deliver a step-change in our ability to leverage data, knowledge, and prediction to find new medicines.

The Onyx Research Data Platform organization is a full-stack shop consisting of product and portfolio leadership, data engineering, infrastructure and DevOps, data / metadata / knowledge platforms, and AI/ML and analysis platforms. It is geared toward:

  • Building a next-generation, metadata- and automation-driven data experience for GSK’s scientists, engineers, and decision-makers, increasing productivity and reducing time spent on “data mechanics.”
  • Providing best-in-class AI/ML and data analysis environments to accelerate our predictive capabilities and attract top-tier talent.
  • Aggressively engineering our data at scale, as one unified asset, to unlock the value of our unique collection of data and predictions in real-time.

Our Compute Platform Engineering team is building a first-in-class platform of toolchains and workflows that accelerate application development, scale up computational experiments, and integrate all computation with project metadata, logs, experiment configuration and performance tracking over abstractions that encompass Cloud and High-Performance Computing. This metadata-forward, CI/CD-driven platform represents and enables the entire application and analysis lifecycle including interactive development and explorations (notebooks), large-scale batch processing, observability and production application deployments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:15 min

Defining classes and building packages with pybind11

Konstantin Bespalov · World Congress 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

Videos

See all

Related articles

See all