Infrastructure Engineer

Digital Intelligence Systems, LLC
New York, NY, United States
25 days ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Data Analysis Microsoft Azure Cloud Computing Configuration Management Identity and Access Management Role-Based Access Control Real-Time Operating Systems Reliability Engineering SQL Databases Information Technology Data Analytics

Job description

Experteer Overview In this role you will lead capacity planning and optimization for a high-demand Azure environment, ensuring resilient, scalable capacity across compute, storage, network, and platform services. You will partner with site reliability teams to meet performance and reliability targets while upholding rigorous controls. This position offers a chance to shape critical cloud infrastructure in a regulated setting through evidence-based practices and continuous monitoring. You will work to balance performance, resilience, and cost while driving capacity governance and improvement. Compensation / Benefits * Own and manage the end-to-end Capacity Management operating model for Azure services in high-criticality projects, including planning, modeling, forecasting, monitoring, tuning, and governance * Ensure sufficient capacity and buffers to meet SLAs, RTOs, RPOs, and regulatory requirements with regional considerations and ongoing monitoring * Collaborate with site reliability engineers to implement capacity practices via IaC, gated change controls, performance baselines, autoscaling, and resilience patterns * Contribute to compliance evidence such as system security plans, control narratives, corrective action plans, and continuous monitoring artifacts * Develop and maintain service-level capacity models across Azure components with buffer standards and validation against demand and failover scenarios * Design and tune autoscaling policies with guardrails on quotas and throttling for stability * Analyze utilization, throughput, and performance metrics to drive tuning actions, reservations, and architectural improvements * Forecast future demand from product roadmaps and business growth, translating into capacity plans and procurement strategies * Participate in change review processes ensuring capacity and security impacts are assessed and documented * Manage cryptographic mechanisms under change control with versioned inventories and validation * Oversee external services supporting capacity for compliance and standards * Enforce region-restriction policies for processing, storage, backups, and DR in high-impact systems * Balance performance, resilience, and cost through rightsizing, tiering, and scheduled scaling * Perform criticality analysis to prioritize capacity needs and align related policies * Validate DR capacity and maintain buffers for failover without impacting steady-state operations * Define, measure, and report capacity KPIs and prepare dashboards for compliance, audits, and decision-making Tasks * Bachelor’s degree in computer science or a related technical field; advanced degrees are a plus * 10+ years of experience in infrastructure capacity and performance engineering across compute, storage, network, and platform services * Experience in regulated environments and familiarity with high-assurance and evidence collection standards * Strong data analysis skills to translate telemetry into actionable insights * Proven process discipline with change/configuration management and cross-team dependencies * Experience coordinating diverse engineering teams across multiple platforms and tools * Proficiency with Azure services and concepts (identity management, SQL, storage, networking, security policies, RBAC) * Excellent communication skills for stakeholder management and executive-level presentation * Ability to translate complex technical requirements into clear plans, milestones, and measurable outcomes Key requirements *

Requirements

for services supporting capacity for compliance and standards * Enforce region-restriction policies for processing, storage, backups, and DR in high-impact systems * Balance performance, resilience, and cost through rightsizing, tiering, and scheduled scaling * Perform criticality analysis to prioritize capacity needs and align related policies * Validate DR capacity and maintain buffers for failover without impacting steady-state operations * Define, measure, and report capacity KPIs and prepare dashboards for compliance, audits, and decision-making Tasks * Bachelor’s degree in computer science or a related technical field; advanced degrees are a plus * 10+ years of experience in infrastructure capacity and performance engineering across compute, storage, network, and platform services * Experience in regulated environments and familiarity with high-assurance and evidence collection standards * Strong data analysis skills to translate telemetry into actionable insights * Proven process discipline with change/configuration management and cross-team dependencies * Experience coordinating diverse engineering teams across multiple platforms and tools * Proficiency with Azure services and concepts (identity management, SQL, storage, networking, security policies, RBAC) * Excellent communication skills for stakeholder management and executive-level presentation * Ability to translate complex technical requirements into clear plans, milestones, and measurable outcomes Key requirements *

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:32 min

Structuring platforms for new services and data analytics

Nevelina Aleksandrova · LIVE

1:48 min

Automating exploratory data analysis within training pipelines

Dora Petrella · World Congress 2023

1:14 min

Evolution of distributed SQL database architectures

Wei Hu Wei Hu · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:36 min

Performing exploratory data analysis to uncover underlying patterns

Julian Joseph · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all