> Markdown version of [/jobs/ext/2629596-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2629596-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer - **Company:** Digital Intelligence Systems, LLC - **Location:** New York, NY, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Data Analysis, Microsoft Azure, Cloud Computing, Configuration Management, Identity and Access Management, Role-Based Access Control, Real-Time Operating Systems, Reliability Engineering, SQL Databases, Information Technology, Data Analytics - **Published:** August 12, 2026 - **Apply:** https://us.experteer.com/career/view-jobs/infrastructure-engineer-new-york-ny-usa-58909926 ## About the Role for services supporting capacity for compliance and standards * Enforce region-restriction policies for processing, storage, backups, and DR in high-impact systems * Balance performance, resilience, and cost through rightsizing, tiering, and scheduled scaling * Perform criticality analysis to prioritize capacity needs and align related policies * Validate DR capacity and maintain buffers for failover without impacting steady-state operations * Define, measure, and report capacity KPIs and prepare dashboards for compliance, audits, and decision-making Tasks * Bachelor's degree in computer science or a related technical field; advanced degrees are a plus * 10+ years of experience in infrastructure capacity and performance engineering across compute, storage, network, and platform services * Experience in regulated environments and familiarity with high-assurance and evidence collection standards * Strong data analysis skills to translate telemetry into actionable insights * Proven process discipline with change/configuration management and cross-team dependencies * Experience coordinating diverse engineering teams across multiple platforms and tools * Proficiency with Azure services and concepts (identity management, SQL, storage, networking, security policies, RBAC) * Excellent communication skills for stakeholder management and executive-level presentation * Ability to translate complex technical requirements into clear plans, milestones, and measurable outcomes Key requirements * ## Description Experteer Overview In this role you will lead capacity planning and optimization for a high-demand Azure environment, ensuring resilient, scalable capacity across compute, storage, network, and platform services. You will partner with site reliability teams to meet performance and reliability targets while upholding rigorous controls. This position offers a chance to shape critical cloud infrastructure in a regulated setting through evidence-based practices and continuous monitoring. You will work to balance performance, resilience, and cost while driving capacity governance and improvement. Compensation / Benefits * Own and manage the end-to-end Capacity Management operating model for Azure services in high-criticality projects, including planning, modeling, forecasting, monitoring, tuning, and governance * Ensure sufficient capacity and buffers to meet SLAs, RTOs, RPOs, and regulatory requirements with regional considerations and ongoing monitoring * Collaborate with site reliability engineers to implement capacity practices via IaC, gated change controls, performance baselines, autoscaling, and resilience patterns * Contribute to compliance evidence such as system security plans, control narratives, corrective action plans, and continuous monitoring artifacts * Develop and maintain service-level capacity models across Azure components with buffer standards and validation against demand and failover scenarios * Design and tune autoscaling policies with guardrails on quotas and throttling for stability * Analyze utilization, throughput, and performance metrics to drive tuning actions, reservations, and architectural improvements * Forecast future demand from product roadmaps and business growth, translating into capacity plans and procurement strategies * Participate in change review processes ensuring capacity and security impacts are assessed and documented * Manage cryptographic mechanisms under change control with versioned inventories and validation * Oversee external services supporting capacity for compliance and standards * Enforce region-restriction policies for processing, storage, backups, and DR in high-impact systems * Balance performance, resilience, and cost through rightsizing, tiering, and scheduled scaling * Perform criticality analysis to prioritize capacity needs and align related policies * Validate DR capacity and maintain buffers for failover without impacting steady-state operations * Define, measure, and report capacity KPIs and prepare dashboards for compliance, audits, and decision-making Tasks * Bachelor's degree in computer science or a related technical field; advanced degrees are a plus * 10+ years of experience in infrastructure capacity and performance engineering across compute, storage, network, and platform services * Experience in regulated environments and familiarity with high-assurance and evidence collection standards * Strong data analysis skills to translate telemetry into actionable insights * Proven process discipline with change/configuration management and cross-team dependencies * Experience coordinating diverse engineering teams across multiple platforms and tools * Proficiency with Azure services and concepts (identity management, SQL, storage, networking, security policies, RBAC) * Excellent communication skills for stakeholder management and executive-level presentation * Ability to translate complex technical requirements into clear plans, milestones, and measurable outcomes Key requirements * ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Data Science on Software Data](https://www.wearedevelopers.com/videos/162-data-science-on-software-data) - [How Data is Shaping our Games](https://www.wearedevelopers.com/videos/176-how-data-is-shaping-our-games) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)