> Markdown version of [/jobs/ext/2711283-enterprise-platform-architect](https://www.wearedevelopers.com/jobs/ext/2711283-enterprise-platform-architect). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Enterprise Platform Architect - **Company:** ONE STOP COLLECTIBLE CORP - **Location:** Palo Alto, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Microsoft Azure, Cloud Engineering, Databases, Disaster Recovery, Distributed Systems, Failover, Fault Tolerance, Load Testing, Working Model 2D, AI Infrastructure, Large Language Models, Information Technology - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/enterprise-platform-architect-evolver-9765715 ## About the Role Experience: 10+ years of experience architecting and building enterprise SaaS platforms or similarly complex production systems. Education: Bachelor's or Master's degree in Computer Science, Engineering, or a related field. Technical Skills: * Deep expertise in distributed systems, reliability, scalability, performance, and cloud architecture. * Strong experience with Azure or other hyperscale cloud platforms. * Strong understanding of databases, APIs, networking, storage, containers, and distributed compute. * Familiarity with AI/LLM infrastructure, model APIs, inference architectures, and AI economics. * Experience with observability, load testing, capacity planning, resilience engineering, and disaster recovery. * Ability to make sound architectural tradeoffs across reliability, performance, complexity, and cost. * Ability to lead architecture and drive execution across multiple engineering teams. ## Description You will work across Engineering, Cloud Infrastructure, DevOps, QA, and AI teams to establish measurable requirements, identify architectural platform bottlenecks and opportunities, and define enhancements. Responsibilities Availability & Resilience * Define measurable availability, resilience, and recovery requirements. * Architect for failures across infrastructure, APIs, databases, external dependencies, and AI models. * Design redundancy, failover, retries, timeouts, graceful degradation, and recovery mechanisms. * Identify and eliminate critical single points of failure. * Establish resilience, failover, and recovery testing with clear production-readiness metrics. * Identify gaps, drive remediation, and validate readiness for enterprise production. Scalability & Performance * Define measurable targets for throughput, concurrency, latency, document size, storage growth, and model capacity. * Architect the platform to scale predictably across customers, workloads, and data volumes. * Lead capacity planning across compute, storage, databases, networking, and AI infrastructure. * Establish load, stress, endurance, and performance-testing standards. * Identify and eliminate architectural and performance bottlenecks. * Maintain performance benchmarks and ensure the platform meets enterprise-scale requirements before production. Cost Efficiency * Define and track platform unit economics, including cost per transaction, workflow, and AI execution. * Establish cost targets and identify the primary drivers of platform economics. * Optimize model selection, routing, caching, batching, and reuse. * Move workloads from expensive LLM reasoning to code, ML, smaller models, or deterministic systems where appropriate. * Improve infrastructure utilization and eliminate unnecessary computation. * Ensure the platform remains economically viable as workload volume and complexity scale. ## Related Videos - [Beyond GPT: Building Unified GenAI Platforms for the Enterprise of Tomorrow](https://www.wearedevelopers.com/videos/1525-beyond-gpt-building-unified-genai-platforms-for-the-enterprise-of-tomorrow) - [System Resilience: Surviving the Software Storm](https://www.wearedevelopers.com/videos/874-system-resilience-surviving-the-software-storm) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Load Testing AI: Aiming at a Moving Target](https://www.wearedevelopers.com/videos/1914-load-testing-ai-aiming-at-a-moving-target) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Building Sovereign AI: Lessons from Deploying Secure RAG Systems using Confidential Computing](https://www.wearedevelopers.com/videos/100108-building-sovereign-ai-lessons-from-deploying-secure-rag-systems-using-confidential-computing) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)