Enterprise Platform Architect

ONE STOP COLLECTIBLE CORP
Palo Alto, United States
2 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
10 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Microsoft Azure Cloud Engineering Databases Disaster Recovery Distributed Systems Failover Fault Tolerance Load Testing Working Model 2D AI Infrastructure
+2 more
Large Language Models Information Technology

Job description

You will work across Engineering, Cloud Infrastructure, DevOps, QA, and AI teams to establish measurable requirements, identify architectural platform bottlenecks and opportunities, and define enhancements.

Responsibilities

Availability & Resilience

  • Define measurable availability, resilience, and recovery requirements.

  • Architect for failures across infrastructure, APIs, databases, external dependencies, and AI models.

  • Design redundancy, failover, retries, timeouts, graceful degradation, and recovery mechanisms.

  • Identify and eliminate critical single points of failure.

  • Establish resilience, failover, and recovery testing with clear production-readiness metrics.

  • Identify gaps, drive remediation, and validate readiness for enterprise production.

Scalability & Performance

  • Define measurable targets for throughput, concurrency, latency, document size, storage growth, and model capacity.

  • Architect the platform to scale predictably across customers, workloads, and data volumes.

  • Lead capacity planning across compute, storage, databases, networking, and AI infrastructure.

  • Establish load, stress, endurance, and performance-testing standards.

  • Identify and eliminate architectural and performance bottlenecks.

  • Maintain performance benchmarks and ensure the platform meets enterprise-scale requirements before production.

Cost Efficiency

  • Define and track platform unit economics, including cost per transaction, workflow, and AI execution.

  • Establish cost targets and identify the primary drivers of platform economics.

  • Optimize model selection, routing, caching, batching, and reuse.

  • Move workloads from expensive LLM reasoning to code, ML, smaller models, or deterministic systems where appropriate.

  • Improve infrastructure utilization and eliminate unnecessary computation.

  • Ensure the platform remains economically viable as workload volume and complexity scale.

Requirements

Experience: 10+ years of experience architecting and building enterprise SaaS platforms or similarly complex production systems.

Education: Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.

Technical Skills:

  • Deep expertise in distributed systems, reliability, scalability, performance, and cloud architecture.

  • Strong experience with Azure or other hyperscale cloud platforms.

  • Strong understanding of databases, APIs, networking, storage, containers, and distributed compute.

  • Familiarity with AI/LLM infrastructure, model APIs, inference architectures, and AI economics.

  • Experience with observability, load testing, capacity planning, resilience engineering, and disaster recovery.

  • Ability to make sound architectural tradeoffs across reliability, performance, complexity, and cost.

  • Ability to lead architecture and drive execution across multiple engineering teams.

Benefits & conditions

  • Competitive Compensation: Tailored to your experience and skill set.

  • Flexible Work Arrangements: Hybrid working model for work-life balance.

  • Career Growth: Opportunities for professional development and leadership roles.

  • Innovative Culture: Work on transformative technologies and make an impact in the AI space.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

2:59 min

Scaling clusters and handling automated replica failover

Jürgen Pilz · World Congress 2023

3:04 min

Database evolution and the funding behind vector databases

Erik Bamberg · LIVE

1:45 min

Defining objectives and methodologies for systematic load testing

Kirill Kulikov · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

5:25 min

Implementing redundancy, failover, and architectural load balancing patterns

Mihaela-Roxana Ghidersa · LIVE

Videos

See all

Related articles

See all