Platform Engineer - AI Agent Infrastructure

SUNNYSTEP PTE. LTD.
Penarth, UK
5 days ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Systems Engineering Audit Trail Automation of Tests Microsoft Azure Backup Devices Business Systems Command-Line Interface Cloud Computing Cloud Engineering
+25 more
Databases Continuous Integration Software Debugging DevOps Programming Tools Disaster Recovery Graph Database Python (Programming Language) Key Management PostgreSQL Machine Learning Shopify TypeScript Management of Software Versions Privacy Controls Large Language Models Prompt Engineering Backend Git Event Driven Architecture Kubernetes Infrastructure Automation Frameworks Front End Software Development Virtual Agents Automation Anywhere

Job description

About the role

We are building Maxify, an AI-native operating platform where agents execute real business workflows through company systems while remaining accountable to source data, permissions, evaluations and measurable outcomes.

We are looking for a hands-on Platform Engineer to build the reusable, multi-tenant infrastructure that makes production AI agents secure, reliable and economical to operate. You will own the shared platform beneath business-specific agents while still shipping practical capabilities end-to-end.

This is not a pure infrastructure, prompt-engineering or research role. You will work across agent runtimes, backend services, data, integrations, deployment and developer tooling.

What you will own

  • Own the multi-tenant agent runtime, orchestration, state and governed memory.
  • Build secure tool execution, authentication, authorization and customer isolation.
  • Design approvals, audit trails, privacy controls and secrets management.
  • Create reusable connectors for APIs, databases and business systems.
  • Build evaluation infrastructure, testing environments and safe release gates.
  • Implement observability for quality, failures, latency, reliability and cost.
  • Establish deployment, versioning, backup, rollback and incident-response mechanisms.
  • Build platform APIs, operator interfaces and developer tooling.
  • Own cloud architecture, infrastructure automation, capacity, availability and disaster recovery.
  • Turn repeated workflow patterns into legible, reusable platform capabilities.

What we are looking for

  • Production experience building and operating multi-tenant backend or platform systems.
  • Strong backend and distributed-systems engineering judgment.
  • Production proficiency in Python and/or TypeScript.
  • Strong PostgreSQL and cloud-hosted data-infrastructure experience.
  • Experience with APIs, queues, asynchronous jobs and event-driven systems.
  • Strong understanding of authentication, authorization, secrets management and data isolation.
  • Experience with CI/CD, infrastructure automation, monitoring and incident response.
  • Experience operating production workloads on AWS, GCP or Azure.
  • Practical knowledge of containers, networking, backups, rollback and disaster recovery.
  • Practical LLM experience including tool use, structured outputs, context management and evaluation.
  • Ability to debug across application, infrastructure and external-service layers.
  • Clear written communication and sound architectural judgment.

Strong signals

  • You have shipped and operated a production agentic or automation platform, not only a demo.
  • You have built multi-tenant systems with explicit isolation and authorization controls.
  • You have designed evaluation, observability, approval or audit systems for AI workflows.
  • You have owned reliability and cloud infrastructure while remaining product-oriented.
  • You use AI development tools extensively while validating outputs and retaining technical ownership.
  • You can explain a system you built end-to-end, what failed and how you prevented recurrence.

Nice to have

  • Experience with OpenAI Agents SDK, LangGraph, OpenClaw or comparable systems.
  • Experience with Supabase or PostgreSQL at production scale.
  • Experience with Lark/Feishu, Shopify, CRM, accounting or enterprise SaaS integrations.
  • Familiarity with knowledge graphs, ontology design or governed memory.
  • Early-stage startup or founding-engineer experience.

Current stack

OpenClaw, Python, TypeScript, Supabase/PostgreSQL, Lark/Feishu APIs and CLI, Shopify and other business-system APIs, Git-based development, automated testing, deployment controls, monitoring and rollback.

What this role is not

  • Prompt writing without software ownership.
  • Pure machine-learning research or model training.
  • Pure DevOps, Kubernetes or cloud administration.
  • Frontend-only or backend-only feature delivery.
  • Building abstractions for hypothetical scale before real workflows require them.

Success in the first 30 days

  • Complete an architecture, security and reliability review and identify the major risks and delivery bottlenecks.
  • Operate the current agent runtime and deployment environment independently.
  • Ship at least one meaningful production platform improvement.
  • Deliver one reusable platform capability with tests, telemetry and rollback.
  • Establish minimum quality, security and observability gates.
  • Document the reusable pattern for onboarding the next customer or workflow.
  • Produce a prioritized 60-day platform roadmap based on evidence from the first month.

These outcomes are subject to timely access and platform readiness.

How we work

  • We move quickly and build for permanence.
  • We prefer small, reversible decisions over speculative architecture.
  • Documentation, testing, evaluation and observability are part of the product.
  • AI accelerates the work; it does not remove engineering accountability.
  • We measure success through customer outcomes, revenue impact, cost reduction and lower operational attention.

Requirements

  • Production experience building and operating multi-tenant backend or platform systems.
  • Strong backend and distributed-systems engineering judgment.
  • Production proficiency in Python and/or TypeScript.
  • Strong PostgreSQL and cloud-hosted data-infrastructure experience.
  • Experience with APIs, queues, asynchronous jobs and event-driven systems.
  • Strong understanding of authentication, authorization, secrets management and data isolation.
  • Experience with CI/CD, infrastructure automation, monitoring and incident response.
  • Experience operating production workloads on AWS, GCP or Azure.
  • Practical knowledge of containers, networking, backups, rollback and disaster recovery.
  • Practical LLM experience including tool use, structured outputs, context management and evaluation.
  • Ability to debug across application, infrastructure and external-service layers.
  • Clear written communication and sound architectural judgment.

Strong signals

  • You have shipped and operated a production agentic or automation platform, not only a demo.
  • You have built multi-tenant systems with explicit isolation and authorization controls.
  • You have designed evaluation, observability, approval or audit systems for AI workflows.
  • You have owned reliability and cloud infrastructure while remaining product-oriented.
  • You use AI development tools extensively while validating outputs and retaining technical ownership.
  • You can explain a system you built end-to-end, what failed and how you prevented recurrence., * Experience with OpenAI Agents SDK, LangGraph, OpenClaw or comparable systems.
  • Experience with Supabase or PostgreSQL at production scale.
  • Experience with Lark/Feishu, Shopify, CRM, accounting or enterprise SaaS integrations.
  • Familiarity with knowledge graphs, ontology design or governed memory.
  • Early-stage startup or founding-engineer experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:45 min

Fusing developer experience and platform engineering for agentic SDLC

Julia Kordick Julia Kordick · World Congress 2026 Europe

4:46 min

Updates on developer events and introducing the Shopify platform

Chris Heilmann +2 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:24 min

Building client-facing AI agents for engineering teams

Alfonso Graziano Alfonso Graziano · Coffee With Developers

5:01 min

Fueling major tech companies and the open source ecosystem

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

Videos

See all

Related articles

See all