Site Reliability Engineer
Moonpay
Municipality of Madrid, Spain
4 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
EnglishJob location
Remote
Municipality of Madrid, Spain
Tech stack
Github
PostgreSQL
Load Testing
Node.js
Octopus Deploy
Platform as a Service (PAAS)
Performance Tuning
Redis
Reliability Engineering
Next.js
TypeScript
Virtual Machines
Datadog
Data Logging
Google Cloud Platform
React
Delivery Pipeline
Kubernetes
Service Stack
Job description
MoonPay builds the infrastructure that powers the decentralized economy, enabling anyone to buy, sell, and trade crypto using everyday payment methods. The Site Reliability Engineering team provides a resilient, secure production-ready platform, ensuring safe, repeatable deployments and operational excellence., * Increase the resiliency and reliability of our current PaaS solution, including improving infrastructure as code maintainability.
- Build dashboards, monitoring, and alerting mechanisms with Datadog.
- Perform load testing and performance tuning of production services.
- Lifecycle and maintenance of Kubernetes clusters.
- Implement new technologies on Kubernetes to scale with business needs.
- Develop and integrate automation solutions to improve reliability and simplify recovery.
- Design and track metrics for site uptime and performance, ensuring high visibility.
- Own deployment pipelines and continuously improve monitoring and alerting capabilities.
- Collaborate closely with other engineering functions to provide timely feedback.
- Support the engineering team in delivering software faster and more safely.
Requirements
- Strong system administration skills and familiarity with containers, virtual machines, and Linux terminals.
- Platform engineering or SRE experience at a fast-growing tech company.
- Experience with the tech stack or ability to cross-train quickly (TypeScript, Node.js, React, GCP, Postgres, Redis, Datadog, ArgoCD, Kubernetes, GitHub).
- Experience working in a regulated industry.
- Ability to guide developers on monitoring, logging, and scale.
- Experience on complex projects and collaborative work with Security, Data, and Engineering teams.
- Drive reliability and recovery processes and understand best practices in reliability engineering.
Technology stack
- TypeScript, Node.js, React, Next.js, Google Cloud Platform, Postgres, Redis, BullMQ, Datadog, ArgoCD, Kubernetes, GitHub.
Benefits & conditions
- Competitive salary package.
- Equity package and equity bonus for high performance.
- Unlimited holidays with autonomy to choose working days.
- Hybrid working schedule - remote or office.
- Private health-care benefits.
- Enhanced parental leave.
- Annual training budget.
- Home-office setup allowance.
- Remote working allowance for utilities.
- Monthly budget for MoonPay products and fee-free crypto transactions.
- Employee referral program.
- Regular remote company off-sites and hackathons.