AWS Database Engineer / Cloud DBA
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+20 more
Requirements
Do you have experience in Technical troubleshooting support?, Weβre seeking a Lead Platform Engineer (Monitoring & Observability) to join a high performing Monitoring Engineering team within a fast paced financial technology organization. In this role, you will apply SRE principles to design, build, and evolve monitoring and observability capabilities that ensure the reliability, performance, and operability of core applications and infrastructure. You will partner closely with application, platform, and development teams to implement data driven alerting, SLO/SLA-based monitoring, telemetry pipelines, dashboards, correlations, and automated remediation. Your work will directly improve system reliability, reduce MTTR, and enhance enterprise wide operational insight. This role requires strong analytical thinking, systems engineering discipline, and a proactive approach to identifying risks, preventing incidents, and driving continuous improvement across the production ecosystem.
Experience 5+ years of experience in software, systems, or reliability engineering roles, with multiple years of hands on experience owning production observability, monitoring, and SLOs in distributed systems., * Deep experience building scalable, reliable monitoring and observability solutions, including instrumentation, alerting, dashboarding, and configuration across large, complex environments.
- Hands on expertise and proficency with modern monitoring and observability tools, (e.g., OpsRamp, Grafana, Elastic, CloudWatch, Azure Monitor BigPanda (AIOps), and strong knowledge of metrics, logs, traces, and OpenTelemetry.
- Strong scripting and programming capability (Bash, PowerShell, and one or more languages such as Python, C-family, or JavaScript) to automate telemetry, alerting, and platform workflows.
- Strong expertise with cloud platforms (AWS and/or Azure) and container orchestration systems (Kubernetes, Docker).
- Deep hands on experience with Elastic Observability (APM, Logs, Metrics, Traces)
- Understanding of distributed systems fundamentals, including networking, security, databases, DevSecOps principles, and performance/capacity engineering.
- Strong communication skills, with the ability to clearly explain complex technical topics to both technical and non technical audiences.
- Exceptional problem solving and troubleshooting abilities, especially in high pressure or time sensitive environments.
- Effective prioritization and multitasking, able to manage competing deadlines while maintaining quality and focus.
- Proven cross functional collaboration, working seamlessly with diverse teams in large, complex IT environments and driving continuous improvement across systems.
Preferred Qualifications
- Experience with CI/CD pipelines and tools like Jenkins, GitHub, GitLab CI, or CircleCI
- Experience querying, manipulating, and visualizing time series data.
- Familiarity with Infrastructure as Code tools (e.g., Ansible, Terraform).
- Knowledge of microservices architecture and event-driven systems.
- Working knowledge of REST APIs, JSON, and ServiceNow.
- Experience with cloud monitoring-particularly AWS or Azure.
Mandatory Skills: Required Skills Deep experience building scalable, reliable monitoring and observability solutions, including instrumentation, alerting, dashboarding, and configuration across large, complex environments. Hands-on expertise and proficency with modern monitoring and observability tools, (e.g., OpsRamp, Grafana, Elastic, CloudWatch, Azure Monitor BigPanda (AIOps), and strong knowledge of metrics, logs, traces, and OpenTelemetry. Strong scripting and programming capability (Bash, PowerShell, and one or more languages such as Python, C-family, or JavaScript) to automate telemetry, alerting, and platform workflows. Strong expertise with cloud platforms (AWS and/or Azure) and container orchestration systems (Kubernetes, Docker). Deep hands-on experience with Elastic Observability (APM, Logs, Metrics, Traces) Understanding of distributed systems fundamentals, including networking, security, databases, DevSecOps principles, and performance/capacity engineering. Strong communication skills, with the ability to clearly explain complex technical topics to both technical and non-technical audiences. Exceptional problem-solving and troubleshooting abilities, especially in high-pressure or time-sensitive environments. Effective prioritization and multitasking, able to manage competing deadlines while maintaining quality and focus. Proven cross-functional collaboration, working seamlessly with diverse teams in large, complex IT environments and driving continuous improvement across systems.
Desired Skills: Preferred Qualifications Experience with CI/CD pipelines and tools like Jenkins, GitHub, GitLab CI, or CircleCI Experience querying, manipulating, and visualizing time-series data. Familiarity with Infrastructure as Code tools (e.g., Ansible, Terraform). Knowledge of microservices architecture and event-driven systems. Working knowledge of REST APIs, JSON, and ServiceNow. Experience with cloud monitoring-particularly AWS or Azure.
Benefits & conditions
Alaska Remote $56 an hour
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on indeed.comGood distractions
Talks and stories from around this role β technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Fully Remote Software Engineer Jobs
Is Software Engineering Over-Saturated?
Data Engineer Salary UK