> Markdown version of [/jobs/ext/2733536-senior-platform-engineer-observability](https://www.wearedevelopers.com/jobs/ext/2733536-senior-platform-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Engineer - Observability - **Company:** Capital Group - **Location:** Charlotte, NC, United States - **Experience:** Expert - **Salary:** $136,749.0 - $218,798.0 - **Contract:** Temporary contract - **Skills:** Java (Programming Language), .NET Framework, Artificial Intelligence, Amazon Web Services, C Sharp (Programming Language), Software as a Service, Cloud Engineering, Information Systems, Continuous Integration, Data Cleansing, Python (Programming Language), Open Source Technology, Scrum Methodology, Reliability Engineering, Prometheus, Software Engineering, SQL Databases, TypeScript, Datadog, Data Logging, GitHub Copilot, Grafana, AI Platforms, Kubernetes, Information Technology, Cloudwatch, Terraform, Splunk, Dynatrace - **Published:** September 5, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/18214840?backUrl=%2Fcareer%2F18214840%2FSenior-Platform-Engineer-Observability-North-Carolina-Charlotte ## About the Role * You have a bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent technical experience * You have at least 8 years in Software Engineering, Platform Engineering, SRE, or Cloud Engineering, with a strong software background designing, building, and operating production-grade systems at scale * You have strong software engineering fundamentals, along with a solid foundation in networking, security, and cloud-native architectures, and can apply them to solve complex platform engineering challenges at scale. * You understand the trade-offs between reliability, performance, and cost, and apply sound engineering judgment when designing platform solutions. * You have deep, hands-on observability platform experience-instrumenting applications and building telemetry pipelines and backends across metrics, events, logs, and traces (MELT), grounded in OpenTelemetry standards with a vendor-neutral, portable design mindset-using modern tooling such as a leading SaaS APM/metrics platform (e.g., Datadog), open-source stacks (Prometheus, Grafana, OpenTelemetry Collector), and log platforms (e.g., Splunk, CloudWatch) * You have proven experience delivering unified, single-pane-of-glass observability-correlating metrics, events, logs, and traces (including trace-to-log correlation and consistent tagging/context propagation) so engineers move from signal to root cause in one experience without switching tools * You are an expert in at least one of Python or Go (our primary languages for platform tooling, collectors, and automation) and are comfortable reading and instrumenting services in additional languages such as Java, .NET (C#), and Node.js/TypeScript; strong telemetry query-language skills (e.g., PromQL, SQL, or equivalent) are expected * You develop with AI as part of your craft-using AI-assisted coding tools (e.g., GitHub Copilot or equivalent) and agentic workflows to accelerate development, testing, and refactoring of platform tooling-while applying sound engineering judgment to review, validate, and secure AI-generated code in line with approved enterprise AI platforms and controls * You have strong Infrastructure-as-Code skills (Terraform/OpenTofu) and deliver Observability-as-Code-monitors, dashboards, and alerts as versioned, reusable modules deployed through CI/CD-plus telemetry collection and pipeline engineering (collector/agent fleet management, data hygiene and tagging standards, and cost/cardinality/sampling optimization) on cloud-native AWS (EKS/Kubernetes, containers) * You define reliability and observability standards-golden signals, distributed tracing/context propagation, SLIs/SLOs, structured logging, RUM/synthetics-and apply AIOps/anomaly detection to accelerate detection and root-cause analysis, while leading technical workstreams independently, mentoring engineers, and communicating clearly with technical teams and stakeholders (Agile/SCRUM) ## Description This is a hands-on senior engineering role. You'll engineer telemetry collection and pipelines, automate onboarding so teams get observability out of the box, and deliver Observability-as-Code-monitors, dashboards, and alerts managed as versioned, reusable modules. A central goal is a true single pane of glass: you'll correlate metrics, events, logs, and traces into one unified experience so engineers can move seamlessly from signal to root cause without switching tools. You'll integrate with leading observability backends (for example, a SaaS APM/metrics platform such as Datadog alongside log platforms and open-source stacks like Prometheus and Grafana) while keeping the architecture standards-based and portable, so we're never locked to a single vendor. You'll define golden-signal and SLI/SLO standards, tune cost and cardinality, advance AIOps and anomaly detection, and coach engineering teams to raise observability maturity across the organization. You'll also develop with AI-using AI-assisted coding tools and agentic workflows to build and refactor platform tooling faster, while keeping quality and security high. ## Related Videos - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Effortlessly Scale Prometheus With The Telemetry Data Platform – And Keep your Grafana Dashboards, Too!](https://www.wearedevelopers.com/magazine/3-effortlessly-scale-prometheus-with-the-telemetry-data-platform-and-keep-your-grafana-dashboards-too) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)