> Markdown version of [/jobs/ext/3593646-techops-engineer](https://www.wearedevelopers.com/jobs/ext/3593646-techops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TechOps Engineer - **Company:** Talon.One GmbH - **Location:** Berlin, Germany (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Command-Line Interface, Software as a Service, Cloud Computing, Continuous Integration, Linux, DevOps, Github, Python (Programming Language), PostgreSQL, Operational Databases, Reliability Engineering, Prometheus, Runbook, Systems Integration, Datadog, Scripting, Google Cloud, Grafana, Agentic-AI, Containerization, Kubernetes, Sentry, Api Design, Golang - **Published:** October 6, 2026 - **Apply:** https://startup.jobs/techops-engineer-talonone-10293045 ## About the Role * 2-4 years of experience in TechOps, DevOps, SRE, Production Engineering, or a similar technical role * Experience working with production systems in a SaaS or cloud environment * Comfortable working with Linux, command-line tools, logs, and monitoring * Experience with scripting or automation and a mindset of "if we do it twice, can we automate it?" * A structured approach to troubleshooting and solving operational problems * Proactive attitude towards improving systems, processes, and tooling * Ability to work collaboratively with engineers across different teams * Willingness to learn and build deeper expertise in production systems and reliability, * Familiarity with core SRE concepts, such as Service Level Indicators/Objectives (SLIs/SLOs) and error budgets. * Experience with observability platforms such as Grafana, Datadog, Prometheus, or Sentry * Experience with Kubernetes and GCP * Experience working with APIs, integrations, CI/CD, or infrastructure automatio OUR TECH STACK * Cloud & Infrastructure: Google Cloud Platform (GCP), Kubernetes, containerized workloads * Infrastructure as Code: Terraform / Helm * Observability & Incident Ops: Grafana, Datadog, Sentry, Incident.io * Databases: PostgreSQL * Automation & CI/CD: Go, Python, Bash, GitHub Actions / CI/CD tooling ## Description Our SRE / Production Engineering team is responsible for keeping Talon.One reliable, scalable, and easy to operate. We work closely with engineering teams across R&D to improve how we monitor, release, troubleshoot, and run our production systems. This is a hands-on role for someone who loves solving production-level problems, automating repetitive work, and building pragmatic tooling to make life safer and easier for the engineers around them. ONCE YOU ARE HERE, YOU WILL: * Eliminate Toil: Identify manual or repetitive operational friction across R&D and build clean scripts, automation, and internal tools to solve it permanently. * Pioneer AI-Driven Operations: Design, build, and integrate AI agents to streamline operational workflows, ensuring proper guardrails, monitoring, and human oversight for safe execution. * Level Up Incident Management: Own and optimize our Incident.io workflows, automation, and integrations. Stay closely engaged with incident response and participate in post-incident reviews to identify friction and turn learnings into improvements to tooling, coordination, and processes, without taking on incident responder responsibilities. * Enhance Observability & System Health: Maintain and refine monitoring, alerting, and dashboards across our observability stack. You will dive into logs, metrics, and production data to investigate operational edge cases. * Optimize Workflows & Runbooks: Partner directly with SRE and R&D teams to identify operational pain points, turning messy procedures into clear, automated runbooks. * Support Core Production Systems: Collaborate with SREs on database maintenance tasks, health checks, and release/deployment workflows where production reliability is impacted.