(Senior) Site Reliability Engineer in Berlin or Konstanz
KNIME AG
Berlin, Germany
about 1 month ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on de.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Shift work
Job source
Tech stack
Java (Programming Language)
Amazon Web Services
Amazon Elastic Compute Cloud
Amazon S3
Microsoft Azure
Software as a Service
Computer Programming
Relational Databases
Linux
Distributed Systems
Identity and Access Management
Python (Programming Language)
+19 more
PostgreSQL
Routing
OAuth
OpenID
Reliability Engineering
Software Deployment
Data Logging
Scripting
Load Balancing
Cloud Platform System
Okta
Firewalls (Computer Science)
Amazon Virtual Private Cloud (VPC)
Cloudformation
Amazon Relational Database Service
Kubernetes
Data Analytics
Cloudwatch
Terraform
Job description
- Using code to automate the deployment and operations of large scale SaaS systems. Experience with Kubernetes operators is a plus.
- Building out infrastructure as code using tools such as Helm, Terraform, Amazon CloudFormation and Azure ARM.
- Participating in on-call rotations, incident triage and mitigation, troubleshooting issues in live environments, providing root cause analysis and issue resolution.
- Setting standards for product deployments including reliability, scalability, traceability and monitoring. Communicate with product and development teams to help drive adoption.
- Instrument deployed systems for performance, reliability and cost effectiveness.
- Embeds with product and engineering to lead planning, own dependency risk, and drive consistent adoption of reliability and operational standards across teams., * Platform stability at scale: The multi-tenant SaaS platform maintains its availability SLO across all tenant tiers as the infrastructure grows from its current state toward a globally distributed commercial release.
- Operational excellence embedded in engineering: Every team shipping to the platform follows a shared production readiness standard, reducing escaped defects and repeated incidents.
- Automated, scalable operations: Tenant onboarding, deployments, and incident remediation are pipeline-driven, eliminating manual toil and enabling the team to scale without growing headcount at the same rate.
- Commercial readiness: The platform meets enterprise security and compliance gates, supports consumption-based metering, and can sustain the onboarding velocity required for a commercial launch without operations becoming the bottleneck.
What we offer
- Purpose-driven impact: The opportunity to be a driving force and cloud advocate as we design and build our next generation of service offerings
- Craft & collaboration: Work alongside experienced engineers in a systems-first culture that values simplicity, maintainability, and clean design.
- Learning: Continuous growth through hands-on challenges, peer exchange, and exposure to cutting-edge AI and data analytics topics.
- Flexibility, health and wellbeing: Hybrid working, flexible hours, subsidised sports or yoga courses, physiotherapy, and flu shots at select locations.
Requirements
- You hold one or more current certifications on a cloud platform such as AWS, Kubernetes, Linux or similar technologies
- Strong cloud experience with at least one among AWS and Azure cloud providers. The ideal candidate would master both. You possess in-depth knowledge of VPC, IAM, EKS, ECR, EC2, S3, RDS, CloudWatch and their counterparts in the Azure environments.
- Have experience deploying software systems to a Kubernetes environment. Have a working knowledge of Kubernetes concepts and the ability to craft deployment solutions using common Kubernetes patterns.
- Scripting knowledge in Python, Shell are required, additional programming experience in Go, Java are a plus.
- Systems level knowledge and experience with Linux. Expertise in networking, including security, routing, load balancers, and firewalls.
- Knowledge of best practices around service telemetry, including metrics aggregation, distributed logging, and tracing in large, distributed systems
- Working knowledge of OAuth/OIDC identity providers such as Keycloak
- Working knowledge of relational databases such as Postgres
- Ability to work independently and within a team environment. This includes clear and concise communication across an organization that is geographically and culturally dispersed
Benefits & conditions
Pulled from the full job description Flexible schedule
About the company
KNIME is a leading AI platform that enables organisations to make sense of their data through intuitive, scalable, and collaborative data science. We empower data professionals and business users alike to build, deploy, and manage AI and data workflows that drive better decisions. Hundreds of global enterprises use the KNIME platform including Citi, Bosch and P&G.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on de.indeed.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this roleβ¦