> Markdown version of [/jobs/ext/3535233-release-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3535233-release-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Release & Reliability Engineer - **Company:** ibg llc - **Location:** United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Apple IOS, JIRA, Configuration Management, Software Quality, Continuous Integration, Linux, DevOps, Disaster Recovery, Elasticsearch, Gradle, Web Portals, Job Scheduling, Python (Programming Language), Log Analysis, Apache Maven, Cisco Nexus Switches, Red Hat Enterprise Linux, Regular Expressions, Release Management, Reliability Engineering, Logstash, Ansible, SonarQube, Management of Software Versions, Web Applications, Backup and Restore, YAML, Delivery Pipeline, Grafana, Change Tracking, Backend, Rate Limiting, Git, Gitlab-ci, Information Technology, Npm(Software), Cloud Migration, Teamcity, Kibana, Software Version Control, Jenkins, Servicenow, Artifactory - **Published:** October 2, 2026 - **Apply:** https://www.dice.com/job-detail/acdbc37d-3f72-4cf8-bc84-f3b25f272977 ## About the Role * 2+ years in release engineering, configuration management, DevOps, SRE or platform engineering, including production releases of server-side services * Rigor about change control: every release documented and reconstructable afterwards, never shipped without a tested rollback path * Disaster recovery: standby environments, replication, backup and restore, and failover you have actually run * Strong Python and shell, enough to own and extend release tooling and automate operations against tool APIs * Deep Linux experience, ideally Red Hat, including diagnosing a failing service on an unfamiliar host * CI/CD and artifact management: Jenkins, TeamCity or GitLab CI, Git, Maven, Gradle or npm, Nexus or Artifactory * Configuration as code: YAML models, environment, region and host overrides, everything in version control * Observability and log analysis: Elasticsearch, Logstash and Kibana, Grafana or equivalent, plus scripting and regular expressions for raw logs * Clear written communication, and comfort working across teams in several countries and time zones * BSc/BA in Computer Science, Engineering or a relevant field Good to have * 5+ years of the above, in banking, brokerage or another regulated production environment * Multi-region disaster recovery at scale, including regional failover of stateful services * Reusable CI/CD and deployment pipeline templates across a large application estate * Automating build and release pipelines for customer facing apps (e.g. iOS, Android, Desktop): code signing, beta distribution and store submission * Cloud migration of on-premises release and deployment processes, infrastructure as code, Ansible, containers or orchestration * Internal developer platforms and self-service release tooling * Code quality gates such as SonarQube, job scheduling and change tracking such as ServiceNow or Jira * Exposure to market data, order routing or brokerage systems ## Description We are looking for a Senior Release & Reliability Engineer to deliver and automate the release and deployment of our backend services: approval, deploy, verification and rollback. More than sixty backend services and web applications are released and restarted across production regions in North America, Europe and Asia, behind our web portal, mobile apps and desktop trading platform. The role also owns disaster recovery for that estate: standby regions that stay current, backups that restore and failover that has been rehearsed. You will The release and deployment team, two engineers today, working with every development team in the group as well as product managers, compliance and the infrastructure team. Responsibilities * Deliver releases and deployments for backend services and web applications, from test through production, to whole regions or single hosts * Run the release request lifecycle with the requesting developer: approvals, rollback path, verification, change records * Execute rollbacks under time pressure, and verify the rolled-back state * Set up disaster recovery: standby regions, replication, backup and restore, failover and failback * Run DR drills and regional failover tests, and keep recovery procedures proven * Verify releases with telemetry rather than by eye: dashboards, log queries, error and latency checks, smoke tests * Onboard new services into the release system, including clustered multi-instance topologies * Automate manual release work: pipelines, deployment and configuration templates, validation, verification harnesses, restart jobs * Bring infrastructure, application and release pipelines under one versioned, testable template language, and retire per-project scripts * Help teams meet readiness standards before new traffic reaches production: monitoring coverage, load-test evidence, rate limiting, staggered restarts * Own build and release infrastructure and access: build jobs, artifact promotion and versioning, release permissions * Provide L2 and L3 support for the release path: off-hours windows, weekend coverage, emergency releases, incidents * Mentor engineers on release practice, and keep release and deployment procedures current