Production Support Engineer II - Java / SQL / Automation

Tekshapers Inc
Alpharetta, GA, United States
29 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Build Automation Bash Shell Configuration Management Databases Data Validation Database Applications Relational Databases Database Queries
+21 more
Software Debugging Distributed Systems Identity and Access Management Python (Programming Language) Performance Tuning Windows PowerShell Query Optimization Queueing Systems Reliability Engineering SQL Databases Working Model 2D Enterprise Software Applications Mttr Break Fix Cloudformation Amazon Relational Database Service Information Technology Cloudwatch Api Gateway Amazon Simple Queue Service (SQS) Terraform

Job description

We are seeking a Production Engineer II to support and improve the reliability, performance, and day-to-day operations of critical production systems. This role combines hands-on troubleshooting with automation, strong SQL skills, and operational excellence practices. The ideal candidate is comfortable working in AWS environments, improving observability, reducing toil, and partnering with engineering teams to deliver stable, scalable services., * Provide production support for Java-based services and data-driven applications, including incident triage, root-cause analysis, and remediation.

  • Build automation to reduce manual operational work (alerts, self-healing actions, runbooks, deployment checks, and routine maintenance).
  • Write and optimize SQL for troubleshooting, data validation, reconciliation, and performance analysis; partner with database teams as needed.
  • Improve service reliability through proactive problem management, capacity planning, and performance tuning.
  • Enhance observability (dashboards, logs, metrics, tracing) and improve alert quality to reduce noise and accelerate resolution.
  • Support release/change activities: deployment readiness, rollback planning, post-release verification, and incident prevention.
  • Participate in on-call rotation and lead efforts to reduce recurring issues and mean time to restore (MTTR).
  • Document operational procedures and contribute to continuous improvement (postmortems, corrective actions, standards)., * Reduces recurring incidents through automation and preventative fixes.
  • Improves observability and alerting to shorten time-to-detect and time-to-recover.
  • Delivers measurable operational excellence outcomes (higher stability, fewer escalations, smoother releases).

Working Model / On-Call

  • Participation in an on-call rotation is required.
  • Occasional after-hours support may be needed for high-severity incidents or planned changes.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • 4 to 5 years of experience in production support, site reliability, or operations engineering for enterprise applications.
  • Strong Java fundamentals and ability to debug application behavior using logs, stack traces, and runtime metrics.
  • Strong SQL skills (joins, aggregations, query optimization) and experience working with relational databases in production contexts.
  • Demonstrated experience building automation using scripting (e.g., Python, Bash, PowerShell) and/or CI/CD pipelines.
  • Familiarity with operational best practices: incident management, problem management, change control, and post-incident reviews.
  • Ability to work under pressure, communicate clearly during incidents, and collaborate across teams., * Hands-on experience with AWS services such as CloudWatch, EC2, S3, RDS/Aurora, IAM, Lambda, Systems Manager, EKS/ECS, and/or SNS/SQS.
  • Experience with Infrastructure as Code (e.g., Terraform, CloudFormation) and configuration management.
  • Familiarity with API gateways, message queues/streams, and distributed system troubleshooting.
  • Experience improving operational KPIs (MTTR, change failure rate, availability, incident volume reduction).
  • Financial services or other regulated-industry experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Essential commands for running and testing Terraform configurations

Hennie Francis · LIVE

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

2:48 min

Daily responsibilities and alignment practices for technical engineering leadership

Edoardo Dusi · LIVE

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all