Production Support Engineer

VLINK INC
Bloomfield, CT, United States
8 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Compensation
$128,960.0
Working hours
Regular working hours

Tech stack

Query Performance Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Cloudfront Amazon S3 Application Integration Architecture Application Performance Management Interactive Voice Response Software as a Service Cloud Computing
+41 more
Computer Programming Databases Data Validation Data Synchronization DevOps Design of User Interfaces Human-Computer Interaction Python (Programming Language) Knowledge Management Log Analysis MongoDB Query Optimization Redis Software Tools Cloud Services Runbook Server Administration Software Deployment Software Engineering Systems Integration Amazon Connect Scripting Enterprise Software Applications Cloud Platform System ReactJS System Availability Prompt Engineering Spring-boot Software Troubleshooting Caching Generative AI Indexer Backend Material UI Information Technology Performance Monitor Software Coding Restful APIs Splunk Dynatrace Microservices

Job description

Our Client is seeking a Production Support Engineer who will provide production support for enterprise applications and cloud-based platforms, ensuring high availability, reliability, performance, and stability. The role requires strong troubleshooting and coding skills across React, Java/Spring Boot, Python, APIs, microservices, databases, caching technologies, and AWS. The engineer will work closely with development, infrastructure, DevOps, and business teams to resolve production issues and continuously improve application support and operational efficiency. Your future duties and responsibilities:

  • Provide L2/L3 production support for enterprise applications developed using React, Python, APIs, microservices, MongoDB, Redis, and AWS services.
  • Monitor application availability, performance, and reliability, and proactively identify production risks and service degradation.
  • Having Coding knowledge on React UI, Java/Spring Boot backend service, integration, database, cache, and application performance issues.
  • Support MongoDB operations, including connectivity, data validation, query optimization, indexing, and performance troubleshooting.
  • Monitor and resolve Redis Cache issues related to availability, memory utilization, key expiration, data synchronization, and application connectivity.
  • Support AWS hosted applications and deployments involving Amazon S3, CloudFront, React UI components, and related cloud services.
  • Use Dynatrace and Splunk for log analysis, distributed tracing, alert investigation, performance monitoring, dashboarding, and production diagnostics.
  • Manage production incidents by performing impact assessment, participating in incident bridges, coordinating resolution, completing root cause analysis, and implementing preventive actions.
  • Develop Python based automation for application health checks, log analysis, alert enrichment, operational reporting, and repetitive support activities.
  • Apply Generative AI, prompt engineering, RAG, and AI assisted tools to accelerate incident analysis, generate incident summaries, support troubleshooting, and improve knowledge management.
  • Support application releases by completing readiness checks, validating deployments, monitoring post release performance, and coordinating rollback or remediation activities when required.
  • Communicate incident status, risks, technical findings, and recovery progress clearly to business stakeholders, development teams, infrastructure teams, and leadership.
  • Participate in rotational on call support and continuously improve application stability, support efficiency, monitoring coverage, and incident prevention.

Requirements

  • At least 6 8 years of overall IT experience, including at least 4 years of L2 production support for enterprise and business critical applications.
  • Having coding knowledge for supporting applications developed using React, Python, REST APIs, microservices, MongoDB, Redis Cache, and AWS services.
  • Experience troubleshooting end to end production issues across the UI, backend services, APIs, integrations, databases, caching layers, and cloud infrastructure.
  • Experience designing and building an enterprise alerting framework for proactive monitoring, alert correlation, notification, escalation, and incident prevention.
  • Experience supporting Doctor Tools and clinician facing applications, including application availability, integrations, workflow issues, and production performance.
  • Strong experience managing incident and change queues, including ticket prioritization, assignment, SLA tracking, technical analysis, change validation, stakeholder communication, and timely closure.
  • Experience supporting MongoDB, including connectivity, query performance, indexing, data validation, and production troubleshooting.
  • Working knowledge of Redis Cache, including availability, memory utilization, key expiration, synchronization, connectivity, and performance issues.
  • Practical experience using AI tools, Generative AI, prompt engineering, and RAG for incident analysis, troubleshooting, operational automation, incident summarization, and knowledge management.
  • Experience supporting, migrating, and operating AI enabled IVR and contact center solutions, including Kore.ai, Sierra AI for Health Services, Amazon Connect Outbound, Doctor Tools, and related AI tools, covering integrations, monitoring, incident resolution, production stability, and continuous improvement.
  • Experience supporting AWS hosted applications involving Amazon S3, CloudFront, application deployments, monitoring, and related AWS services.
  • Strong experience using Dynatrace and Splunk for log analysis, distributed tracing, dashboarding, alert investigation, performance monitoring, and root cause analysis.
  • Proven experience managing major incidents, incident bridges, problem management, root cause analysis, corrective actions, and preventive measures.
  • Experience developing Python automation scripts for health checks, log analysis, alert enrichment, operational reporting, and repetitive support activities.
  • Experience supporting production releases, deployment validation, post release monitoring, rollback coordination, runbooks, SOPs, and knowledge documentation.
  • Ability to participate in rotational on call support and communicate effectively with business stakeholders, client teams, development teams, infrastructure teams, and leadership.
  • Education: Bachelor’s degree in computer science or related field., Amazon CloudFront, Amazon Simple Storage Service (S3), Amazon Web Services (AWS), Analysis Skills, Application Hosting, Application Integration, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Caching, Call Centers, Candidate Screening, Change Management, Cloud Applications, Cloud Computing, Communication Skills, Computer Programming, Computer Science, Continuous Improvement, Corrective Action, Data Quality, Database Technology, DevOps, Diversity, Documentation, Enterprise Applications, Establish Priorities, Healthcare, High Availability, High Reliability, Identify Issues, Knowledge Management, Leadership, Memory Hardware, Microservices, MongoDB, On Call, Operational Strategy, Operational Support, Performance Analysis, Problem Solving Skills, Production Management, Production Support, Python Programming/Scripting Language, Query Optimization, React.js, Redis, Reporting Dashboards, Risk Analysis, Root Cause Analysis, Scripting (Scripting Languages), Service Level Agreement (SLA), Software Administration, Software Development, Software Engineering, Splunk, Standard Operating Procedures (SOP), Technical Analysis, Technical Support, Time Management, User Interface/Experience (UI/UX), Voice Response Systems

About the company

VLink, founded in 2006, is a leading global provider of software engineering services with next-gen technologies and best-in-class talent. Our Headquarters are in the U.S, and we have offices in 7+ countries from North America-Europe to APAC, with expansion plans in the Middle East. With over 1,000 employees working globally, VLink has helped SMBs, and large enterprises achieve their business goals, and gained the trust of Fortune-250 companies. VLink is ‘Great Place to Work CertifiedTM’ and has been a consistent winner as- Best Places to Work in CT. Trust, collaboration, and accountability are the three elements that are at the core of VLink’s work culture. We value our professionals, providing comprehensive benefits and the opportunity for growth., Started in 2006, VLink has built a solid foundation of providing end-to-end project delivery services, IT services, and talent acquisition solutions to various clients of all sizes -from small, medium to large Fortune 500 companies. We have a stellar history of providing continuous and superior quality services to our customers. This is a testament to the quality of our employees, who are our greatest asset. We believe providing our employees with the tools, training, and processes will not only improve their skills but provide a high caliber of service to our clients. This employee-centric approach, coupled with our financial stability and retention policies and procedures, has resulted in an employee turnover rate well below industry averages. In addition, over the past twelve years, VLink has created a robust database of thousands of pre-screened candidates who may be actively seeking new opportunities. Presently, VLink has over 400+ employees working on-site at client sites in the United States, and in our offshore delivery centers in India and Indonesia.

Company Size: 100 to 499 employees

Industry: Computer/IT Services

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all