Application Support Engineer
Zact Inc.
United States
3 days ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
8 years minimum
Working hours
Regular working hours
Job source
Tech stack
Query Performance
Application Programming Interfaces (APIs)
JIRA
Bash Shell
Cloud Computing
Databases
Database Queries
Monitoring of Systems
Issue Tracking Systems
Intrusion Detection Systems
Python (Programming Language)
Knowledge Management
+17 more
Log Analysis
MySQL
Site Reliability Engineering Practices
SQL Databases
Systems Integration
Transaction Data
Data Logging
Google Cloud
Cloud Platform System
Macros
Cloud Monitoring
Grafana
Build Management
Restful APIs
Zendesk
Dynatrace
Microservices
Job description
- Lead triage of complex, multi-service incidents - from first signal to root cause, with a documented timeline throughout
- Use GCP Cloud Logging, Cloud Monitoring, and distributed tracing to trace failures across microservice boundaries
- Write MySQL queries to investigate transaction, card, and account data issues in production read replicas - following strict change-control at all times
- Own post-incident reviews (PIRs): root cause, contributing factors, timeline, and preventive actions
- Act as the senior technical escalation for banking client issues - calm, precise, and solution-oriented
Monitoring & observability
- Design and maintain the team’s alerting strategy in GCP Cloud Monitoring - setting thresholds, reducing alert noise, and ensuring the right signals fire at the right severity
- Build and maintain dashboards (Grafana, GCP Monitoring, or equivalent) covering transaction throughput, API error rates, service latency, and MySQL query performance
- Define and track SLIs/SLOs for key product flows - payments authorisation, settlement processing, card management - and surface degradation early
- Instrument new services in partnership with engineering: ensure every new microservice ships with adequate logging, structured log fields, and correlation IDs
Back-office tooling & automation
- Help design and build the internal back-office tools and support functions the team needs to operate effectively - this is not off-the-shelf tooling configuration; it includes scoping, designing, and writing the tools from scratch where needed
- Example: an internal investigation dashboard that pulls GCP log context and MySQL transaction data for a given incident ID in one view
- Example: automated log correlation scripts that surface the root cause of a known error pattern in under 2 minutes
- Example: MySQL query templates that auto-populate with a transaction ID and return the full investigation snapshot
- Example: alert-to-ticket automation that pre-populates incident records with log context and affected service
- Write and maintain automation tooling in Python and/or Bash - scripts, schedulers, data reconciliation jobs
- Build internal CLI tools or runbook-linked scripts that junior analysts can invoke safely without senior oversight
- Identify manual reconciliation or reporting tasks that can be replaced by scheduled queries or GCP Cloud Functions
- Own the tooling roadmap for the support function - prioritise what gets built next based on toil volume and incident frequency
Client technical support
- Serve as the senior technical point of contact for all customer-facing product issues that escalate beyond first-line resolution
- Write clear, jargon-free incident communications for banking client operations teams during and after incidents
- Work directly with client technical teams (bank IT, card ops) to diagnose integration issues, file format mismatches, or API usage errors
- Maintain a client-facing known-issues log and service status communication cadence
Documentation & knowledge management
- Build and own the team’s runbook library - step-by-step investigation procedures for every known failure pattern across the product suite
- Write and maintain a MySQL query reference for common support investigations (transaction lookup, card status, spend limit checks, settlement reconciliation)
- Document every new incident type in a known-issues catalogue: symptom, root cause, resolution, recurrence trigger, and prevention status
- Produce onboarding materials that bring a new junior analyst to independent productivity within 6 weeks
- Maintain a living architecture reference showing how services connect, what each service owns, and which logs to check first for each component
Team leadership - hands-on
- Lead a small, global support team - you are a working lead, not a manager removed from the queue; you will handle tickets alongside the analysts based on volume and complexity
- Mentor and develop two junior application support analysts - review their investigations, coach their SQL and log techniques, and build their confidence on live incidents
- Run daily standups and weekly incident reviews across time zones - keeping a distributed team coordinated and nothing falling between shifts
- Define team support processes: ticket triage criteria, severity classification, escalation thresholds, SLA tracking
- Partner with engineering and product to represent the support team’s perspective - surfacing recurring issues, requesting observability improvements, and advocating for supportability in new releases
Technology Environment
- Cloud platform: Google Cloud Platform (GCP)
- Log & monitoring: GCP Cloud Logging, Cloud Monitoring, Pub/Sub; Grafana
- Database: MySQL - production read replicas for investigation; strict change-control for any writes
- Architecture: Microservices - REST APIs with correlation ID tracing across service boundaries
- Scripting & automation: Python (primary), Bash - for tooling, automation, and log parsing
- Ticketing: Internal ticketing system (Jira or equivalent)
- Product domain: Commercial card programs - card issuance, authorization, settlement, spend controls, reporting
- Client environment: Banking clients (community and regional banks) - regulated, SLA-bound, audit-sensitive
Requirements
- 8+ years in application support, platform operations, or a technical operations role with a team leadership component
- MySQL proficiency - writing complex queries, reading execution plans, diagnosing slow queries, understanding index behaviour
- GCP experience - hands-on with Cloud Logging (log queries, log-based metrics), Cloud Monitoring (alerting policies, uptime checks), and ideally Cloud Functions or Pub/Sub
- Microservice troubleshooting - comfortable tracing a failure across 3-5 services using correlation IDs and structured logs
- Python or Bash scripting - you have written automation tools that saved real time, not just one-off scripts
- Methodical incident investigation - you document as you go, form hypotheses, test them, and do not guess
- Strong written communication - incident updates, PIRs, client communications, and runbooks are all part of your normal output
- Experience building or significantly contributing to a team knowledge base or runbook library
Nice to Have
- Experience with distributed tracing tools - Cloud Trace, Jaeger, Zipkin, or equivalent
- Grafana dashboard authoring
- Exposure to SRE practices: SLO definition, error budgets, toil reduction
- Familiarity with GCP Cloud Functions, Cloud Scheduler, or Workflows for automation
- Experience administering Zendesk - configuring ticket views, SLA policies, macros, triggers, automations, and reporting; familiarity with Zendesk API or webhook integrations is a bonus
- Background in a regulated or compliance-sensitive environment - fintech, banking, healthcare, or telecoms
- Familiarity with commercial card concepts (authorisation flows, settlement, BIN management, spend controls) - helpful but fully trainable
Note on Domain Experience
Prior fintech or banking experience is helpful but not a requirement. What matters is depth in GCP, MySQL, and microservice troubleshooting - and the instinct to automate. We provide structured onboarding into the product domain and a growing runbook library to support ramp-up.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on startup.jobs
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
LM
Luis Minvielle
over 2 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
about 1 month ago
LM
Luis Minvielle
Is Software Engineering Over-Saturated?
over 2 years ago
LM
Luis Minvielle
Fully Remote Software Engineer Jobs
over 2 years ago
LM
Luis Minvielle
7 Cloud Computing Trends Coming in 2025 for Developers
over 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago