Senior Sre/Devops Engineer
Role details
Job location
Tech stack
Job description
V-Thru is looking for a Senior SRE/DevOps Engineer to build and scale our infrastructure and deployment systems as we grow our QCommerce platform.This role will own the reliability, observability, and automation of our production systems, working closely with our engineering team to accelerate delivery while maintaining high availability.You'll have the autonomy to shape our infrastructure strategy and implement modern DevOps practices in a fast-paced startup environment.Core ResponsibilitiesInfrastructure & PlatformDesign, build, and maintain scalable cloud infrastructure supporting our backend servicesOwn CI/CD pipelines and deployment automation, enabling fast and safe releasesImplement infrastructure as code practices for reproducible and version-controlled environmentsOptimize infrastructure costs while maintaining performance and reliability standardsEvaluate and integrate new infrastructure technologies and servicesSite Reliability & MonitoringEnsure high availability and performance of production systems handling real-time operationsDesign and implement comprehensive monitoring, alerting, and observability solutionsLead incident response, conduct root cause analysis, and drive systematic reliability improvementsEstablish SLIs, SLOs, and error budgets for critical servicesParticipate in on-call rotation and own escalation processesDeveloper Experience & AutomationBuild tools and automation that improve engineering team productivity and deployment confidenceStreamline developer workflows from local development through production deploymentImplement automated testing, security scanning, and quality gates in CI/CD pipelinesCreate and maintain documentation for infrastructure, runbooks, and operational proceduresReduce manual toil through intelligent automationPerformance & OptimizationMonitor and optimize application and infrastructure performanceIdentify and resolve bottlenecks in services, databases, and message queuesImplement caching strategies and optimize resource utilizationConduct capacity planning and scaling exercisesBalance cost optimization with performance requirementsSecurity & ComplianceImplement security best practices across infrastructure and deployment pipelinesManage secrets, credentials, and access control systemsEnsure compliance with relevant standards and conduct security auditsStay current with security vulnerabilities and coordinate remediation effortsRequired Qualifications5+ years of experience in SRE, DevOps, or infrastSo ructure engineering rolesStrong expertise with containerization (Docker) and orchestration technologiesProven experience building and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or similar)Deep knowledge of cloud platforms (AWS, GCP, or Azure) and infrastructure as code toolsStrong Linux systems administration and networking fundamentalsExperience with monitoring and observability tools (Prometheus, Grafana, New Relic, Sentry, or similar)Proficiency in at least one scripting/programming language (Python, Go, Bash, or similar)Track record of improving system reliability and team velocityStrong problem-solving skills and ability to work independentlyPreferred QualificationsExperience in startup or high-growth e-commerce/logistics environmentsKnowledge of Kubernetes and service mesh technologiesFamiliarity with Firebase, RabbitMQ, and message queue systemsExperience with database performance tuning and optimizationBackground working with PHP, Go, or Node.Js applicationsExperience implementing canary deployments and progressive deliveryUnderstanding of SRE principles and practices from industry leadersContributions to DevOps/infrastructure open source projectsWhat Success Looks LikeDeployment frequency increases while deployment failures decreaseMean time to recovery (MTTR) for incidents shows measurable improvementInfrastructure costs are optimized without sacrificing reliabilityEngineering team reports improved developer experience and deployment confidenceSystem observability provides clear visibility into service health and performanceAutomated processes replace manual operational tasksOn-call burden is reduced through better automation and reliabilityWhat We OfferOpportunity to build infrastructure foundation for a growing QCommerce platformAutonomy to make technical decisions and shape infrastructure strategyWork with modern technologies and introduce best practicesDirect impact on engineering team productivity and product reliabilityCollaborative environment with experienced engineering leadership#J-*****-Ljbffr
Requirements
escalation processesDeveloper Experience & AutomationBuild tools and automation that improve engineering team productivity and deployment confidenceStreamline developer workflows from local development through production deploymentImplement automated testing, security scanning, and quality gates in CI/CD pipelinesCreate and maintain documentation for infrastructure, runbooks, and operational proceduresReduce manual toil through intelligent automationPerformance & OptimizationMonitor and optimize application and infrastructure performanceIdentify and resolve bottlenecks in services, databases, and message queuesImplement caching strategies and optimize resource utilizationConduct capacity planning and scaling exercisesBalance cost optimization with performance requirementsSecurity & ComplianceImplement security best practices across infrastructure and deployment pipelinesManage secrets, credentials, and access control systemsEnsure compliance with relevant standards and conduct security