> Markdown version of [/jobs/ext/232893-ai-support-engineer-application](https://www.wearedevelopers.com/jobs/ext/232893-ai-support-engineer-application). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Support Engineer- Application - **Company:** Accrete, Inc. - **Location:** Indianapolis, IN, United States - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Batch Processing, Databases, Data Validation, Database Queries, Software Debugging, Monitoring of Systems, Issue Tracking Systems, Runbook, SQL Databases, Datadog, Grafana, Integration Frameworks, Database Monitoring, Cloudwatch, Kibana, Data Pipelines - **Published:** May 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=561678a9da4290a5 ## About the Role Do you have experience in SQL?, Must Have * Strong understanding of APIs and HTTP status codes * Experience with monitoring tools/logs (Datadog, CloudWatch, Grafana, Kibana, etc.) * Basic knowledge of SQL (queries, data validation checks) * Ability to work with dashboards, alerts, and incident tracking systems * Experience in incident management / production support environments Good to Have * Exposure to AWS services (CloudWatch, Lambda basics, etc.) * Understanding of data pipelines and batch processing systems * Familiarity with observability tools and logging systems Behavioral Competencies * Ability to stay calm under pressure during incidents * Strong communication and coordination skills * High level of ownership and follow-through * Ability to work in a 24x7 support environment with rotational shifts ## Description We are looking for an Application Support Engineer (L1/L2) to ensure the stability, reliability, and smooth functioning of our production systems. This role acts as the first line of defense for system monitoring and incident response, ensuring that issues are identified early, resolved quickly, and escalated appropriately. The ideal candidate should be comfortable working in a high-availability, fast-paced environment, handling alerts, monitoring data pipelines, and ensuring seamless platform operations., Monitoring & System Health * Monitor production systems using tools such as Datadog, CloudWatch, and internal dashboards * Track system health across APIs, data pipelines, databases, and third-party integrations * Identify anomalies and validate alerts to reduce false positives Incident Management & Response * Respond to system alerts in real-time (failures, latency spikes, downtime) * Perform initial incident triage and identify impacted components * Execute predefined runbooks and recovery actions (job restarts, retries, etc.) Escalate issues to engineering teams when required Data Pipeline Monitoring * Monitor scheduled jobs and workflows (e.g., Dagster, SageMaker, batch pipelines) * Identify missing, delayed, or failed data processes * Trigger re-runs or escalate issues to relevant teams Third-Party & Vendor Monitoring * Monitor failures in external APIs, proxies, and vendor systems * Coordinate with internal teams for resolution * Track and highlight recurring vendor-related issues Database Monitoring * Perform basic database health checks including: + Connection issues + Slow queries + Replication lag + Storage utilization * Raise alerts for any anomalies Runbook Execution & Documentation * Follow standard operating procedures and runbooks for known issues * Maintain clear logs of actions taken during incidents * Ensure proper closure and documentation of incidents Reporting & Shift Handover * Maintain incident logs and reports * Provide structured shift handovers to ensure continuity * Highlight recurring issues and patterns for further analysis What You Will NOT Be Responsible For (To set the right expectations clearly) * No deep debugging or code-level fixes * No infrastructure changes * No ownership of alert configurations (handled by SRE/Engineering teams) ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Navigating the AI Wave in DevOps](https://www.wearedevelopers.com/videos/853-navigating-the-ai-wave-in-devops) - [Add Location-based Searching to Site with ElasticSearch](https://www.wearedevelopers.com/videos/77-add-location-based-searching-to-site-with-elasticsearch) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 162: AI careers, MCP, AWS best practices & floppy sweaters](https://www.wearedevelopers.com/magazine/571-dev-digest-162-ai-careers-mcp-aws-best-practices-floppy-sweaters) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)