Senior Data Engineer

General Motors
Austin, TX, United States
1 day ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$125,000.0 - $191,500.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Geographic Information Systems Application Programming Interfaces (APIs) Artificial Intelligence Data Analysis Application Performance Management Automation of Tests Microsoft Azure Cloud Computing Code Review Information Systems Computer Engineering
+64 more
Continuous Delivery Continuous Integration Information Engineering Data Security Data Structures DevOps Distributed Computing Environment Distributed Systems Github JSON Python (Programming Language) Key Management Machine Learning Metadata Repositories NumPy Octopus Deploy Object-Oriented Software Development Operational Databases Performance Tuning Reliability Engineering Cloud Services Prometheus DataOps Search Technologies Software Engineering SQL Databases Data Streaming Privacy Controls Azure Service Bus Datadog Data Processing Azure Data Factory Cloud Monitoring Pytorch Retrieval-Augmented Generation Large Language Models Grafana Multi-Agent Systems Prompt Engineering Apache Spark Model Validation Generative AI Pandas Event Driven Architecture Containerization Data Lakes Scikit Learn Integration Tests Kubernetes Infrastructure Automation Frameworks Information Technology Apache Flink Apache Kafka Azure AKS Graphql Data Management Machine Learning Operations Video Streaming Restful APIs Terraform Grpc Data Pipelines Dynatrace Databricks

Job description

This role is categorized as hybrid. This means the successful candidate is expected to report to Austin Technical Center three times per week, at minimum [or other frequency dictated by the business if more than 3 days]., Vehicle Data Engineering is looking for a Senior Data Engineer to design, build, and operate data products that transform connected-vehicle signals into trusted insights, health information, and proactive customer experiences. This role will work across real-time streaming, cloud data platforms, APIs, analytics, and artificial intelligence to deliver secure, reliable, and explainable data capabilities at scale.

The ideal candidate is a hands-on technical leader who can move from architecture to production implementation, improve engineering standards, and partner effectively with product, software, vehicle, cloud, analytics, and data-governance teams.

What You’ll Do

  • Design and develop production-grade batch and real-time data pipelines for connected-vehicle telemetry, trip and session data, diagnostic signals, and vehicle-health indicators.
  • Build streaming applications that ingest, enrich, validate, deduplicate, curate, and publish event-driven data for downstream services, notifications, reporting, and analytics.
  • Develop reliable data products using Apache Flink, Apache Spark Structured Streaming, Java, Python, and SQL.
  • Work with Azure services including Azure Kubernetes Service, Event Hubs, Azure Data Explorer, Azure Key Vault, Azure Databricks, Azure Monitor, and Application Insights.
  • Design and maintain data contracts, schemas, APIs, and event models using GraphQL, REST, gRPC, JSON, and cloud-event patterns.
  • Apply artificial intelligence and machine learning to data engineering problems such as anomaly detection, data-quality triage, predictive health signals, intelligent operations, and engineering productivity.
  • Build or integrate generative artificial intelligence capabilities, including large language model applications, embeddings, vector search, retrieval-augmented generation, agentic workflows, prompt engineering, evaluation, and safety guardrails.
  • Create automated tests, performance benchmarks, integration tests, and validation checks for high-volume data and event-driven systems.
  • Establish observability with OpenTelemetry, Datadog, Grafana, Prometheus, dashboards, monitors, service-level objectives, and actionable alerts.
  • Secure data in transit and at rest and apply privacy, consent, retention, lineage, access-control, and regional compliance requirements to vehicle and location data.
  • Automate infrastructure and delivery using Kubernetes, Helm, Argo CD, Terraform, continuous integration, continuous delivery, and infrastructure-as-code practices.
  • Participate in architecture reviews, code reviews, incident response, root-cause analysis, operational readiness, and on-call support as needed.
  • Mentor engineers, raise technical standards, document design decisions, and contribute to a culture of quality, ownership, and continuous improvement.

Requirements

  • Bachelor’s degree in computer science, computer engineering, data engineering, information systems, or a related technical field, or equivalent experience.
  • 5+ years of professional experience in data engineering, software engineering, distributed systems, or a related field.
  • Strong hands-on experience with Java or Python, SQL, object-oriented design, data structures, algorithms, and automated testing.
  • Experience designing and operating production data pipelines using Apache Flink, Apache Spark, Kafka, Azure Event Hubs, or comparable streaming technologies.
  • Experience with cloud-native development on Microsoft Azure and containerized workloads running on Kubernetes.
  • Experience with Databricks, Delta Lake, distributed data processing, data modeling, and performance optimization.
  • Experience designing APIs and event-driven systems using GraphQL, REST, gRPC, asynchronous HTTP clients, or equivalent technologies.
  • Experience with schema evolution, data contracts, data-quality validation, lineage, observability, and privacy-aware data handling.
  • Demonstrated experience applying machine learning, artificial intelligence, or generative artificial intelligence in a production engineering, analytics, or data-product environment.
  • Working knowledge of large language models, embeddings, vector databases or vector search, retrieval-augmented generation, prompt design, model evaluation, and responsible artificial intelligence practices.
  • Ability to troubleshoot complex distributed systems and communicate technical decisions clearly to both technical and nontechnical audiences.
  • Ability to work effectively in a collaborative, agile, cross-functional environment.

What Can Give You a Competitive Advantage (Preferred Qualifications)

  • Master’s degree in computer science, data science, artificial intelligence, machine learning, or a related field.
  • Experience building multi-agent or agentic systems for data analysis, data operations, engineering support, or customer-facing insights.
  • Experience with Databricks artificial intelligence and machine learning capabilities, MLflow, model registries, feature stores, vector search, or model-serving platforms.
  • Experience designing retrieval-augmented generation systems, including chunking, embedding strategies, hybrid retrieval, reranking, grounding, citation, offline evaluation, online evaluation, and hallucination mitigation.
  • Experience applying generative artificial intelligence to data observability, incident summarization, root-cause analysis, schema mapping, documentation generation, or pipeline remediation.
  • Experience with time-series data, vehicle telemetry, geospatial data, diagnostics, predictive maintenance, anomaly detection, or other high-volume sensor data.
  • Experience with Azure OpenAI Service or comparable large language model platforms and with securing enterprise AI workloads.
  • Experience with Python data and machine-learning libraries such as pandas, NumPy, scikit-learn, PyTorch, or equivalent tools.
  • Experience with Terraform, Helm, Argo CD, GitHub Actions, Azure DevOps, or comparable DevOps platforms.
  • Experience with OpenTelemetry, Datadog, Grafana, Prometheus, distributed tracing, service-level objectives, and production reliability engineering.
  • Knowledge of data catalogs, governance platforms, consent management, privacy engineering, and regional data-retention requirements.
  • Strong technical leadership, mentoring, influencing, documentation, and cross-functional communication skills.
  • Demonstrated initiative, sound judgment, accountability, curiosity, and ability to simplify complex problems.

Benefits & conditions

  • The expected base compensation for this role is: $125,000 - $191,500. Actual base compensation within the identified range will vary based on factors relevant to the position.
  • Bonus Potential: An incentive pay program offers payouts based on company performance, job level, and individual performance.
  • Benefits: GM offers a variety of health and wellbeing benefit programs. Benefit options include medical, dental, vision, Health Savings Account, Flexible Spending Accounts, retirement savings plan, sickness and accident benefits, life insurance, paid vacation & holidays, tuition assistance programs, employee assistance program, GM vehicle discounts and more.

About the company

We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team., General Motors is committed to being a workplace that is not only free of unlawful discrimination, but one that genuinely fosters inclusion and belonging. We strongly believe that providing an inclusive workplace creates an environment in which our employees can thrive and develop better products for our customers.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

1:11 min

Evaluating architectural trade-offs between REST and gRPC

Sakshi Nasha Sakshi Nasha · Europe 2026 Virtual

Videos

See all

Related articles

See all