Data Analyst

AUTNHIVE INC.
United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$70,364.0 - $84,739.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Cloud Storage Continuous Integration Extract Transform Load (ETL) Data Visualization Database Queries Decision Support Systems Virtual Private Networks (VPN) JSON Python (Programming Language) Meta-Data Management
+26 more
Metadata Standards Performance Tuning Raw Data Power BI Security Information and Event Management SQL Databases Systems Integration Tableau (Software) Unstructured Data Workflow Management Systems Data Logging Feature Engineering Large Language Models Grafana Advanced Reports Generative AI Git Data Layers Pandas Pyspark Data Lineage Data Analytics Oracle Cloud Infrastructure Looker Analytics Software Version Control Data Pipelines

Job description

  • Design, build, and maintain scalable data pipelines to collect, clean, transform, and validate structured & unstructured data from multiple internal and external sources.
  • Collaborate with data engineers, cybersecurity teams, AI/ML & LLM/Agent teams, and business stakeholders to define data needs and ensure data accuracy, security, and reliability.
  • Analyze cybersecurity datasets including SIEM logs, CVEs, vulnerabilities, and asset inventories to identify trends, anomalies, and actionable insights that strengthen security analytics.
  • Develop and maintain data models, semantic layers, ontology relationships, lineage documentation, and metadata standards to support governance, interoperability, and transparency.
  • Automate data workflows and scheduling using Airflow, Prefect, or Temporal; support CI/CD processes for data including validation, testing, and version control.
  • Integrate and manage datasets across OCI and hybrid environments, ensuring secure connectivity via VPNs, private endpoints, and cloud storage services.
  • Build and maintain dashboards and BI insights using Power BI, Looker, or Grafana to enable data-driven decisions for leadership and technical teams.
  • Support GenAI/LLM initiatives by preparing datasets, curating data pipelines, performing feature engineering, and enriching metadata to enable high-quality model training.
  • Contribute to AI-driven analytics (AI for BI) by integrating LLMs and Agent-based workflows with BI tools (Power BI, Looker, Tableau) to automate reporting and deliver deeper insights.
  • Continuously improve data quality, documentation, and analytical processes through performance optimization, stakeholder feedback, and best-practice implementation.

Requirements

Do you have experience in Version control?, * 5+ years of experience in SQL & Python (Pandas / PySpark) for data querying, analysis, transformation, and automation.

  • Should have experience working with an AI team and should know how the data are prepared for GenAI solution.
  • Should be able to perform end-to-end feature engineering, including extraction, selection, transformation, and enrichment of raw data to make datasets suitable for training ML/GenAI models.
  • Strong hands-on experience with ETL & orchestration tools such as Airflow, Prefect, or Temporal to build robust data workflows.
  • Expertise in data modeling, cleaning, transformation, and handling structured/unstructured formats, including JSON schema understanding.
  • Good understanding of cybersecurity datasets such as SIEM logs, CVEs, vulnerability & asset inventory data.
  • Experience with OCI or similar cloud platforms, including Object Storage, Autonomous DB, Logging, hybrid integrations via VPN & private endpoints.
  • Knowledge of semantic modeling & ontology-based data relationships, metadata management, and lineage tracking.
  • Experience working in AI-enabled analytics environments (e.g., AI for BI), integrating GenAI/LLM-based Agents with BI tools such as Power BI, Looker, or Tableau to automate insight generation, reporting workflows, and data-driven decision support.
  • Hands-on with visualization tools (Power BI / Looker / Grafana), delivering performance dashboards and insights.
  • Experience working on at least one production-grade GenAI/LLM solution (e.g., RAG, AI Agents, automated workflows) and collaboration with AI teams for data readiness.
  • Familiarity with CI/CD for data pipelines, including testing, validation, and version control tools (e.g., Git, DVC), and experience with Agentic workflows.

Benefits & conditions

$70,363.72 - $84,739.10 a year - Full-time

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · WWC 2024

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:31 min

Essential AI and human skills for future teams

Alexander Weißhaupt Alexander Weißhaupt +1 · WWC 2025

Videos

See all

Related articles

See all