Sr Data Engineer - Sports Analytics Platform (Python, Web Scraping, Statistical Modeling) - FT - Worldwide

Arc Full-time
United States
1 day ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Shift work
Job source

Tech stack

JavaScript (Programming Language) Airflow Amazon Web Services Computer Vision User Authentication Microsoft Azure Cron Web Scraping Github Python (Programming Language) PostgreSQL NumPy
+11 more
Parsing Selenium ReactJS Pandas Scikit Learn Statistics Packages Playwright Production Code Data Management Entity Resolution Streamlit Framework

Job description

Jane Cervantes is in direct contact with the company and can answer any questions you may have. Email Jane Cervantes, Recruiter

About the project

An NCAA Division I women’s volleyball program is building a private player-evaluation platform for roster, recruiting and transfer-portal decisions. The core idea: rate players on efficient, high-value production, not the stats that look best in highlights.

The coaching staff has explored the concept but has no production code, so this is a ground-up build. You’ll work directly with the head coach and the team’s analytics lead, who know exactly what they want to see.

This is a full-time, focused engagement. We’re looking for someone who can commit 40 hours a week until it’s delivered, not someone splitting time across projects.

Phase 1: Core platform

  • Scrapers for season stats from about 340 Division I athletics sites. The sites run on multiple platforms, some rendered in JavaScript. Data refreshes at least daily in season.
  • Player matching: one record per player across seasons, schools, transfers and name variations, joined to roster data (position, class, height).
  • A PostgreSQL database of players, teams, conferences and seasons that supports historical comparison.
  • A position-specific value rating (points added per set) with conference-strength adjustments. Staff can change weights without code changes.
  • A secure, mobile-friendly dashboard: rankings, filters, player rating breakdowns, weight controls, and a chart showing undervalued players.
  • Deployment on client-owned cloud accounts, with documentation and a maintenance guide.

Phase 2: Match data

  • Parse licensed VolleyMetrics / DataVolley match files.
  • Build an expected-value model that values every contact by its effect on rally win probability.
  • Merge the results into the dashboard. The data stays on the program’s systems.

Phase 3 (optional, may be a separate hire): Practice movement tracking

  • Camera-based tracking of player movement, such as blocker closing speed and transition timing.
  • Computer vision experience is a bonus, not a requirement.

Requirements

  • 5+ years of experience; strong Python (pandas, NumPy)
  • Production web scraping: requests/httpx, BeautifulSoup/lxml, Playwright or Selenium, including maintaining scrapers when sites change
  • Entity resolution / record linkage on messy data
  • PostgreSQL and scheduled pipelines (cron, GitHub Actions, Prefect or Airflow)
  • Solid statistics and probability (statsmodels or scikit-learn)
  • A dashboard framework (Streamlit, Dash or React) with authentication
  • AWS, GCP or Azure deployment; Git/GitHub
  • Able to own a build end to end and communicate clearly with non-technical stakeholders, * Sports analytics or sports betting experience, especially win-probability or expected-value models
  • Volleyball knowledge, DataVolley/VolleyMetrics, or R
  • Computer vision / pose tracking (for Phase 3)
  • University SSO integration

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Loading talks and stories from around this role…