Data Scientist

Intercontinental Exchange
Jacksonville, FL, United States
15 days ago
Apply on dejobs.org
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Airflow Amazon Web Services Microsoft Azure Big Data Computer Programming Continuous Integration Data as a Services Information Engineering Data Migration File Transfer Federal Information Processing Standards (FIPS)
+25 more
Python (Programming Language) Machine Learning NumPy Operational Databases Query Optimization Reference Data Power BI Cloud Services Tensorflow DataOps SQL Databases Tableau (Software) Parquet Google Cloud Sql Optimization Pytorch Apache Spark Git Pandas Core Data Scikit Learn Information Technology Data Analytics Software Version Control Databricks

Job description

Intercontinental Exchange, Inc. (ICE) presents an opportunity for a full-time Data Scientist to join the Data Analytics team. The team owns the quality, enrichment, and delivery of the property and real estate reference data that powers ICE’s Fixed Income and Data Services products, and increasingly contributes to enterprise artificial intelligence and machine learning initiatives as part of ICE’s AI Center of Excellence. The Data Scientist will work across the full data lifecycle, from profiling and validation through modeling, delivery, and production support.

In the near term, the role centers on ensuring that large real estate datasets, including deed, assessment, and address records, are accurate, well matched, and fit for use in downstream analytics such as home price indices and portfolio insights. Over time, the Data Scientist will also apply their skills to a broader set of AI and machine learning projects across the enterprise. The ideal candidate is a strong, adaptable generalist who is comfortable moving between hands-on data operations and applied model development, applies sound statistical judgment, and takes ownership of recurring deliveries to internal teams and external clients.

This position requires technical proficiency and strong problem solving, along with an eager attitude, professionalism, and solid communication skills. Clear written and oral communication is important, as the successful candidate will interact frequently with data engineering, product, and client-facing teams across the enterprise to meet business goals.

Responsibilities

On any day, the candidate could be doing any or all of the following:

  • Own, validate, and maintain recurring production data feeds and aggregated property and real estate datasets (for example deed, assessment, and address records), confirming data quality and soundness before each internal or client delivery.
  • Build, modernize, and automate SQL and Databricks (Spark) workloads, including converting legacy match and append and record-linkage processes into production-grade automated jobs.
  • Plan and run data migration and platform rollout testing, including home price index and geography changes, quantifying differences between data versions and assessing impact on deliverables and customers.
  • Develop, validate, and interpret statistical and predictive models, and build visualizations that turn analysis into portfolio and market insights.
  • Contribute to enterprise AI and machine learning initiatives within ICE’s AI Center of Excellence, from prototyping through productionizing models and generative AI solutions using frameworks such as TensorFlow or PyTorch.
  • Partner with data engineering, product, and client-facing teams to move validated data into production and to translate business requirements into technical solutions.
  • Communicate methods, findings, and limitations clearly to technical and non-technical audiences, and respond to internal and external client questions on data and methodology.
  • Document workflows and data definitions, participate in code and query reviews, and mentor junior team members.

Requirements

  • Advanced degree preferred (MS or PhD) in a quantitative field such as computer science, statistics, mathematics, or economics, or equivalent experience.
  • Strong programming skills in Python (or R) with core data science libraries (for example pandas, NumPy, scikit-learn), and advanced SQL for profiling, complex joins, and query optimization.
  • Hands-on experience with Databricks and Spark, including building and maintaining scheduled production jobs and modernizing legacy SQL processes.
  • Solid grounding in data quality, validation, and reconciliation, including record matching or entity resolution and quantifying differences between data versions; familiarity with real estate or property reference data (for example deed, assessment, parcel, FIPS, and APN) is an advantage.
  • Experience preparing and delivering recurring data products to clients, for example via secure file transfer, with delivery validation.
  • Familiarity with machine learning and, ideally, generative AI frameworks (for example TensorFlow, PyTorch, or large language model tooling), with interest in applying them to new problems.
  • Working knowledge of cloud data platforms (AWS, Azure, or Google Cloud), big data formats such as Parquet, and data engineering, version control, and CI/CD tools (for example Airflow and Git); experience with BI tools such as Tableau or Power BI.
  • Excellent written and oral communication, with the ability to explain technical concepts to technical and non-technical audiences; experience in an applied, agile research and development environment is a plus.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

1:34 min

Bringing diverse skills to industrial data science roles

Katja Träumner

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all