Mid Data Analyst

Capco
United States
17 days ago
Apply on job-boards.greenhouse.io
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours
Languages
English

Tech stack

Adaptable Database Systems Big Data Databases Extract Transform Load (ETL) Data Transformation Data Systems Distributed Data Store Python (Programming Language) Machine Learning NumPy SciPy SQL Databases
+6 more
Data Processing Data Ingestion Pandas Pyspark Information Technology Data Management

Job description

We are looking for an experienced Data Analyst to provide services supporting the development of advanced machine learning models within a global financial services environment.

In this role, you will work closely with data scientists, model developers, and business stakeholders to deliver high-quality data solutions that power analytical and modeling initiatives. You will be responsible for sourcing, transforming, analyzing, and preparing large-scale datasets while contributing to the design and implementation of robust ETL processes., * Support the development of advanced financial machine learning models by sourcing, aggregating, and preparing data from multiple internal and external sources

  • Clean, transform, validate, and analyze complex datasets to ensure high-quality inputs for modeling activities
  • Design, develop, and optimize ETL pipelines to streamline data ingestion and processing workflows
  • Work with large-scale databases and distributed data environments to extract and manage data efficiently
  • Collaborate with cross-functional teams including model developers, data scientists, and business stakeholders
  • Identify data quality issues and implement solutions to improve data accuracy, consistency, and accessibility
  • Contribute to continuous improvement of data management and analytical processes, * Opportunity to work on challenging projects within the financial services sector
  • Exposure to modern data, analytics, and machine learning initiatives
  • Collaborative and supportive team environment
  • Access to professional development opportunities and a global network of experts

Requirements

  • 4-5+ years of experience as a Data Analyst
  • Strong proficiency in SQL and Python, including:
  • Pandas
  • NumPy
  • SciPy
  • PySpark
  • Solid understanding of data wrangling, data transformation, and ETL processes
  • Experience working with large and complex databases
  • Strong analytical and problem-solving skills
  • Banking or financial services domain knowledge is highly desirable
  • Bachelor’s degree in Computer Science, Engineering, Mathematics, or a related field
  • Excellent communication and stakeholder management skills
  • Detail-oriented, collaborative, and business-focused mindset
  • Fluent English, both written and spoken

About the company

At Capco Poland, we’re not just another consultancy-we’re the dynamic force driving digital transformation in the financial world. As a global leader in technology and management consulting, we help clients tackle complex challenges across banking, payments, capital markets, wealth, and asset management.

Our culture is fast-paced, flexible, and entrepreneurial. We move quickly, think creatively, and put people at the center of everything we do. We’re passionate about delivering value to our clients while creating opportunities for our people to grow and thrive.

We’re proud to be:

  • Leaders in banking, payments, capital markets, wealth, and asset management consulting
  • Champions of an agile, innovative, and collaborative work environment
  • Committed to attracting and developing exceptional talent

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on job-boards.greenhouse.io
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:54 min

Development history of scientific computation libraries and PyViz tools

Radovan Kavický · LIVE

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:33 min

Refactoring data science workflows using Rapids QDF and Pandas

Paul Graham Paul Graham · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all