Data Modeler

Mphasis
Berkeley Heights, NJ, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Batch Processing Big Data Computer Programming Information Engineering Data Governance Data Integrity Extract Transform Load (ETL) Data Structures Data Stores Data Visualization Data Warehousing
+25 more
Document Management Systems Distributed Computing Environment Distributed Data Store Distributed Systems Python (Programming Language) Metadata Meta-Data Management MongoDB NoSQL Performance Tuning Software Tools Application Data SQL Databases Data Streaming Database Optimization Apache Spark Data Strategy Event Driven Architecture Information Technology Data Lineage Low Latency Apache Flink Apache Kafka Physical Data Models Stream Processing

Job description

We are seeking a Big Data Modeler to design and optimize data schemas across our modern streaming and NoSQL ecosystem. You will bridge the gap between complex business requirements and high-performance technical execution, ensuring our platforms for real-time ingestion, stream processing, and document storage are scalable and governed

Responsibility: Stream & Batch Modeling: Design conceptual, logical, and physical data models for streaming data platforms using Kafka and Flink Distributed Processing Architecture: Architect data structures optimized for Spark batch processing and in-memory event streaming, ensuring data integrity and low latency. NoSQL Data Modeling: Design flexible schemas, indexes, and aggregation pipelines in MongoDB to support high-speed, unstructured, and semi-structured application data. Data Governance & Metadata: Establish data governance standards, metadata management, and data lineage for both event streams and historical data stores. Performance Optimization: Collaborate with data engineers to tune data models for high-throughput, high-volume workloads and distributed computing constraints Cross-functional Collaboration: Partner with software developers, data scientists, and business analysts to translate business use cases into scalable data strategies.

Requirements

Experience: 5 to 8+ years of experience in data modeling, data warehousing, or big data architecture. Core Technologies: Proven ability to model, design, and work with distributed data ecosystems: Kafka: Designing event-driven architecture, topics, schemas, and event payloads. Flink & Spark: Structuring data for stateful stream processing and large-scale batch ETL pipelines. MongoDB: Document modeling, indexing strategies, and querying nested data structures. Programming: Proficiency in SQL along with coding experience in Python, Java, or Scala. Modeling Tools: Familiarity with modern data modeling and visualization tools (e.g., Hackolade, erwin).

Education: Bachelor s degree in Computer Science, Data Engineering, Information Technology, or a related field Behavioral Skills: Good Communication skills Flexible to rotational shifts, 5 days WFO Team Player Ability to work in a changing environment Strong problem solving and analytical skills Ability to work independently or within a team Manage day-to-day challenges and communicate developmental risks with the technical team

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

2:01 min

Migrating existing applications from MongoDB to Postgres

Nikita Shamgunov Nikita Shamgunov · WWC 2024

3:28 min

Defining big data and machine learning fundamentals

Ayon Roy · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

1:28 min

Building shared Java modules and analyst targeting platforms

Chris Heilmann +2 · LIVE

Videos

See all

Related articles

See all