> Markdown version of [/jobs/ext/497760-ml-research-scientist-deep-learning-transformer-architectures](https://www.wearedevelopers.com/jobs/ext/497760-ml-research-scientist-deep-learning-transformer-architectures). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # ML Research Scientist -Deep Learning & Transformer Architectures - **Company:** MILLENNIUM - **Location:** New York, NY, United States - **Salary:** $150,000.0 - $200,000.0 - **Contract:** Permanent contract - **Skills:** C++ (Programming Language), Encodings, Computer Programming, Cursor (Graphical User Interface Elements), Programming Tools, Information Theory, Python (Programming Language), Machine Learning, Tokenization, Pytorch, Large Language Models, Deep Learning, Information Technology - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4bda2febb3d42220 ## About the Role Do you have experience in Statistics?, This is a long-term research project with significant computational resources. The successful candidate will have a PhD in machine learning or a related field and demonstrated ability to implement Transformer architectures from first principles., * PhD in Machine Learning, Computer Science, Statistics, Applied Mathematics, or a related field with a focus on deep learning * Demonstrated ability to implement Transformer architectures from scratch (not just finetuning pre-trained models) * Deep understanding of attention mechanisms, positional encodings, tokenization strategies, and training dynamics * Expert-level PyTorch skills including custom modules, training loops, mixed-precision, and multi-GPU training * Strong mathematical foundations: linear algebra, probability theory, optimization, information theory * Experience training models at scale (100M+ parameters) * Strong programming skills in Python and C++ for performance-critical components * Self-directed researcher capable of defining and executing a multi-month research agenda * Familiarity with Al-assisted development tools (Cursor, Claude Code) Preferred Skills / Experience * Experience applying deep learning to financial data or time-series forecasting * Familiarity with tokenizatlon approaches for continuous or non-text data * Published research in top ML venues (NeurlPS, ICML, ICLR) or equivalent industry experience * Knowledge of market microstructure and intraday trading dynamics * Experience with model compression, quantization, and inference optimization ## Description We are seeking an exceptional ML research scientist with deep expertise in Transformer architectures and large-scale model training. You will design, implement, and train a custom decoder-only Transformer from scratch -not fine-tune an existing LLM, but build a purpose* built architecture for financial time-series., * Design and implement a custom decoder-only Transformer architecture optimized for tokenized financial time-series data * Develop a novel tokenization scheme for intraday market data: price movements, volume, order flow, and cross-sectional features * Implement efficient training pipelines using PyTorch with mixed-precision training, gradient checkpointing, and multi-GPU parallelism * Design attention mechanisms adapted to financial data: temporal attention patterns, cross-asset attention, and multi-scale representations * Build evaluation frameworks for next-token prediction accuracy, signal quality, and trading performance * Implement inference optimization for low-latency production deployment: model quantization, KV-cache, speculative decoding * Conduct rigorous ablation studies to validate architecture choices and training methodology * Collaborate with the team to integrate model predictions into the live trading pipeline * Document research methodology, experimental results, and architectural decisions ## Related Videos - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [When Value Becomes Programmable: A Developer's Guide to Tokenization](https://www.wearedevelopers.com/videos/100276-when-value-becomes-programmable-a-developer-s-guide-to-tokenization) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [The Best Large Language Models on The Market](https://www.wearedevelopers.com/magazine/319-the-best-large-language-models-on-the-market) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)