> Markdown version of [/jobs/ext/2689418-research-engineer-ai-native-online-datastore-systems](https://www.wearedevelopers.com/jobs/ext/2689418-research-engineer-ai-native-online-datastore-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Research Engineer - AI-Native Online Datastore Systems - **Company:** The Client - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $218,400.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Automation of Tests, C++ (Programming Language), Software Quality, Databases, Data Distribution Service, Data Stores, Data Synchronization, Distributed Systems, Python (Programming Language), Machine Learning, Meta-Data Management, Runbook, Software Engineering, Systems Integration, Toolchain, AI Infrastructure, Software Repository, Data Storage Technologies, Real Time Systems, Large Language Models, Multi-Agent Systems, Caching, Build Management, Information Technology, Database Tools and Utilities - **Published:** September 3, 2026 - **Apply:** https://www.themuse.com/jobs/tiktok/senior-research-engineer-ainative-online-datastore-systems ## About the Role Minimum Qualification(s): - Research background (required - one of the following): - PhD in Computer Science, Artificial Intelligence, or a related field; or - Thesis-based Master's degree with a research focus in AI, ML, software engineering, or systems. - Proficiency in Python and Go (or C++ / Java), with the ability to translate research prototypes into production systems; AI-Native coding mindset. - Deep expertise in LLM application engineering: RAG, tool use / function calling, structured outputs, prompt / context engineering; proven track record of engineering research into shipped systems. - Solid distributed systems foundation; familiarity with online storage, caching, data sync, or metadata management is a plus. Preferred Qualification(s): - Peer-reviewed publications or top-tier conference acceptances in agent systems, LLM applications, AIOps, or SE4ML. - Research-grade experience designing agent evaluation, offline benchmarks, or experimentation frameworks. - Multi-agent orchestration and guardrail design for high-risk operations. - Global online storage infrastructure background (multi-region active-active, cross-region sync, compliance governance). ## Description We are TikTok's Online Datastore Systems team, responsible for designing, building, and governing the core online storage infrastructure that powers our global business-including databases, caches, data synchronization, metadata management, storage governance, and global data distribution. Our mission is to deliver online data storage services with ultimate performance, reliability, and intelligence for hundreds of millions of users and countless business scenarios worldwide. The team is undergoing a pivotal paradigm shift: from manual oncall and hand-crafted governance toward AI-native development and operations. We already have early practices in production-an Oncall Agent (knowledge base + runbooks + tiered write-permission model), a storage replica governance Skill, and an Agent Box toolchain-but we need someone to design the AI infrastructure layer from scratch, establish an evaluation framework, and drive org-wide adoption. This is a greenfield opportunity. You won't be maintaining existing AI systems-you'll be the team's first dedicated AI-Native Research Engineer, balancing research exploration with hands-on engineering: studying how agents work best in online data storage contexts, and turning those insights into reusable infrastructure and workflows. Here, you will tackle world-class challenges in globalization, multi-region active-active, compliance, and cost efficiency-while using AI to redefine how storage systems are built and operated. Responsibilities - Design and build the AI Context & Knowledge Layer: Architect a centralized context layer for the online datastore systems team, integrating knowledge bases, runbooks, code repositories, and real-time system state so AI agents have grounded, traceable, team-specific domain knowledge. Continuously iterate on knowledge structure and retrieval strategies to improve answer quality and evidence-chain completeness. - Research and implement Agentic Workflows: Explore agent applications across the development lifecycle (AI-assisted coding, automated testing, PR pre-review, deployment validation) and the operations lifecycle (oncall triage, structured troubleshooting, storage governance task submission). Design multi-agent orchestration frameworks with tiered permission and safety guardrails (read-only / confirm-then-execute / human-execute). - Establish Agent Evaluation & Experimentation: Design offline eval sets and core metrics (hit rate, evidence-chain completeness, misoperation rate). Validate agent effectiveness through shadow mode or A/B experiments. Build a reproducible evaluation methodology to drive continuous improvement. Share findings through tech talks or internal write-ups. - Storage Platform Skill-ification & Toolchain Integration: Encapsulate core online storage capabilities (metadata queries, workflow troubleshooting, storage replica governance, DDL changes, etc.) as reusable Skills. Integrate with gdpa-cli, lark-cli, and real-time query tools. Build standardized AI development environments. - Drive AI-Native Adoption & Enablement: Accelerate team-wide adoption through pairing, workflow demos, architecture reviews, and best-practice documentation. Balance AI-driven speed with code quality and system safety. Serve as the technical advocate for AI-native transformation. ## Related Videos - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [OLAP for AI Applications and why you should care](https://www.wearedevelopers.com/videos/100212-olap-for-ai-applications-and-why-you-should-care) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1146-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)