The TikTok Data Ecosystem Team has the vital role of crafting and implementing a storage solution for offline data in TikTok's recommendation system, which caters to more than a billion users. Their primary objectives are to guarantee system reliability, uninterrupted service, and seamless performance. They aim to create a storage and computing infrastructure that can adapt to various data sources within the recommendation system, accommodating diverse storage needs. Their ultimate goal is to deliver efficient, affordable data storage with easy-to-use data management tools for the recommendation, search, and advertising functions.
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth.
Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals.
Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted.
Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Successful candidates must be able to commit to at least 3 months long internship period.
Responsibilities
- Design and implement unified offline and real-time data architectures for large-scale recommendation system feature storage
- Build streaming lakehouse systems supporting feature pipeline processing, model training, and real-time inference workloads
- Design optimized data formats and high-throughput data loading interfaces for deep learning model training frameworks
Minimum Qualifications:
- Currently pursuing a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Information Systems, or a related technical discipline
- Basic understanding of distributed systems, preferably in storage, stream processing, or ML infrastructure areas
- Proficiency in Java/Scala/C++ programming, with strong debugging and problem-solving abilities
Preferred Qualifications:
- Hands-on experience with Apache Flink or Lakehouse technologies (Apache Paimon, Iceberg, Delta Lake, Hudi)
- Familiarity with feature storage and training data pipelines, and their integration with PyTorch
- Knowledge of columnar file formats (Parquet, ORC, Lance) and their use in ML data loading