Build and maintain robust data pipelines to integrate data collection, preprocessing, embedding, and retrieval processes. Ensure smooth and scalable operations across the system to support large-scale data handling and model integration.
Collaborate with Al researchers and data scientists to understand data requirements and deliver high-quality datasets for model training and evaluation.
Build and maintain data architectures
Ensure data is pre-processed, cleaned, and transformed in preparation for machine learning models.
Optimize data workflows to improve performance and reduce latency in Al model training and inference.
Develop and maintain automated ETL (Extract, Transform, Load) processes to support data availability and consistency.
Work with cloud platforms (AWS, GCP, Azure) to manage and scale data storage and processing systems.
Ensure data governance, security, and compliance with data privacy regulations.
Continuously monitor and improve the performance of data pipelines to meet the evolving needs of Al projects.
Hunt on autopilot
New Software Developer jobs in China, in your inbox.
The boar re-runs this search every day and emails you what’s new — scraped straight from company career sites, no recruiter spam. Free.