Research Scientist, Data

Added
1 day ago
Type
Full time
Salary
Upgrade to Premium to se...

Related skills

aws sql python hadoop spark

πŸ“‹ Description

  • Own large-scale data pipelines for ML model training
  • Curate and manage diverse multimodal datasets for pre-/mid-training
  • Build scalable data ingestion, labeling, filtering, augmentation, and storage
  • Ensure data quality, privacy, and ethical compliance
  • Optimize data processing for distributed training pipelines
  • Prototype production-ready methods for dataset creation and management

🎯 Requirements

  • 5+ years building data pipelines for ML in research or model training
  • Strong data engineering and ML data curation for multimodal models
  • Experience with distributed data systems (Spark, Hadoop, Ray, etc.)
  • Production-grade data infrastructure for ML pipelines
  • Tools for labeling, filtering, deduplication, QA, dataset management
  • Python, SQL, PySpark; cloud platforms AWS/GCP/Azure

🎁 Benefits

  • Competitive salary and substantial equity
  • Full health benefits, 401k matching, and more
  • Collaborative, mission-driven team with growth opportunities
  • Flexible on-site/remote hybrid (HQ in Palo Alto, CA)
Share job

Meet JobCopilot: Your Personal AI Job Hunter

Automatically Apply to Engineering Jobs. Just set your preferences and Job Copilot will do the rest β€” finding, filtering, and applying while you focus on what matters.

Related Engineering Jobs

See more Engineering jobs β†’