Search

Research Crawling Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Job DescriptionAbout the Role

This is a hands-on engineering role focused on building and operating large-scale web crawlers and data acquisition systems that power dataset creation for frontier AI model training. You will work as an extension of AI research teams, helping to collect, clean, and curate web data at a scale that few organizations can match. The work directly shapes the pretraining and inference pipelines used by some of the most advanced AI labs in the world.

What You'll Do

  • Build and maintain large-scale web crawlers across diverse domains including social media, travel, and multi-language sites.
  • Design high-throughput, fault-tolerant data collection systems capable of handling millions to billions of URLs per day.
  • Navigate anti-bot systems, rate limits, and JavaScript-heavy sites, finding creative solutions when standard protocols fall short.
  • Develop pipelines for cleaning, deduplication, filtering, and normalization of web data at TB to PB scale.
  • Construct and maintain datasets for research and model training in close collaboration with research teams.
  • Monitor crawl performance, coverage, and data quality, iterating quickly as web environments change.
  • Optimize infrastructure for cost, latency, and reliability across cloud and bare-metal environments.

What We're Looking For

  • 3 or more years building and operating web crawlers at scale (2 or more years considered for candidates with a PhD).
  • Proficiency in one or more of: Go, Rust, Python, Java, or C++.
  • Experience running data pipelines at TB or greater scale.
  • Deep knowledge of HTTP, networking, and browser behavior.
  • Hands-on experience with distributed systems or parallel processing.
  • Experience with headless browsers such as Playwright, Puppeteer, or Chrome DevTools Protocol.
  • Familiarity with proxy systems, IP rotation, or request orchestration.
  • Experience with data quality evaluation, scoring, or benchmarking at scale.
  • Experience running crawling or data workloads on cloud platforms (AWS, GCP) or bare-metal infrastructure.
  • Background in NLP pipelines, ML dataset curation, or AI lab work is a strong plus.

Compensation & Benefits

Salary range: $160,000 to $250,000 USD annually. Visa sponsorship is not available for this role.

Location

This role is fully remote. The primary location is Los Angeles, CA, United States, though candidates based in other major US cities are welcome.

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...