Data Collection & Web Scraping Intern (6 months)

About

DATA&DATA helps luxury brands understand online market dynamics. We aggregate and analyze large-scale data from across the web to provide actionable insights into the pricing, availability, and visibility of high-end consumer goods.

Our core asset is our data: collected at scale from hundreds of sources, checked, normalized, and turned into analytics for some of the world's most iconic luxury brands. We're a small team, which means data collection work here is concrete and immediately useful.

Job Description

As a Web Scraping & Data Collection Intern, you'll be responsible for feeding and monitoring the data that powers our platform. Concretely, you'll:

  • Configure and run scrapers on new and existing websites using our in-house scraping framework

  • Investigate target websites to identify undocumented or public APIs, and figure out the most reliable way to collect their data

  • Write Python scripts and notebooks to collect, extract, and structure data from a wide variety of sources

  • Run quality checks on collected data: completeness, consistency, detecting when a source breaks or silently changes

  • Run and monitor existing collection pipelines, and flag or fix issues when a source stops behaving

  • Answer ad hoc questions on our database with SQL queries (volumes, coverage, anomalies) for internal and client-facing needs

Preferred Experience

You must be enrolled in a school or university able to provide a convention de stage for the full 6 months.

Must-haves:

  • Working knowledge of Python: requests, playwright (or Selenium), plus the basics (loops, functions, files, JSON, virtual environments)

  • Basic SQL: you can write a SELECT with a WHERE, a GROUP BY, and a simple JOIN without needing to look everything up

  • Comfort reading HTML and using browser dev tools (inspecting the DOM, reading the network tab)

  • Git basics for version control

  • Autonomy: you're comfortable investigating a problem on your own before asking

  • Fluency in English (written & spoken); French is a bonus

Nice-to-haves:

  • Prior experience with web scraping, in any context (personal projects count)

  • Familiarity with anti-bot mechanisms, proxies, or headless browsers

  • Experience with pandas for quick data checks

What you get:

  • Real, messy, large-scale data from day one — the kind you can't get from a course project

  • Genuine technical depth on scraping: reverse-engineering APIs, dealing with sites that don't want to be scraped, keeping collection reliable at scale

  • Autonomy and ownership over the sources you handle

  • A flat structure, flexible hours, casual dress, no bureaucracy

Recruitment Process

  • Initial screening – We review your resume and any additional materials you submit

  • Phone interview – A call to discuss your background, your motivation, and a few basic technical questions

That's it. No take-home assignment, no multi-round process.

Additional Information

  • Contract Type: Internship (Between 5 and 7 months)
  • Location: Paris
  • Education Level: Bachelor's Degree
  • Occasional remote authorized