How we work

Scoped in writing. Shipped in your repo. Checked before delivery.

The same process whether you hire an engineer or order a dataset: a written scope, a sample before the full run, quality gates on every delivery, and a thirty-day fix window.

Process

Six steps, every job.

  1. Scope

    You send URLs and fields. We probe the site, name the protection stack, and reply in writing within 24 hours: doable or not, approach, timeline, price.

  2. Sample

    A real sample from the live site, in your target format, usually within 3 days. You check fields and quality before committing to the full run.

  3. Build

    Scrapers are written in Python (Scrapy, Playwright, Crawlee) in a repo you can read, with the schema from the approved sample enforced in code.

  4. Run and validate

    The full run goes through the quality gates below. Anything that trips a threshold is investigated before data leaves our side.

  5. Deliver

    CSV, JSON, Sheets, a bucket, a warehouse table or a webhook. Each delivery comes with a short run report: rows, duplicates removed, failures, run time.

  6. Fix and maintain

    Thirty-day fix window on every data project. For recurring runs and hired engineers, monitoring and fixes are part of the monthly price.

Quality gates

Checked before it leaves our side.

Row counts alone prove nothing. Every delivery runs through the same checks, and the run report says which ones fired.

Schema test

Every row is validated against the schema from the approved sample: required fields present, types correct, enumerations in range. Rows that fail are quarantined and reported, never silently dropped.

Deduplication

Natural keys (URL, SKU, listing ID) are defined per site and enforced across pages, sort orders and runs.

Row-count variance alert (>20%)

A run that returns 20% more or fewer rows than the previous run or the estimate is held and checked before delivery. It is usually a site change, a block, or a real change in the data; we tell you which.

Empty-field rate alert

Per-column empty rates are compared against the sample. A field that was 2% empty and is now 40% empty means a selector broke or the site is serving placeholders.

Block-rate monitoring

403s, captchas and challenge pages are counted per run. A rising block rate is fixed before it becomes missing data.

Spot checks against the live site

A handful of rows from every delivery are compared by hand against the live page, field by field.

Reporting

You always know what happened yesterday.

Daily: one line in your tracker

What shipped, what is blocked, what is next. Posted at the end of each working day by every engineer on your project. No meeting needed.

// example, not a real runtue · pyle_product_urls: pagination fixed, 4,120 rows · blocked: none · next: detail pages

Weekly: a written report

Sites worked on, scrapers shipped or fixed, rows delivered, blocks handled, plan for next week. Two to four paragraphs every Friday, read by a lead before it reaches you.

// example, not a real runweek 41 · 3 sites live · 2 fixes · 86k rows delivered · 1 new protection (DataDome) assessed

Legal stance

Public data, reviewed per job.

We would rather decline a job than deliver something you cannot use. These rules apply to every quote.

  • Public data only: we extract what the site shows to any visitor. No scraping behind someone else's login or paywall.
  • We review robots.txt and the terms of service for every job before quoting, and tell you if we decline.
  • No personal data beyond what the site displays publicly. GDPR- and CCPA-aware delivery: field minimisation, deletion on request.
  • Request rates sized so the target site is never degraded.
  • Your credentials, your account: login-gated work uses accounts you own and have the right to automate.
  • Code and data are yours. We sign your NDA before any access is shared.

Questions

Process, answered.

What does the daily update look like?

One line per engineer per day in your tracker or chat: what shipped, what is blocked, what is next. No meetings required; if you run a weekly call we join yours.

What is in the weekly report?

Sites worked on, scrapers shipped or fixed, rows delivered, blocks encountered and how they were handled, and the plan for next week. Written, two to four paragraphs, sent every Friday.

What happens if the site changes after delivery?

Within thirty days of a data project delivery we fix and re-run at no charge. After that, or for recurring runs, maintenance is included in the monthly price.

Which tools and languages do you use?

Python for nearly everything: Scrapy and httpx for HTTP, Playwright for browser work, Crawlee where it fits. Apify for hosted Actors. Delivery through boto3, the Google APIs or plain SQL. When you hire an engineer, we use your stack.

Can we see the code?

On a hired-engineer engagement the code lives in your repo from day one. On a data project you receive data by default; the scraper source can be included for a fee quoted up front.

Want to see the process on your site?

Send a URL and the fields. The written scope arrives within 24 hours.