← All projects

Python · Scrapy · FastAPI

Walmart Catalogue & Pricing Scraper

Collects product, price, availability and review data for chosen categories and sellers, normalises it and serves it through a private API and scheduled CSV exports.

PythonScrapyFastAPIPostgreSQLDockerAWS
Collected catalogue with normalised fields
Crawl, clean and serve pipeline
Run report: coverage and failures

Screens are illustrative layouts of the product, not client screenshots.

Overview

What it is

A data-collection service for Walmart marketplace listings. It crawls chosen categories, keywords and sellers, cleans and de-duplicates the results and serves them through a private API and scheduled exports.

The problem

Why it was needed

The client needed catalogue and pricing data in a consistent shape, but raw page data was messy, changed often and arrived with duplicates across crawls.

Normalisedtyped, consistent fields across runs
De-duplicatedno repeats between crawls
API + CSVdata ready for BI or repricing

Features

What We Built

Targeted crawls

By category, keyword, seller or product list.

Normalisation

Consistent price, availability and attribute fields.

De-duplication

Stable product keys across runs.

Private API

FastAPI endpoints with filtering and pagination.

Scheduled exports

CSV drops to storage on a schedule.

Run reports

Coverage, failures and timing for every crawl.

Architecture

How It Fits Together

  1. Crawl definitionscategories, sellers, keywords
  2. Scrapy spiderscollection and retries
  3. Cleaning pipelinetyping and de-duplication
  4. PostgreSQLproducts and prices
  5. FastAPIprivate data API
  6. Exports + reportsCSV, run summaries

Spiders only collect; a separate pipeline cleans and validates, so bad data is caught before it reaches the API.

Roadmap

From Idea to Launch

A typical delivery roadmap for a project of this kind.

  1. 1

    Scope

    3-5 days
    • Target categories, sellers and fields
    • Volume and freshness requirements
    • Review terms and allowed use
  2. 2

    Spiders

    1-2 weeks
    • Category and keyword crawlers
    • Retry and throttling
    • Pagination handling
  3. 3

    Pipeline

    1 week
    • Cleaning and typing
    • De-duplication keys
    • PostgreSQL schema
  4. 4

    API and exports

    1 week
    • FastAPI endpoints
    • Scheduled CSV exports
    • Auth for API clients
  5. 5

    Operations

    1 week
    • Run reports and alerts
    • Docker and AWS deployment
    • Handover

Want something like this?

Tell us what it should do and we'll come back with a stack, a scope and a timeline.