Case studies · Case study 07 · Data architecture · AI

A 2.9-million-product catalog built from merchants' messy feeds, and a public MCP server that lets Claude and ChatGPT search it.

A marketplace for local commerce imports the product feeds of thousands of merchants across about fifteen markets and normalizes them into one catalog, searchable by market and language. In 2026 that catalog was exposed through a public Model Context Protocol server: an AI assistant queries it directly, with no API key and no custom integration. This is how it is built, what broke, and what I would change.

2.9 Mproducts in the unified catalog, measured in August 2026
~15markets and languages, each with its own catalog view
6MCP tools: search, suggest, categories, cities, merchants, merchant detail
0API keys needed: public, read-only, any assistant can call it

Engagement sheet

Role

Founder and architect; a team of four developers across front-end, back-end, Android and iOS

Stack

In-house PHP framework, MongoDB for storage and full-text search, Redis for caching, Python jobs for cleaning and exports

Sources

Merchants' XML feeds in Google Shopping style, with every variation in encoding, entities and missing fields that implies

Cadence

Automated import cycle about every ninety minutes; monthly re-import per merchant on the full plan; exports to Google Shopping and Facebook feeds

MCP server

Streamable HTTP transport, protocol 2025-03-26, tools only, no authentication, catalog scoped per market

Status

In production, used daily by merchants and by AI assistants

Context

Merchants do not write product data for machines. A feed arrives with HTML entities inside titles, prices without currency, categories in the merchant's own words, services and digital goods mixed with physical products. Multiply that by thousands of merchants in fifteen languages and the catalog is only as good as the pipeline that cleans it.

Problem

Three problems at once. Ingest dirty feeds without a human in the loop and without losing items. Keep one catalog that still behaves as fifteen national ones, because a shopper in Lisbon must not see Tokyo prices. And, from 2026, make the whole thing usable by language models, which are the new front door for product search, without building a separate integration for each assistant.

What I did

  • A tolerant parser: feeds are downloaded, cleaned of broken entities, forced to UTF-8, validated and parsed item by item, with field fallbacks in Google Shopping style (title, category, image, price, currency derived from the locale when missing, shipping cost, type). Nothing is rejected for a bad character.
  • Catalog scoped per market and language from the start: one database per site, categories mapped onto each merchant's own tree, keywords derived from titles, text indexes on title, content and author.
  • A daemon with locks and timeouts so that imports never overlap and a stuck job is killed rather than queued; the same cycle produces the Google Shopping and Facebook exports.
  • A mass reclassification job that reads items in batches and sorts them into products, services and digital files with weighted multilingual rules (English, Italian, German, Spanish, French), logging and resolving identifier collisions.
  • Language-model agents in production for the parts rules cannot do: cleaning catalogs, preparing and publishing content on Facebook and Instagram, daily operational reports.
  • A public MCP server with six tools: full-text search with filters by country, city, category, price and delivery, autocomplete, category and city lists with counts, merchant search with a fallback on OpenStreetMap, merchant detail. Read-only, no keys, so that any assistant can use it on day one.

Result

About 2.9 million products from thousands of merchants, searchable per market in about fifteen languages, kept fresh by an unattended cycle. The MCP server has made the catalog a tool that Claude, ChatGPT or any compatible assistant can call directly, without a developer on either side. For a small team, that is distribution bought with architecture instead of with marketing.

What I would do differently

  • Classify at import, not afterwards. Products, services and digital goods were told apart too late, and the mass reclassification had to handle collisions that a stricter first pass would have avoided.
  • Make the market explicit in every call. The MCP server has a default market, and it is almost never the one the assistant wants; the tool descriptions now say so, but the right design is a required parameter.
  • Full-text search on the database's own indexes was the right call at one million products and is a trade-off at three. The next step is a dedicated search engine, chosen before the catalog forces the decision.

Get the PDF by email

Want this case as a PDF, to keep or forward?

The PDF arrives by email within a minute, as an attachment. No calls, no automated sequences.

Case studies

A catalog, a data pipeline or an AI integration that has outgrown the team that built it?

Write me ten lines about the data, the volumes and where it hurts. Within 48 hours you get a free preliminary opinion: what to keep, what to rebuild, with which budget.