Engineering
Notes from the team building Favora: the crawlers, the catalogue, and the models that make fashion discovery personal. Real systems, live numbers.
2M+ product images in the catalogue
5,148,711 pages crawled in the last 30 days
366,242 items catalogued, one schema
100+ brands crawled, enriched and indexed
Writing
The Favora dataset
Two million product images, 366,242 catalogued items, five million pages crawled every month. How we turn the fashion internet into one machine-readable catalogue, and why we build it ourselves.
Read the piece
The stack
Python and Rust pipeline workers on Temporal. Postgres for the catalogue, Typesense for keyword, faceted and vector search. An LLM labels and embeds every item, gated on a hand-labelled gold set. SvelteKit on the web, Expo in the app.
Working on consumer AI, search, or crawling at scale? We are hiring. For everything else: humans@favora.ai