Reading this
Rows marked HIT are articles the customer really bought
in the target week. The tags name the retrieval strategies that proposed each article:
r1_repurchase exact repeat, r6_product_variant the same garment in
another colour or size, r2b_global and r2_popularity bestsellers,
r5_category bestsellers in a category this customer shops.
B2 is the baseline this project had to beat — a customer's recent purchases topped up with
bestsellers, reaching MAP@12 0.02557 against the model's 0.03296. Customers buy roughly three
articles a week out of 105,542, so single-digit hit counts are the scale of the problem rather
than a defect.
Data. Transactions, customer records and product metadata come from the
H&M Personalized Fashion Recommendations dataset published on Kaggle by
the H&M Group, used here for research and portfolio demonstration only, not for any
commercial purpose. Customer identifiers are the dataset's own anonymised hashes; no real
personal data is displayed. Product names and descriptions are H&M's. This page is an
independent project, not affiliated with or endorsed by H&M.